Deployment topology & sizing¶
How the Bedrock components combine into a deployment, how coverage cells map routers to geography, and what actually drives scaling — including a clear account of what needs HA (the Directory, and the web recorders that durably archive position and voice) versus what is relay infrastructure (the routers).
Current deployment versus PETRA 1.0 target
This page's component counts and per-router storage caveats describe current code. PETRA 1.0 statically provisions Relay + Storage roles and requires at least one in each fully capable tactical partition. It does not add election or production HA. The normative role, validation, retention, capacity, and recovery rules are in the Zenoh and DDIL data-plane contract.
PETRA 1.0 DDIL operating states¶
Before a run, record one horizon (default seven days), maximum retained bytes/values and
value size, sustained/peak ingress, reconnect concurrency, provisioned replica count and
placement, disk reserve, and post-expiry recovery procedure. Those values are the capacity
envelope for the lossless-within-horizon claim; the example e2-small host is not a
validated envelope.
| State | Operator expectation |
|---|---|
| Connected | Live traffic, query catch-up, authority refresh, and snapshot generation are available. |
| Core disconnected | A partition with surviving Relay + Storage continues using valid cached authority inside the horizon; no new authority is minted. |
| Alternate relay/storage | Endpoints use a statically configured surviving path; there is no promotion or election. |
| All storage unavailable | Permitted non-retained live peer traffic may continue, but delayed delivery/catch-up is unavailable. Eligible retained content stays queued; LIVE_PUBLICATION is dropped/reported and never enters an outbox. Sent is only upstream put success, not proof of storage. Treat this as degraded, not healthy. |
| Restored within horizon | Reconnect and outbox drain are automatic and idempotent; duplicate storage replies are verified and semantically merged. |
| Restored after horizon | Transient history is visibly incomplete. A cold client first obtains a fresh Web head assertion, then declared durable read models apply the exact Web-signed authoritative snapshot; old key epochs are not required. |
The deployed current-state caveats below remain important until the implementation roadmap is complete: current Server stores are not replicated or backfilled, and ordinary Nodes do not yet host the provisioned storage role.
Components¶
A deployment is built from a handful of process types. Only the router scales horizontally; the Directory is a singleton; remote-api bridges exist only alongside Web.
| Component | Count | Scales? | Notes |
|---|---|---|---|
| Directory | 1 | No — singleton | The identity authority: mints tokens, signs the revocation list, holds the trust root. Single point of failure — there is no documented replication, standby, or HA. If it is down, new logins and token issuance stop; already-cached tokens on routers and clients keep working. |
| Server / router | 1+ | Yes — horizontal | Content-blind Relay + Storage. Carries a set of coverage cells and keeps a bounded durable store-and-forward buffer in RocksDB within the offline horizon. It remains restartable and expendable because it holds no group key or authority-minting key; its stored copies are not a Core source of truth. |
| Gateway | 0+ | Optional | Stateless interop bridge. Connects to one or more routers as a Directory-authenticated client (service IdentityToken) and subscribes. Not on the critical path — the mesh keeps working if a gateway is down. |
| Web / Android / Node | many | — | Clients, each pointed at a static router endpoint (no load balancer or service discovery). Android/Node speak the router's native transport (quic/tls); browsers can't — they reach the mesh only through a remote-api bridge (see below). The web tier also runs server-side recorders that durably archive to its database — the position recorder (web:app/domains/tracks/position_zenoh_recorder.ts) and the voice recorder (VoiceZenohRecorder); these are the one durable record of positions and of voice respectively (both are non-durable in the mesh transport; clients show only a ~10-minute live position trail and hear voice live only). |
| Feed ingestion (NiFi + HTTP seam) | 0+ | Optional | Two ways external producers push data in, neither on the core critical path. NiFi writes AIS/ADS-B/GeoJSON rows into track_hits directly via the least-privilege nifi_ingest Postgres role. Separately, the HTTP force-tracking ingest seam (POST /api/feed/*, HS256 JWT) lets external force-tracking producers push detections/targets into the web tier (see Machine Ingest API). Both are optional — a deployment with no external feeds runs without either. |
| Zenoh remote-api bridges | 1 public live bridge per web, plus 1 private query bridge per router | — | @eclipse-zenoh/zenoh-ts speaks the zenoh-plugin-remote-api WebSocket rather than native Zenoh transport. Caddy exposes the public bridge as browser-facing wss on port 7448; it uses the configured ordered router paths for live traffic. Each Docker-network-only query bridge is pinned to one router, and the Web backend sends retained reads to every one so no router-local store is silently omitted. The bridges are stateless relays of sealed bytes and grant no application authority. |
What needs HA¶
Durability is concentrated in two places; everything else is relay or client-owned. That is the whole reason "make the routers highly available" is not the goal.
- Directory — the one true SPOF. It is the identity authority (token signing, revocation list, trust root) with no replication or standby in source. Down ⇒ no new logins or token issuance — but already-issued tokens on routers and clients keep working, so the live mesh keeps running. This is the component whose availability matters most. Source ships no HA mechanism; design redundancy around it (see Status & Roadmap).
- Web recorders — the position and voice archive. The web tier runs two server-side recorders that durably persist to its database: the position recorder (
web:app/domains/tracks/position_zenoh_recorder.ts) and the voice recorder (VoiceZenohRecorder). Both record streams that are otherwise non-durable in the mesh transport — positions relay through the mesh with clients showing only a ~10-minute live trail, and voice is non-durable in the live mesh transport (no router persists it). The web voice recorder is the only hop that durably archives voice (to SQL, for replay/ORK); so voice is live-only in the mesh but durably archived in the web DB. If position or voice history matters to you, these recorders are the thing worth protecting. - Routers — relays, redundant by federation, not by HA. A router forwards live traffic and keeps short-lived local buffers. For live continuity you don't make one router highly available — you run federated peers on the same cells and let clients reconnect. A client outbox (web/Android) preserves only queue-eligible retained content across reconnect and restart; non-retained LIVE traffic is never queued. After upstream
putsuccess, however,Sentis not proof of storage: losslessness requires that a provisioned Storage copy accepted the message and survives.
The one caveat: per-router durable history is not replicated
A router holds bounded store-and-forward copies in its local RocksDB, and there is no router-to-router storage backfill. The outbox is sender-side: it covers what you send, not what you receive. A client that reconnects to a different peer after a failure sees that peer's retained window, not the dead router's. The Server copy is never the Core authority or an indefinite archive.
Coverage cells¶
Routing in Bedrock is cell-first: every position, heartbeat, drawing, and target is
tagged with the geographic cell it happened in, and the cell is the first segment of the
message address (waypoint/<cell>/...).
- Cells are geohash-5. A cell is exactly five characters from the geohash base32 alphabet (digits plus lowercase letters, minus
a,i,l,o) — enforced byGEOHASH5_PATTERNin the Directory (directory:app/domains/shared/server_token_service.ts). A geohash-5 cell is roughly a few km on a side at mid-latitudes (the create-server runbook uses ~5 km × ~5 km as a rule of thumb). - Deny-by-default ACL. A router's native transport access control denies every peer,
action, and key, then adds only the configured closed subjects and key families. A
ServerTokencoverage cell never creates endpoint publication authority; Common's capability and proof gates remain authoritative for storage and receiving subscribers. - Coverage is a live transport bound. Each
ServerToken.coverage_cellsentry adds the exact cell's ordinary position and heartbeat publication families to the coarse transport allowlist. Registered retained drawing and target transport families span canonical cells so Web can publish and record the COP; their Directory capability, proof, classification, and receiver/storage checks remain mandatory. - 256-cell cap. A single ServerToken may carry at most 256 coverage cells (
MAX_COVERAGE_CELLS = 256,directory:app/domains/shared/server_token_service.ts). The source comment notes this covers ~6,400 km² at the widest geohash-5 cell size (~24 km × ~24 km near the equator). - Manual assignment at enrollment. Cells are chosen by the operator and baked into the ServerToken when the server is registered — there is no automatic geography-to-cell allocation tool. Two routers covering neighbouring areas simply own disjoint cell lists.
PETRA 1.0 also has a separate, connected Directory provisioning set for each endpoint.
Ordinary field endpoints receive only their relationship-derived exact-cell drawing and
target authority. Web's automatic machine profile instead receives Common's registered
all-cell drawing and target grants plus the all-cell position read grant required for the
COP. Neither profile is copied from, inferred from, or widened by a router's
ServerToken.coverage_cells; the ServerToken controls coarse transport reachability while
Directory provisioning controls what the authenticated subject may publish or receive.
Changing an ordinary endpoint's cells requires Directory connectivity. Position and
heartbeat movement use the separately bounded, non-retained LIVE_PUBLICATION rule with a
fixed authenticated opaque source.
Cells are set when you register a server. See Add a server (Step 2 "Decide what to register" and "How geofencing works").
Fog of war — and the trusted web client's exemption¶
Cell-first routing is fog of war: a field device (Android, a drone, a sensor) scopes its geocentric subscriptions to its own cell, so it only receives positions and drawings for where it is. Sensor map objects are drawing values, not a second subscription family. That's deliberate — a captured handset shouldn't expose the whole operating picture.
The web client is the Common Operating Picture. Its server-side recorders
and browser sessions declare only selectors present in their current Directory-issued
authority, including the registered all-cell COP families and exact relationship-derived
channel scopes. It does not subscribe to waypoint/** or receive a certificate-scoped
application-authority exception.
This keeps transport posture uniform: Web, Android, nodes, gateways, and peer routers all match the same generic Zenoh ACL subject. The router ACL admits only the closed family registry, including coverage-cell position/heartbeat publication and the registered all-cell retained families; Directory capabilities and signed envelopes remain authoritative for principal, classification, retained query, and content checks. Optional mTLS is a deployment-wide admission choice, not an HQ/edge distinction.
Topology patterns¶
Single router¶
Smallest possible deployment: one Directory, one router holding all the coverage cells. Every client connects to the same endpoint. No federation, no redundancy.
flowchart TD
D["Directory<br/>(identity authority)"]
R["Router<br/>(all cells)"]
C["Clients<br/>(web / android / node)"]
D -. tokens .-> R
D -. tokens .-> C
C --> R
Use it for: a single site, a demo, or a pilot where one box comfortably carries the whole operating area.
Single-site HA¶
Multiple routers covering the same geography — i.e. identical coverage cells — with clients configured with one ordered static endpoint list. Each router still routes the same cells.
flowchart TD
D["Directory"]
R1["Router A<br/>(cells X, Y, Z)"]
R2["Router B<br/>(cells X, Y, Z)"]
C1["Clients (group 1)"]
C2["Clients (group 2)"]
D -. tokens .-> R1
D -. tokens .-> R2
C1 --> R1
C2 --> R2
R1 <--> R2
Static alternate path, not HA of one box
Two routers on the same cells give live continuity when both are configured: if one fails, first-party clients reconnect through the next ordered locator without application reconfiguration. This is not standby promotion, election, or state replication. Received chat/drawing history is per-router and is not backfilled, so a client reconnecting to the peer sees the peer's history, not the dead router's (see What needs HA). Positions are non-durable in the mesh — routers relay them and clients keep a ~10-minute live trail. Directory-signed ServerTokens validate configured router identities; they do not add or reorder client routes.
Use it for: a single site that wants more than one box live so one failure doesn't take everyone offline — accepting that received durable history is per-router.
PETRA 1.0 demo runout endpoint order¶
The Infrastructure #125 demo profile supplies Android with one device-scoped ordered locator list. Operators stage and verify the exact router identities while connected, then exercise path loss without changing Android configuration:
primary quic/<primary-host>:<port>
primary tls/<primary-host>:<port>
alternate quic/<alternate-host>:<port>
alternate tls/<alternate-host>:<port>
The same list is available after a cold restart inside the offline horizon. Stop the primary to exercise the alternate; stop both to record the documented live/delayed- delivery degradation. This resettable, co-located fixture is not production HA and must not be described as automatic promotion or lossless storage failover.
Multi-site / federated¶
One router per region, each carrying different coverage cells, linked into a mesh by an explicit static peer list. Cross-region traffic relays over the federation hop.
flowchart LR
D["Directory"]
RE["Router EU<br/>(EU cells)"]
RU["Router US<br/>(US cells)"]
CE["EU clients"]
CU["US clients"]
D -. tokens .-> RE
D -. tokens .-> RU
CE --> RE
CU --> RU
RE <==>|"transport.connect<br/>private-CA TLS"| RU
Use it for: a deployment spanning regions where each region has its own router and operators in one region need to see relevant traffic from another.
Federation¶
Routers federate router↔router. The link is not auto-discovered — the operator declares it:
- Static peer list. Each router lists the peers it dials in
transport.connect(server:src/config.rs), a list of Zenoh locators. Add a locator to connect, prune it to disconnect. - Private-CA TLS with application peer identity. Each dialer validates the remote router's server leaf against
transport.tls.peer_ca_pathand presents its own leaf. Strict inbound mTLS is optional and deployment-wide. The Directory-signedServerTokencarries peer identity and coverage authority. Federation does not re-sign application envelopes — payloads relay byte-identical across the hop. - Operator chooses the shape. Because peering is explicit, the operator decides whether routers form a chain, a star, or a full mesh. There is no topology controller.
Sizing¶
What drives scaling¶
There is no single "users per box" number in the source. Scale is driven by three independent pressures:
- Geography → more routers / more cells. Wider or more-fragmented coverage means more geohash-5 cells, and (past the 256-cell cap per ServerToken, or past one box's reach) more routers.
- Redundancy → more routers. Wanting more than one router live for a given area means duplicating its cells onto additional boxes (see Single-site HA) — bounded by the state caveat, not by a config limit.
- Operators / devices → bigger revocation list + Directory load. More principals and devices grow the revocation list every router polls and the token-issuance load on the (single) Directory.
The hard, source-defined limits a deployment runs into:
| Limit | Value | Where |
|---|---|---|
| Coverage cells per ServerToken | 256 (MAX_COVERAGE_CELLS) |
directory:app/domains/shared/server_token_service.ts |
| ServerToken TTL | default 30 days, max 365 days | directory:app/domains/shared/server_token_service.ts (DEFAULT_SERVER_TOKEN_TTL_DAYS, MAX_SERVER_TOKEN_TTL_DAYS) |
| Chat retention (durable, per-router) | 24 h TTL, swept hourly | server:src/store_ttl.rs (CHAT_TTL_MS) |
| Report retention (durable, per-router) | 48 h TTL | server:src/store_ttl.rs (RECORD_REPORT_TTL_MS) |
| Other retained-family horizon | 7 d, equal to the Common outbox horizon | server:src/store_ttl.rs (OFFLINE_RESYNC_HORIZON_MS) |
| Default retained-store bounds | 1,000,000 values / 10 GiB / 1 MiB envelope plus 8-byte frame per value | server:src/config.rs (StateStorageSection) |
| Default concurrent replay / handoff buffer | 8 queries / 16 values per query | server:src/config.rs (StateStorageSection) |
| Revocation-list poll cadence | default 300 s | server:src/config.rs (default_revocation_sync_interval_secs) |
The rack is three mini PCs, not a validated production size
The PETRA small rack runs on three mini PCs on one switch: two Beelink SER5 (the privileged core, and the routers) and one Beelink SER8 with a Radeon 780M for the media/tile servers, which is the only box doing GPU transcode (VAAPI over /dev/dri, not NVIDIA). The split across three machines is the trust model made physical — privileged core, expendable routers, isolated out-of-model asset servers — not a capacity decision. There is no autoscaling and no load balancer, and no production sizing has been validated in source. Treat this as a starting point, not a recommendation.
Browser bridge fan-in¶
Every browser on a deployment funnels through one public remote-api bridge on the web VM (see Components). This is a fan-in worth understanding before sizing the web tier. The private per-router query bridges are not browser endpoints and do not participate in this live fan-in; the Web backend uses them only to send authenticated retained queries to every configured local store.
flowchart LR
B1["browser 1"] -->|WS| BR
B2["browser 2"] -->|WS| BR
BN["browser N"] -->|WS| BR
BR["Bridge<br/>(one runtime,<br/>N sessions)"] -->|1 TLS link| R["Router → mesh"]
Each browser WebSocket is its own Zenoh session inside the bridge, but all N sessions share the bridge's single runtime and its one uplink to the router. That asymmetry sets the performance shape:
- Upstream is deduplicated. Fifty browsers all subscribing
waypoint/<cell>/positionsdeclare ~one aggregated subscription upstream — upstream interest scales with the number of distinct keyexprs, not the user count. - Downstream fans out per browser. Each inbound sample is copied to every matching browser WebSocket, all serialized and written on the single web VM. Downstream cost ≈
samples × subscribers. This is the web tier's scaling chokepoint, and voice is the worst case.
Voice is full-mesh — fan-out scales with channel size, not a mixer
Voice has no server-side mixing or SFU. Each speaker publishes its own Opus stream (waypoint/global/voice/<opaque-channel>/<opaque-principal>/<opaque-session>) at 50 frames/s (20 ms frames), and every participant subscribes to the exact capability-derived channel selector (waypoint/global/voice/<opaque-channel>/**). So in an N-person channel the bridge fans each active speaker's 50 frames/s to the other N−1 browsers: roughly 50 × (concurrent speakers) × (N−1) frame deliveries per second, all on the one bridge / one web VM. Voice is push-to-talk, so steady state is usually 1–2 concurrent speakers; the absolute worst case (everyone keyed at once) is 50 × N × (N−1) — ~19,000 deliveries/s for N=20. Either way voice saturates the web tier long before map traffic does — size (and load-test) the box running web and its bridge against your largest channel and a realistic concurrent-talker count, not against average load. (On the rack that is core-01, which also carries Postgres, the Directory and NiFi — a mini PC, not a validated voice size. See the note above.)
The selector remains wildcard only beneath one opaque channel, so it matches every authorized speaker session in that channel — including the operator's own published key. There is no application-level self-filter; whether a speaker receives its own frames back is governed by Zenoh subscriber locality. For sizing this ±1 is immaterial — treat the local fan as ~N subscribers per channel. Source: web:inertia/features/comms/voice.ts (50 fps), …/transport/subscribers/voice_audio_subscriber.ts (capability-derived per-channel ** subscription).
Single point of failure for realtime. The bridge is per-web; if it restarts, every browser on that web drops its session at once and reconnects (the client retries automatically). The web app itself — HTTP/SSR, login — is unaffected, and the position recorder runs server-side, so the durable position archive never depends on the bridge.
Identity survives the multiplexing. The bridge presents one generic transport session to the router, so the router ACL only gates it coarsely by keyexpr. Per-operator authorization is unaffected because it lives at the app layer: each browser packs its own AuthEnvelope with its own IdentityToken and verifies inbound envelopes itself — the bridge only ferries sealed bytes. The tradeoff: you cannot apply per-browser transport-level ACL, since to the router they are one session.
Scaling the web tier. Vertically, give the web VM more CPU/bandwidth headroom for voice fan-out. Horizontally, run more than one web VM — each gets its own bridge and hostname, and to the mesh each bridge is just another client. There is no shared-bridge clustering; capacity grows by adding webs, not by pooling one.
Worked example¶
A 2-site, ~50-operator deployment: operators split across Europe and the US, with one router per region and a gateway feeding an external system.
Shape:
- 1 Directory — the single identity authority for both sites. Mints all operator/device tokens and signs the one revocation list both routers poll.
- 2 routers, peer-linked:
- Router EU — coverage cells over the EU operating area.
- Router US — coverage cells over the US operating area (a disjoint cell list).
- Each lists the other in
transport.connect, forming a 2-node TLS mesh with the same posture. - 1 gateway — connects to one (or both) routers as a Directory-authenticated service client and bridges to the external system. If it dies, the mesh is unaffected.
- Clients — web and Android operators, each pointed at the router for their region.
How it connects and what federates:
- Before anything, the PKI and the Directory exist, and each box meets its prerequisites — see Before you begin.
- Each router is registered against the Directory with its region's cells — see Add a server (Step 2 sets the coverage cells; Step 3 mints and installs the ServerToken). The EU and US routers get different cell lists.
- An EU operator authenticates against the Directory, receives a token, and connects to Router EU. Their position/heartbeat/drawing traffic is tagged with EU cells and routes through Router EU.
- A US operator does the same against Router US with US cells.
- What federates: the closed global families (chat, voice, channel projections, router liveliness/token, command/ack, records, snapshots, and revocation) plus registered cell-scoped traffic. Ordinary field position/heartbeat publication follows router coverage; Web's Directory-issued all-cell drawing/target publication and position-read profile is the bounded HQ exception. A router's retained store is local and is not replicated to its peer.
Known limits & gaps¶
These are the current reality, not future plans. Plan around them; do not assume an HA mechanism the source doesn't implement. The framing is in What needs HA — these are the sharp edges of it.
- Directory is a single point of failure. No replication, standby, or HA in source. Directory down ⇒ no new logins or token issuance (cached tokens keep working, so the live mesh continues). This is the availability priority.
- Per-router retained history is not replicated. Each router keeps only its bounded offline-resync window in local RocksDB. Failover can therefore lose the dead router's received copies; Core/Web authority and archives provide the longer-lived record.
- No automated failover. Nothing promotes a standby or reroutes clients automatically; client redirection is operator-managed (DNS / static endpoint lists). Live continuity comes from running federated peers, not from HA of one box.
- No autoscaling and no load balancer. The rack is a fixed set of boxes; clients use static router endpoints.
- No per-router capacity SLA. Source defines no "clients per router" or throughput model. Capacity must be measured, not read off a number.
- No multi-region cell-assignment tooling. Coverage cells are chosen and assigned manually at enrollment; there is no algorithm or tool to allocate geohash cells across regions.
For where these sit relative to planned work, see Status & Roadmap.
Verified against server@ab688f0, directory@9c5e565, web@80e3ec2, infrastructure@b3849c0.