Troubleshooting¶
A cross-cutting, symptom-first guide to the things that actually go wrong on a running deployment. The individual runbooks each carry an "If something goes wrong" section for their own step; this page collects the failures that span steps — where the cause is in one component and the symptom shows up in another — so you can start from what you're seeing and work back to the fix.
For each symptom: the likely cause, how to check, and the fix (with a link to the runbook or reference that owns the detail).
A client can't connect¶
The user opens the app or the web client and never reaches the map — a TLS warning, a login error, or a silent hang.
| Likely cause | How to check | Fix |
|---|---|---|
| Certificate not trusted — the client doesn't have the Root CA, or the server presents only its leaf, not the chain. | Connecting shows an "untrusted" / certificate warning before any login screen. | Distribute the Root CA to clients and make sure the server presents leaf + intermediate (not just the leaf). See Set up the certificate chain → If something goes wrong. |
| Hostname mismatch — the server's certificate SAN doesn't list the name the client dialed. | A hostname-mismatch TLS error naming the host. | Reissue the server leaf with the right hostname in the SAN. See Set up the certificate chain. |
| A browser hostname doesn't resolve — the site or the Directory is unreachable. | dig auth.bedrockdefence.com returns nothing on the client path. |
Device-facing URLs use owned names. Rack hosts map them through managed /etc/hosts; off-LAN managed clients need active WARP and the matching Cloudflare private-hostname route. Do not substitute .lan, which is internal-only. See Stand up the Directory. |
| Login fails (web) — the Directory is unreachable from the web container, or the web service token is missing. | Login redirect loops, or fails closed. | Two distinct causes: the rack CA never reached the container, so every HTTPS call to the Directory fails (Node only warns about a missing NODE_EXTRA_CA_CERTS, so the app starts and then fails); or web_service_token_b64 is unset, so the revocation cache never primes. See Stand up the web client → If something goes wrong. |
| User isn't onboarded — "user not found" / access denied. | Login reaches the Directory but is rejected for that person. | Onboard them first. See Onboard an operator. |
| The user was revoked — their tokens are now refused everywhere. | Their Directory page shows a red revoked badge. | Revocation is permanent. If it was a mistake, onboard them again from scratch. See Revoke a principal and A device or operator can't be revoked below. |
The authoritative gate on the client→router hop is server-cert TLS plus app-layer envelope verify — the client presents no client cert. If TLS is clean but messages still don't flow, the problem is upstream of connection: see the next two sections. (See Security model → Transport TLS.)
Messages aren't arriving¶
The client connects and the map loads, but positions, chat, or drawings from other users never show up — or only some do.
| Likely cause | How to check | Fix |
|---|---|---|
| No server covers the field user's cell — routing is cell-first; coverage bounds ordinary field position/heartbeat publication. | Affected field users are operating in an area whose cell isn't in any server's coverage list. Their positions do not appear, while Web's separate HQ all-cell authority may still work. | Add the missing cell(s) to the server's coverage and reissue its ServerToken, then restart the box. See Add a server → How geofencing works and the wire protocol key namespace. |
| The web client has no bridge configured — the browser opens no session. | web_router_endpoints is blank, or the zenoh-bridge container is not running. |
Point the browser at the remote-api bridge, not the router's :7447. On the rack web_router_endpoints is app.bedrockdefence.com and the app expands a bare domain to wss://<domain>:7448, which Caddy proxies to the bridge; re-run playbooks/core.yml. See Stand up the web client → Server / relay. |
The web client points at the router, not the bridge — browser zenoh-ts speaks only the zenoh-plugin-remote-api WebSocket, not the router's native transport. |
Browser console loops 1006 / "disconnected from remote-api-plugin"; web_router_endpoints points at :7447 or the bridge isn't running. |
Use the bridge's wss://…:7448, and confirm the bridge sidecar is up and web_bridge_connect_endpoints reaches a live router. See Stand up the web client → Server / relay. |
A classification gate is dropping the message — retained ingest/replay must lie within the subject, capability, and ServerToken.clearance ceilings; every LIVE receiver separately checks the subject and capability ceilings before decode. |
Server audit shows a retained classification denial, or a receiving endpoint rejects the LIVE value after verification. | For retained traffic, confirm the Server's Classification is high enough; edit the device, Reissue Credentials, and redeploy if needed. For LIVE traffic, inspect the sender capability and receiver clearance rather than treating the native Relay as the gate. See Add a server → Step 2 and Security model → Classification. |
| A device just had its group key rotated — devices pick up the new epoch on their next background poll; a brief window can look like dropped traffic. | Trouble appears right after a key rotation and clears within a few minutes. | Wait for the next poll; the previous epoch stays valid through the bounded grace window. If it persists, see Tokens are rejected. See Rotate keys → If something goes wrong and wire protocol → Group-key rotation and backfill. |
It's a duplicate / replay drop, not a loss — the receive pipeline rejects stale frames (outside the ±60 s window) and repeated (principal, nonce) pairs. |
Only old or re-sent frames are missing; live traffic is fine. | Expected behaviour, not a fault. See Security model → Receive-side gates. |
A server won't federate¶
Two servers are up but won't link and share traffic — users on one relay can't see users on the other.
| Likely cause | How to check | Fix |
|---|---|---|
| Peer certificate not trusted — each router must trust the CA signing the other's server leaf. | One or both dialers reject the peer's certificate at link time. | Configure the shared Root / intermediate as transport.tls.peer_ca_path. See Set up the certificate chain → If something goes wrong and wire protocol → Federation. |
| A server's ServerToken is missing or expired — the box can't authenticate itself into the mesh. | The server starts but rejects traffic, or won't carry some messages. | Reissue the ServerToken from the device's page (Reissue Credentials), redeploy all three files, and restart. See Add a server → If something goes wrong. |
| The box won't start at all — usually TLS certs or the encrypted data-directory mount. | The server process exits on boot; health check on 9090 never answers. |
Check the TLS cert/key paths and the data-directory mount first. See Add a server → If something goes wrong. |
Forwarded IdentityTokens are byte-identical across hops, so the Directory signature stays
verifiable end-to-end — routers never re-sign identity. If federation is up but identity
claims are being rejected across the link, that's a token problem, not a federation one — see
the next section.
Tokens are rejected¶
A connection or message is refused at the identity gate — the token doesn't verify, has expired, or signs under the wrong key.
| Likely cause | How to check | Fix |
|---|---|---|
Token expired — expires_at_ms is in the past (gate 3). |
The principal worked until a deadline, then stopped. Service devices that don't self-refresh need reissuing. | For a service device (server/gateway/web token) Reissue Credentials from its page; operators get fresh passes on next login. See Add a server → Step 3 and Security model → Receive-side gates. |
Clock skew — issued_at_ms is outside the ±60 s replay window (gate 2). |
Fresh frames from one host are rejected as stale/future while others work. | Fix the host's clock (NTP). The window is fixed at 60 s for fresh frames. See Security model → Receive-side gates. |
| Signing key rotated, key set stale — the Directory signed under a fresh key the verifier hasn't cached yet. | Rejections start right after a signing-key rotation; the old key is still served until its tokens expire. | This resolves as verifiers refresh the Directory key set. The Directory keeps the previous signing key served until its tokens expire — don't force-expire it early. See Rotate keys. |
| Group-key epoch retired — a device is still using a previous epoch past its grace cutoff. | Trouble persists well past a group-key rotation (beyond the 1–60 min grace). | The device must pick up the current epoch — within the configured horizon it backfills; past the horizon, class-2 state uses a current-key authoritative snapshot and transient history remains incomplete. See Rotate keys → If something goes wrong. |
| Snapshot recovery blocked — post-horizon client has no fresh Web head assertion. | UI reports recovery blocked or stale head age; cached snapshots may exist but no current head is verified. | Restore a path to Web and verify its snapshot authority/revocation state. Do not force-apply a storage-only snapshot; a cold client cannot distinguish it from rollback. |
| Wrong / missing service token on a box — the server or Web machine has no valid IdentityToken to authenticate to Directory. | Web cannot refresh its group key/capabilities/revocation state; a Server cannot poll /api/revoked-principals or renew its credential generation. Servers are deliberately denied /api/group-key. |
Reissue and atomically redeploy the machine credentials. For Web, mint a new service identity from Devices; for Server, deploy the complete three-file generation. See Stand up the web client and Add a server. |
The Directory's public signing key is served at /api/.well-known/directory-key — confirming
that endpoint returns the expected key is the first check when token verification fails broadly.
(See Stand up the Directory → How to know it worked.)
A device or operator can't be revoked¶
You're trying to cut off access and the revoke action isn't available, or the revocation doesn't seem to take effect.
| Likely cause | How to check | Fix |
|---|---|---|
| The Revoke button is greyed out — the system won't let you revoke the last active principal or device, to prevent a full lockout. | The button is disabled on the page. | Add another active principal or device first, then revoke this one. See Revoke a principal. |
| The Revoke button is missing entirely — either it's already revoked, or you're viewing as a non-admin. | Check for a red revoked badge; if it's not there, your login may lack the Admin role. | If already revoked, you're done. Otherwise ask an Admin — only the Admin role sees the Revoke control. See Revoke a principal → If something goes wrong. |
| You're looking under the wrong menu — people are under Principals, devices are under Devices (both top-level). | You can't find the record because you're in the wrong list. | Principals and Devices are separate top-level menu items. Pick the right one. See Revoke a principal and Onboard a device. |
| Revocation has not propagated yet — the change is instant in Directory, but cached clients and Relay + Storage learn it on their next signed-snapshot refresh. | The principal may still appear connected briefly, but its values stop being accepted once each verifier receives the new snapshot. | Wait for the bounded revocation refresh. Receivers then reject the principal/device and Relay + Storage rejects or purges its retained values; a token-only native Relay does not identify and force-close the application session. Loss/capture also triggers group-key rotation for newly sealed content. See Revoke a principal → How to know it worked and Security model → Revocation. |
| You need to revoke only one device, not the operator — a device key is compromised but the person stays active. | You want to keep the operator and kill one device. | The RevocationList carries both levels — revoke the device (under Devices) to drop that device's sign-key without touching the operator. See Security model → Revocation. |
Revocation is one-way and permanent — there is no un-revoke. If you revoke the wrong record, the only path back is to onboard them again from scratch (a fresh profile, re-register their security key). See Revoke a principal.
See also¶
- Before you begin — prerequisites, roles, and the end-to-end checklist.
- Stand up a deployment — the deployment order and the four "is it up?" checks.
- Security model — the receive-side gates, revocation, and classification.
- Wire protocol — the cell-first namespace, transport, and group-key rotation.
- Operator training index — every runbook's own "If something goes wrong" section.
Verified against directory@6fcd1201 / server@0a794550 / web@f5683354 / infrastructure@5ac55aeb on 2026-09-02 — browser-hostname and bridge symptoms were rechecked against the owned-name private routes, rack L4 boundary, Caddy :7448 publication, Web endpoint expansion, and the current identity/revocation receive paths.