Skip to content

Troubleshooting

A cross-cutting, symptom-first guide to the things that actually go wrong on a running deployment. The individual runbooks each carry an "If something goes wrong" section for their own step; this page collects the failures that span steps — where the cause is in one component and the symptom shows up in another — so you can start from what you're seeing and work back to the fix.

For each symptom: the likely cause, how to check, and the fix (with a link to the runbook or reference that owns the detail).


A client can't connect

The user opens the app or the web client and never reaches the map — a TLS warning, a login error, or a silent hang.

Likely cause How to check Fix
Certificate not trusted — the client doesn't have the Root CA, or the server presents only its leaf, not the chain. Connecting shows an "untrusted" / certificate warning before any login screen. Distribute the Root CA to clients and make sure the server presents leaf + intermediate (not just the leaf). See Set up the certificate chain → If something goes wrong.
Hostname mismatch — the server's certificate SAN doesn't list the name the client dialed. A hostname-mismatch TLS error naming the host. Reissue the server leaf with the right hostname in the SAN. See Set up the certificate chain.
A browser hostname doesn't resolve — the site or the Directory is unreachable. dig auth.bedrockdefence.com returns nothing on the client path. Device-facing URLs use owned names. Rack hosts map them through managed /etc/hosts; off-LAN managed clients need active WARP and the matching Cloudflare private-hostname route. Do not substitute .lan, which is internal-only. See Stand up the Directory.
Login fails (web) — the Directory is unreachable from the web container, or the web service token is missing. Login redirect loops, or fails closed. Two distinct causes: the rack CA never reached the container, so every HTTPS call to the Directory fails (Node only warns about a missing NODE_EXTRA_CA_CERTS, so the app starts and then fails); or web_service_token_b64 is unset, so the revocation cache never primes. See Stand up the web client → If something goes wrong.
User isn't onboarded — "user not found" / access denied. Login reaches the Directory but is rejected for that person. Onboard them first. See Onboard an operator.
The user was revoked — their tokens are now refused everywhere. Their Directory page shows a red revoked badge. Revocation is permanent. If it was a mistake, onboard them again from scratch. See Revoke a principal and A device or operator can't be revoked below.

The authoritative gate on the client→router hop is server-cert TLS plus app-layer envelope verify — the client presents no client cert. If TLS is clean but messages still don't flow, the problem is upstream of connection: see the next two sections. (See Security model → Transport TLS.)


Messages aren't arriving

The client connects and the map loads, but positions, chat, or drawings from other users never show up — or only some do.

Likely cause How to check Fix
No server covers the field user's cell — routing is cell-first; coverage bounds ordinary field position/heartbeat publication. Affected field users are operating in an area whose cell isn't in any server's coverage list. Their positions do not appear, while Web's separate HQ all-cell authority may still work. Add the missing cell(s) to the server's coverage and reissue its ServerToken, then restart the box. See Add a server → How geofencing works and the wire protocol key namespace.
The web client has no bridge configured — the browser opens no session. web_router_endpoints is blank, or the zenoh-bridge container is not running. Point the browser at the remote-api bridge, not the router's :7447. On the rack web_router_endpoints is app.bedrockdefence.com and the app expands a bare domain to wss://<domain>:7448, which Caddy proxies to the bridge; re-run playbooks/core.yml. See Stand up the web client → Server / relay.
The web client points at the router, not the bridge — browser zenoh-ts speaks only the zenoh-plugin-remote-api WebSocket, not the router's native transport. Browser console loops 1006 / "disconnected from remote-api-plugin"; web_router_endpoints points at :7447 or the bridge isn't running. Use the bridge's wss://…:7448, and confirm the bridge sidecar is up and web_bridge_connect_endpoints reaches a live router. See Stand up the web client → Server / relay.
A classification gate is dropping the message — retained ingest/replay must lie within the subject, capability, and ServerToken.clearance ceilings; every LIVE receiver separately checks the subject and capability ceilings before decode. Server audit shows a retained classification denial, or a receiving endpoint rejects the LIVE value after verification. For retained traffic, confirm the Server's Classification is high enough; edit the device, Reissue Credentials, and redeploy if needed. For LIVE traffic, inspect the sender capability and receiver clearance rather than treating the native Relay as the gate. See Add a server → Step 2 and Security model → Classification.
A device just had its group key rotated — devices pick up the new epoch on their next background poll; a brief window can look like dropped traffic. Trouble appears right after a key rotation and clears within a few minutes. Wait for the next poll; the previous epoch stays valid through the bounded grace window. If it persists, see Tokens are rejected. See Rotate keys → If something goes wrong and wire protocol → Group-key rotation and backfill.
It's a duplicate / replay drop, not a loss — the receive pipeline rejects stale frames (outside the ±60 s window) and repeated (principal, nonce) pairs. Only old or re-sent frames are missing; live traffic is fine. Expected behaviour, not a fault. See Security model → Receive-side gates.

A server won't federate

Two servers are up but won't link and share traffic — users on one relay can't see users on the other.

Likely cause How to check Fix
Peer certificate not trusted — each router must trust the CA signing the other's server leaf. One or both dialers reject the peer's certificate at link time. Configure the shared Root / intermediate as transport.tls.peer_ca_path. See Set up the certificate chain → If something goes wrong and wire protocol → Federation.
A server's ServerToken is missing or expired — the box can't authenticate itself into the mesh. The server starts but rejects traffic, or won't carry some messages. Reissue the ServerToken from the device's page (Reissue Credentials), redeploy all three files, and restart. See Add a server → If something goes wrong.
The box won't start at all — usually TLS certs or the encrypted data-directory mount. The server process exits on boot; health check on 9090 never answers. Check the TLS cert/key paths and the data-directory mount first. See Add a server → If something goes wrong.

Forwarded IdentityTokens are byte-identical across hops, so the Directory signature stays verifiable end-to-end — routers never re-sign identity. If federation is up but identity claims are being rejected across the link, that's a token problem, not a federation one — see the next section.


Tokens are rejected

A connection or message is refused at the identity gate — the token doesn't verify, has expired, or signs under the wrong key.

Likely cause How to check Fix
Token expired — expires_at_ms is in the past (gate 3). The principal worked until a deadline, then stopped. Service devices that don't self-refresh need reissuing. For a service device (server/gateway/web token) Reissue Credentials from its page; operators get fresh passes on next login. See Add a server → Step 3 and Security model → Receive-side gates.
Clock skew — issued_at_ms is outside the ±60 s replay window (gate 2). Fresh frames from one host are rejected as stale/future while others work. Fix the host's clock (NTP). The window is fixed at 60 s for fresh frames. See Security model → Receive-side gates.
Signing key rotated, key set stale — the Directory signed under a fresh key the verifier hasn't cached yet. Rejections start right after a signing-key rotation; the old key is still served until its tokens expire. This resolves as verifiers refresh the Directory key set. The Directory keeps the previous signing key served until its tokens expire — don't force-expire it early. See Rotate keys.
Group-key epoch retired — a device is still using a previous epoch past its grace cutoff. Trouble persists well past a group-key rotation (beyond the 1–60 min grace). The device must pick up the current epoch — within the configured horizon it backfills; past the horizon, class-2 state uses a current-key authoritative snapshot and transient history remains incomplete. See Rotate keys → If something goes wrong.
Snapshot recovery blocked — post-horizon client has no fresh Web head assertion. UI reports recovery blocked or stale head age; cached snapshots may exist but no current head is verified. Restore a path to Web and verify its snapshot authority/revocation state. Do not force-apply a storage-only snapshot; a cold client cannot distinguish it from rollback.
Wrong / missing service token on a box — the server or Web machine has no valid IdentityToken to authenticate to Directory. Web cannot refresh its group key/capabilities/revocation state; a Server cannot poll /api/revoked-principals or renew its credential generation. Servers are deliberately denied /api/group-key. Reissue and atomically redeploy the machine credentials. For Web, mint a new service identity from Devices; for Server, deploy the complete three-file generation. See Stand up the web client and Add a server.

The Directory's public signing key is served at /api/.well-known/directory-key — confirming that endpoint returns the expected key is the first check when token verification fails broadly. (See Stand up the Directory → How to know it worked.)


A device or operator can't be revoked

You're trying to cut off access and the revoke action isn't available, or the revocation doesn't seem to take effect.

Likely cause How to check Fix
The Revoke button is greyed out — the system won't let you revoke the last active principal or device, to prevent a full lockout. The button is disabled on the page. Add another active principal or device first, then revoke this one. See Revoke a principal.
The Revoke button is missing entirely — either it's already revoked, or you're viewing as a non-admin. Check for a red revoked badge; if it's not there, your login may lack the Admin role. If already revoked, you're done. Otherwise ask an Admin — only the Admin role sees the Revoke control. See Revoke a principal → If something goes wrong.
You're looking under the wrong menu — people are under Principals, devices are under Devices (both top-level). You can't find the record because you're in the wrong list. Principals and Devices are separate top-level menu items. Pick the right one. See Revoke a principal and Onboard a device.
Revocation has not propagated yet — the change is instant in Directory, but cached clients and Relay + Storage learn it on their next signed-snapshot refresh. The principal may still appear connected briefly, but its values stop being accepted once each verifier receives the new snapshot. Wait for the bounded revocation refresh. Receivers then reject the principal/device and Relay + Storage rejects or purges its retained values; a token-only native Relay does not identify and force-close the application session. Loss/capture also triggers group-key rotation for newly sealed content. See Revoke a principal → How to know it worked and Security model → Revocation.
You need to revoke only one device, not the operator — a device key is compromised but the person stays active. You want to keep the operator and kill one device. The RevocationList carries both levels — revoke the device (under Devices) to drop that device's sign-key without touching the operator. See Security model → Revocation.

Revocation is one-way and permanent — there is no un-revoke. If you revoke the wrong record, the only path back is to onboard them again from scratch (a fresh profile, re-register their security key). See Revoke a principal.


See also


Verified against directory@6fcd1201 / server@0a794550 / web@f5683354 / infrastructure@5ac55aeb on 2026-09-02 — browser-hostname and bridge symptoms were rechecked against the owned-name private routes, rack L4 boundary, Caddy :7448 publication, Web endpoint expansion, and the current identity/revocation receive paths.