Skip to content

Stand Up the Directory

What this is: deploying the Directory — the identity authority that issues every IdentityToken and signing permit in the system. Nothing else can work until this exists: servers, users, and web clients all verify their tokens against the Directory's public signing key. When you'd do it: first bring-up of the rack, or rebuilding the Directory from scratch. How long it takes: minutes for the converge; the first-admin /setup flow takes about five more once the site is reachable.

This is the first thing you stand up — before servers, before onboarding any users. See the system overview and PKI model for how the Directory fits into the trust chain.

On the PETRA rack it runs as a container on core-01, alongside its Postgres and the Caddy that fronts it. playbooks/core.yml stands it up.

Who can do this: a deployment engineer with rack access — an SSH key in ssh_authorized_keys and the vault password on their laptop (see the infrastructure repo's Onboarding section). The first-admin bootstrap is part of the same job: because there are no Admins yet, this creates the very first one via the one-time /setup flow, not the normal admin dashboard, which doesn't exist yet.

The shape of it

flowchart LR
    A["ansible-playbook<br/>playbooks/core.yml"] --> B["postgres + directory<br/>containers on core-01"]
    B --> C["Open /setup<br/>Enter the setup key"]
    C --> D["First person registers<br/>their FIDO passkey"]
    D --> E["They become<br/>the first Admin"]
    E --> F["Directory is live —<br/>sign in and start<br/>adding servers + users"]

Before you start

  • Your controller is set up: bin/bootstrap-controller.sh has put the vault password on your laptop, and your SSH key is on the boxes.
  • playbooks/foundation.yml has run at least once on core-01 — it creates the Compose directories and brings up the registry the Directory image is pulled from.
  • A FIDO2 hardware security key for the first Admin. It is the only authentication method the admin UI accepts.
  • A way to reach https://auth.bedrockdefence.com — directly over the rack LAN with the owned name resolved to core-01, or through the managed Cloudflare private-hostname path.

Step 1 — Converge

ansible-playbook playbooks/core.yml

playbooks/core.yml orders Postgres → Directory → web, so the database is up and healthy before the Directory starts; the Compose fragment makes that ordering explicit with a depends_on on Postgres' healthcheck.

Nothing needs supplying by hand. Two app secrets are generated on the box on the first converge and then never regenerated — roles/service/directory writes them with force: false:

Secret File on core-01
APP_KEY /etc/bedrock/directory/secrets.env
DIRECTORY_WEBAUTH_CLIENT_SECRET same file

The database password is never written into that file — it is read live from the database_url fact on every run.

Where the signing key actually lives

The Ed25519 key that signs every IdentityToken is held in the database, in directory_signing_keys, not in a file. The Directory generates it itself on first boot and keeps rotation history, so a retired key stays served as previous long enough for tokens signed by it to expire naturally rather than breaking in flight.

DIRECTORY_SIGNING_KEY_PATH still appears in the container's environment, but it is a one-time seed path, consulted only on a cold boot with an empty table so that an older file-based deployment's key can be carried over. On a fresh rack that file does not exist, so a fresh keypair is generated and the database is authoritative from then on. The file is never written back.

So what you back up is Postgres. /var/lib/bedrock/postgres-data on core-01 holds the identity root of the entire deployment — lose it and every IdentityToken, ServerToken and enrolled device has to be re-issued. Backing up a key file gets you nothing, because there isn't one.

Rotation is a deliberate act through the admin UI, not something you do on the filesystem — see Rotate keys.

Docker lab path (non-production)

For local development, the directory repo has a docker-compose.yml that starts a Directory container alongside a plain Postgres 17 container:

cp .env.example .env
# Fill in APP_KEY — run: node ace generate:key
docker compose up --build

The entrypoint runs migrations automatically and the Directory starts on port 3333. Leave DIRECTORY_SIGNING_KEY_PATH alone unless you are deliberately importing a key from an older deployment; with no file there, the Directory mints its own on first boot.

Step 2 — Run the first-time /setup

This step only works once. The /setup route checks whether any principals exist in the database. When the database is empty, setup is open. The moment the first Admin is created, the setup key is consumed and /setup returns 404 for every future request.

How the setup key works: on boot, when no principals exist, the Directory generates a random 32-byte setup key (displayed as 64 uppercase hex characters) and prints it to the application log. It has no time-based expiry: it remains valid in memory until the first Admin is created or the Directory process restarts. After a restart, use the newest key in the logs.

  1. Find the setup key in the container log, on core-01:
ssh ubuntu@core-01.bedrock.lan
docker logs directory | grep -i setup

On the Docker lab path: docker compose logs directory | grep -i setup.

  1. Open https://auth.bedrockdefence.com/setup in a browser. Browsers warn about the certificate until the rack CA is trusted on that machine — see environments/rack/README.md in the infrastructure repo. That warning is expected, not a fault.

  2. Enter the setup key from the logs and your display name. This calls POST /setup/fido/begin, which validates the key and starts a WebAuthn registration ceremony.

  3. Register your FIDO2 passkey when the browser prompts you. On success the browser calls POST /setup/fido/complete, the Directory creates the first principal with roles admin and operator, and the setup key is consumed.

  4. You're redirected to /login. From this point /setup is permanently closed — it returns 404 because principals now exist.

The three setup routes:

Route What it does
GET /setup Shows the setup page (404 if principals exist)
POST /setup/fido/begin Validates the setup key, starts FIDO registration
POST /setup/fido/complete Verifies the passkey, creates the first Admin

Step 3 — Sign in as the first Admin

Navigate to /login and authenticate with the passkey you just registered. You land on the admin dashboard at /admin/dashboard. From here you can:

  • add more principals (people) and devices
  • register servers — each must be enrolled with the Directory before it can carry traffic, see Add a server
  • mint the machine identities the rest of the rack waits on: the router's atomic service-token, ServerToken, and replay-signing-key generation, and the web tier's service token and signing key
  • manage signing keys via /admin/keys

A first rack bring-up is expected to come up partially: until those machine identities exist, the routers and the web service identity skip themselves. That is by design — mint them here, paste them into the vault, and re-run playbooks/site.yml.

How to know it worked

  • The admin dashboard loads at /admin/dashboard and you can sign in with your passkey.
  • GET /api/.well-known/directory-key returns the Directory's public signing key (lowercase hex) — the endpoint everything else uses to verify IdentityTokens.
  • On core-01, docker logs directory shows migrations completed and no startup errors.

If something goes wrong

  • /setup returns 404 immediately. A principal already exists — someone ran setup before you, or a previous attempt partly succeeded. Check the principals table (see Connect to the database), then either sign in as that Admin or follow the reset procedure in the directory repo.

  • Setup key rejected ("Invalid or expired setup key"). The key is single-use and is replaced whenever the Directory process restarts. Check that you copied all 64 hex characters from the latest container start. If in doubt, docker restart directory on core-01, then use the newly logged key.

  • Database not reachable. The Directory waits on Postgres' healthcheck, so this usually means Postgres itself did not come up. Check docker compose -f /etc/bedrock/compose/docker-compose.yml ps on core-01.

  • Certificate warning in the browser. The rack CA isn't trusted on that machine. This is expected — the rack mints its own certificates and has no public CA. See environments/rack/README.md.

  • auth.bedrockdefence.com doesn't resolve. This is the owned, browser-facing name. Rack services resolve it through the managed /etc/hosts block; an attached operator laptop needs the reviewed LAN name-to-address path, while an off-LAN managed device needs active WARP and the Cloudflare private-hostname route. Do not substitute a .lan service name — that suffix is retained only for rack-internal management names.

  • "legacy signing key at … is N bytes, expected 32". Something that is not a 32-byte Ed25519 key is sitting at DIRECTORY_SIGNING_KEY_PATH on a cold boot with an empty table. The Directory fails loudly at seed time rather than inserting a bad row. Remove the stray file and let it mint its own.

See also


Verified against directory@6fcd1201 / infrastructure@5ac55aeb on 2026-09-02 — the browser origin and WebAuthn RP were checked against inventories/rack/group_vars/all/vars.yml, roles/service/directory/templates/directory.yml.j2, playbooks/core.yml, and Directory's setup routes/runtime env. Device/browser Directory traffic uses auth.bedrockdefence.com; rack management SSH remains on core-01.bedrock.lan.