Operations
The kit ships as a compose stack you can run on one VPS (or a home server) and reach from anywhere
through Tailscale. This page is the buyer-facing version of infra/README.md — the runbook lives
there, next to the files.
First boot
Section titled “First boot”# one-time: Docker, Tailscale, a `deploy` user, ufw, the clone, .env with generated secretscurl -fsSL https://raw.githubusercontent.com/<you>/<repo>/main/infra/scripts/provision.sh \ | sudo REPO_URL=https://github.com/<you>/<repo>.git EXPOSE=tailscale bash
su - deploy # or: sudo -iu deploycd /opt/director# fill DOMAIN (public) or WEB_URL / API_URL (tailnet) — provision.sh already set CADDY_MODE and# generated BETTER_AUTH_SECRET, POSTGRES_PASSWORD and SERVICE_TOKEN; see Configurationdocker compose up -d --buildThen create the first platform administrator, from the checkout on the box (provision.sh installs
Bun for the deploy user and wrote a DATABASE_URL for the compose Postgres, published on loopback
only, into .env):
bun installADMIN_PASSWORD='…' bun run admin:create -- --email you@example.com --name 'You'The password comes from ADMIN_PASSWORD, or from an interactive prompt when that is unset — never
from the command line, where it would stay in shell history and be visible in ps to every other
user of the box. The script runs through the kit’s own auth config (no bunx, no CLI download, no
network), so it cannot drift from the pinned better-auth and it works inside the api container
too. The checkout is simply the easier place: it has the real .env, and Postgres is published on
127.0.0.1:${POSTGRES_PORT} (5433 in the shipped .env) for exactly this. On a box you set up by
hand, prefix the command with
DATABASE_URL=postgres://director:<POSTGRES_PASSWORD>@127.0.0.1:5433/director.
Sign in as that user and open /admin.
Always run docker compose from the repository root. .env sets COMPOSE_FILE=infra/compose.yml,
so no -f is needed and interpolation (DOMAIN, CADDY_MODE, POSTGRES_PASSWORD) comes from the
same file. POSTGRES_PASSWORD, SERVICE_TOKEN and BETTER_AUTH_SECRET are required: compose stops
with set … in .env instead of starting on a placeholder. Migrations run automatically (migrate
one-shot); api and worker start after it exits 0.
Optional services are compose profiles. Put them in COMPOSE_PROFILES in .env, not on the command
line: infra/scripts/deploy.sh runs up --remove-orphans, which removes the containers of every
profile that is not active in that invocation.
Every service runs with no-new-privileges and cap_drop: ALL (Caddy adds back only
NET_BIND_SERVICE) under a 512-process cap, and Postgres gets shm_size: 1gb so pgvector scans can
parallelise.
| Profile | Service | What |
|---|---|---|
backup |
backup |
nightly pg_dump (+ uploads), S3 copy, heartbeat — infra/backup |
worker |
worker |
job worker as its own process (WORKER_ENABLED=false on the API). Its container health check is disabled by design: it shares the API image but serves no HTTP, so the image’s HEALTHCHECK could never pass |
monitoring |
uptime-kuma |
uptime + alerting on 127.0.0.1:3001 |
observability |
lgtm |
Grafana + Tempo + Loki + Prometheus on 127.0.0.1:3030; set OTEL_EXPORTER_OTLP_ENDPOINT=http://lgtm:4318 |
# after enabling the backup profile: prove it works now, do not wait for 02:30 (or set BACKUP_ON_START=true)docker compose run --rm backup bun src/cli.ts backupCaddy modes
Section titled “Caddy modes”subdomains |
single-origin |
|
|---|---|---|
| Use | a public domain | a tailnet / LAN host |
| Layout | DOMAIN, docs., app., api. |
/ app, /api /rpc /uploads /healthz /readyz API, /docs docs |
| TLS | Caddy + Let’s Encrypt | terminated upstream (tailscale serve / funnel) |
| Cookies | COOKIE_DOMAIN=.DOMAIN |
same origin, no CORS |
Both Caddyfiles set HSTS / nosniff / X-Frame-Options / Referrer-Policy / Permissions-Policy,
write JSON access logs, forward one X-Forwarded-For and an X-Request-Id, and hold requests for
up to 15 s while an upstream restarts (active health checks on /readyz and /healthz).
Request bodies are capped at the edge: 32 MB on /uploads/*, 1 MB everywhere else, over
which Caddy answers 413 before anything reaches an upstream. Header, body-read and idle timeouts
are set; there is deliberately no write timeout, because it would cut a jobs.stream SSE
connection mid-flight. The two static sites get their own Content-Security-Policy from Caddy — the
docs’ one allows cdn.jsdelivr.net for the interactive API reference, the marketing site’s allows no
third party at all; both still need 'unsafe-inline' for styles and Astro’s inline scripts, which is
the ceiling until Astro’s own CSP support is adopted. The access log deletes sig, token,
code, state and email from the query string it records, so an emailed reset link never lands in
a log file.
Caddy’s ports are published on CADDY_BIND (default 0.0.0.0). Docker-published ports bypass ufw,
so with EXPOSE=tailscale provision.sh sets CADDY_BIND=127.0.0.1: only tailscale serve/funnel,
which proxy to loopback, reach Caddy.
sudo tailscale serve --bg 80 # tailnet onlysudo tailscale funnel --bg 80 # public, no open portsWhat the API and the app send
Section titled “What the API and the app send”Every API response carries Content-Security-Policy: default-src 'none'; frame-ancestors 'none' and
Cache-Control: private, no-store: nothing the API answers is a document that may load anything, be
framed, or be kept in a shared cache. Two exceptions are deliberate — /healthz and /readyz set no
cache header (machines poll them), and /api/auth/reference is exempt from the CSP because Better
Auth’s Scalar page loads a CDN script; that page and /api/auth/open-api/generate-schema are 404 in
production anyway.
Behind Caddy’s caps the API bounds bodies twice more: /rpc/* and /api/v1/* answer 413
payload_too_large over 1 MiB, and the local driver’s presigned PUT requires a Content-Length
(411 length_required), refuses a declared or actual size over MAX_UPLOAD_BYTES
(413 too_large), and returns no ETag. CORS allows Last-Event-Id so an EventSource polyfill can
resume jobs.stream. Downloads set Content-Disposition per RFC 6266, so a document named in
Chinese, Cyrillic or with an emoji downloads on both storage drivers.
Three limiters guard the procedure surface and all three are in-process: the shared 300 requests
per minute per IP, documents.search’s 30 per minute per user, and the five concurrent
jobs.stream connections one account may hold. They are sized for one API container — running
several replicas behind Caddy needs a shared store before any of the three means what it says.
apps/web adds Cross-Origin-Opener-Policy: same-origin to every page and
Cache-Control: private, no-store to /app and /admin. Its service worker caches the build and
the prerendered public pages only — never a navigation, which would be one user’s HTML with their
session inlined — answers a failed navigation with /offline, and clears its caches on sign-out.
robots.txt disallows /app and /admin.
Backups
Section titled “Backups”# COMPOSE_PROFILES=backup belongs in .env, not on the command line: deploy.sh runs# `up --remove-orphans`, which drops the containers of every profile not active in that invocation.docker compose up -d # with COMPOSE_PROFILES=backup setdocker compose run --rm backup bun src/cli.ts backup # one backup, nowdocker compose run --rm backup bun src/cli.ts list # disk + bucket; exits 1 on a stale setdocker compose run --rm backup bun src/cli.ts restore latest --drill # into a scratch database, then droppedlist exits 1 when the newest restorable set is older than BACKUP_MAX_AGE (36 h), or when
there is none — which is what makes it usable from cron or a health check, and the only thing that
detects a stale backup rather than a failed one. --no-max-age turns the check off while you are
browsing a recovery box. Separately, BACKUP_CATCHUP (on by default) takes one backup at start when
the newest set is older than a schedule interval, because Bun.cron has no catch-up and a reboot
past 02:30 would otherwise skip the night in silence. The drill restores into a scratch database and
asserts the dump’s checksum and the row counts recorded in the manifest — not merely that the
tables came back.
A dump without the .env (or its SOPS copy) is not a restore.
Uploaded documents depend on the storage driver. With the local driver they live on the uploads
volume and the nightly run archives them next to the dump. With the S3 driver (S3_BUCKET set) the
tool backs up no documents at all — they are in your bucket, so turn on object versioning and a
lifecycle rule there and treat that as their backup.
Neither BACKUP_S3_BUCKET nor S3_BUCKET → backups stay on the same disk; the tool warns. Point
BACKUP_HEARTBEAT_URL at Healthchecks.io or an Uptime Kuma push monitor so a missed night pages you.
A real restore needs --yes (without it the CLI prints what it would run and stops); it writes
director_<stamp>_pre_restore.dump into the backups directory first, so a restore of the wrong stamp
is itself recoverable (--no-safety-dump skips it; safety dumps are never pruned).
Before it touches the database the tool refuses three things, and only two of the refusals have an override:
- the dump is re-hashed against the
sha256in the set’s manifest. A mismatch is never overridable — the only safe answer is another set. (--no-manifest-checkwaives a missing manifest, not a failed hash.) manifest.databasemust match the target, so a staging dump cannot land in production (--force-database-mismatch);- no other session may be connected. Stop the writers —
docker compose stop api worker— or pass--terminate-connections.
The restore itself runs --single-transaction --exit-on-error under BACKUP_LOCK_TIMEOUT, so it is
all-or-nothing and fails instead of hanging on a writer’s lock. A very large schema can exhaust
max_locks_per_transaction (out of shared memory): raise that setting, or pass
--no-single-transaction knowing that a failure then leaves the database half-restored. Only one
backup or restore of a database runs at a time (a Postgres advisory lock), so a manual backup cannot
capture a half-restored database and pass verification.
The dump is plaintext — password hashes, sessions, email addresses, Stripe ids, document text —
and every run says so. Encrypt the volume holding BACKUP_DIR, turn on server-side encryption in the
bucket, give the backup a write-only credential (PutObject + HeadObject) and prune by
lifecycle rule rather than by delete permission, and keep object versioning (or object lock) on so a
leaked key cannot erase the history. Each set’s manifest also records rowCounts and
uploadsDegraded, so a run whose uploads archive came out incomplete says so in the set itself and
not only in a log line.
--uploads unpacks the uploads archive as well — it merges into the volume, so a file deleted after that
backup comes back; wipe the volume first if you want the archive exactly. A run -v on
/data/uploads does not override the service’s read-only mount — Compose keeps :ro for that
target path even if the flag says :rw — so mount the volume at another path for the one-shot and
point STORAGE_DIR at it.
Bare metal, from nothing (pg_dump captures neither roles nor the database itself, so Postgres has
to exist before the restore):
# new box → provision.sh → restore the .env you kept with the backups (and the backup files themselves)docker compose up -d postgresdocker compose run --rm backup bun src/cli.ts restore <stamp> --yes# documents too, at a path the service does not mount read-only:docker compose run --rm -v director_uploads:/restore/uploads -e STORAGE_DIR=/restore/uploads \ backup bun src/cli.ts restore <stamp> --yes --uploadsdocker compose up -dVariables and every restore flag: infra/backup/README.md.
Health, logs, traces
Section titled “Health, logs, traces”| Signal | Where |
|---|---|
| Liveness | GET /healthz on the API and the web (process up) — container health checks on api/web/ai/caddy, Caddy’s active checks. The worker container has none: it serves no HTTP, so its health check is disabled on purpose |
| Readiness | GET /readyz on the API (database reachable) — Caddy only routes to a ready API. A 503 answers {"status":"unavailable","checks":{"database":"timeout"}} (or "unreachable") and nothing more; why is in the API’s readiness check failed log line (check, outcome, reason). Concurrent probes share one bounded select 1, so polling a hung database costs one connection, not one per poll |
| Logs | docker compose logs -f api — JSON lines in production (LOG_FORMAT, LOG_LEVEL). Every line written while a request is handled carries its requestId (Caddy’s X-Request-Id), procedure calls also carry procedure (e.g. documents.list), and traceId when tracing is on. A 500 repeats the id to the caller: {"error":"internal_error","requestId":…} — grep that |
| Traces | OTEL_EXPORTER_OTLP_ENDPOINT — the observability profile, or Grafana Cloud / Axiom / Honeycomb (OTEL_EXPORTER_OTLP_HEADERS). A job and the ai.* calls it makes share one trace; the request that queued the job is a separate trace on purpose (a job runs later, usually in another process, so stitching them would make waterfalls misleading) — find it by the jobId in the logs, not by traceparent. Health probes (/healthz, /readyz, the AI /health) are not traced and query strings are stripped from span URLs |
| Uptime | COMPOSE_PROFILES=monitoring in .env (Uptime Kuma, a stock container — add monitors for http://api:3000/readyz, http://web:3000/healthz, http://ai:8000/health and the backup heartbeat); or Better Stack from the outside |
| System page | /admin/system — version, migrations, AI, queue, storage, providers. The AI service answers /health with 200 and ok: false when its configuration cannot serve (no token check, an OCR backend that cannot run there); the faults are listed in problems and shown in red here. Not a 503, because a restart cannot fix an environment — the container health check asserts liveness only |
Deploy from CI
Section titled “Deploy from CI”.github/workflows/deploy.yml builds {api,web,ai,caddy,backup} on every push to main (and v*
tags) and pushes them to ghcr.io/<owner>/<repo> tagged sha-<commit> plus main and latest — a
v* tag gets the version instead of latest. The caddy image bakes the canonical URLs of the two
static sites at build time, so it requires four repository variables — SITE_URL, DOCS_URL,
WEB_URL, API_URL — and fails with a message naming them rather than shipping example.com.
Set the repository variable DEPLOY_HOST and the secret DEPLOY_SSH_KEY and the workflow then logs
the server into GHCR and runs infra/scripts/deploy.sh on it (docker compose pull, up -d --wait,
/readyz), passing IMAGE_REGISTRY so a fork’s .env never has to name the registry by hand. A
server that is only on your tailnet: DEPLOY_VIA_TAILSCALE=true plus a Tailscale OAuth client.
Pulling the images by hand needs a login of your own: IMAGE_REGISTRY must point at your fork’s
namespace, and private packages need docker login ghcr.io with a token that has read:packages
(the workflow’s login uses a token that expires with the job).
Rollback: Actions → Deploy → Run workflow with tag: sha-<previous> (skips the build), or
IMAGE_TAG=sha-abc1234 GIT_REF=abc1234 bash infra/scripts/deploy.sh on the box; deploy.sh pins
the deployed tag and registry in .env, so the rollback survives a later docker compose up -d.
A sha-… tag or a released version (1.2.0 → its git tag’s commit) is deployed with the matching
commit checked out; anything the workflow cannot resolve to a commit fails instead of deploying old
images against current main. Migrations are forward-only — roll back code, not schema.
Secrets
Section titled “Secrets”provision.sh writes a chmod 600 .env with generated BETTER_AUTH_SECRET,
POSTGRES_PASSWORD and SERVICE_TOKEN. In production the API refuses the example placeholders
outright — any change-me… value, and a DATABASE_URL whose password is director, postgres or
change-me… — so generate every secret (openssl rand -hex 32).
To version the file: encrypt with SOPS + age into
infra/env/production.enc.env (see /.sops.yaml and infra/env/README.md); CI decrypts with
the SOPS_AGE_KEY secret. 1Password (op inject) and Doppler work the same way — the
containers only ever see environment variables.
Audits and load
Section titled “Audits and load”CI runs bun audit --prod --audit-level=high and pip-audit on the Python lockfile,
bun run licenses --strict over both dependency trees, cargo fmt --check / clippy / check on
the Tauri crate, the Playwright end-to-end suite (home, axe/WCAG 2.2 AA, locales) against a
production preview of apps/web, docker compose config (every profile) + caddy validate on both
Caddyfiles, the backup tool’s integration suite against a real Postgres 17, and — on pull requests —
a build of all five images. Every action is pinned to a commit SHA, and .github/dependabot.yml
opens one grouped weekly pull request that moves the pins.
docker run --rm -i -e API_URL=http://<api-origin> grafana/k6 run - < infra/loadtest/k6/smoke.js# optional signed-in scenario: -e K6_EMAIL=… -e K6_PASSWORD=…Aim it at the API directly, not through Caddy. The script gives every virtual user its own
X-Forwarded-For address so the per-IP limiter sees many clients — which only works where the API
trusts that header (TRUST_PROXY=true, what infra/compose.yml sets on the api service) and where
Caddy is not in front rewriting it to the real client address; the header must hold exactly one
address (a chain falls back to the socket), and against an API started without TRUST_PROXY=true
every virtual user lands in the same bucket. Tune DATABASE_POOL_MAX so the whole stack stays under
Postgres max_connections (100 by default): DATABASE_POOL_MAX per API and per worker process,
2 for migrate while it runs (an advisory lock plus the migrator), 1 for backup, plus your own
psql.
The full runbook — firewall, troubleshooting, image tags, scaling workers — is infra/README.md.