Deploy gotchas — READ THIS BEFORE (and while) DEPLOYING¶
Audience: the agent/human running or debugging a prod deploy. These are the traps that have each cost hours (some cost days) and were NOT obvious from the code. If a deploy step is red, find it here before you form a theory. The governing rule of this whole page:
THE ONE RULE THAT KEEPS GETTING RE-LEARNED THE HARD WAY: prod secrets live in a RAM-only directory that does NOT persist. Anything that creates a new container on the box (a pipeline run, the D5 probe, a reprocess/reenrich/gi-repair job) MUST re-stage the secrets immediately before it, or the container starts with no credentials and the gateway returns 401. A gateway 401 means the key is MISSING from the container, not that the key is wrong. Re-stage — never re-mint. Deleting/re-minting a live key on a 401 has burned entire evenings (2026-08-18, 2026-08-21) and never once was the fix.
Related: PROD_RUNBOOK.md, docs/adr/ADR-115 (tmpfs secrets), docs/adr/ADR-142 (LiteLLM gateway),
the canonical re-stage action .github/actions/stage-prod-secrets/action.yml (read its header — it
is the best explanation of this whole failure), the gate scripts/tools/check_prod_secret_staging.py.
1. The RAM secrets directory does NOT persist — re-stage before EVERY container creation¶
This is the trap behind almost every "it was working, now it 401s again."
How secrets work here (ADR-115, PODCAST_SECRETS_VIA_FILES=1): LLM keys + Sentry DSNs are
delivered to /dev/shm/podcast-secrets/* on the box — RAM only, never written to disk — and
compose/docker-compose.secrets.yml mounts them at /run/secrets/*. That "never on disk" property
is deliberate and good. Its cost is the thing that bites:
The directory is not durable. It is reaped when the SSH session that staged it ends (the box's
systemd-logind reaps a uid≥1000 user's /dev/shm on logout). So:
Fixed at the cause 2026-09-30 — read this before trusting the rest of §1. The reap is logind's default
RemoveIPC=yes, and it took all three secret dirs, not onlypodcast-secrets: at 06:14:54 that day the lastdeploysession closed andoperator-secrets/player-secrets, staged four minutes earlier, went with it. Running containers keep their copies, so nothing looked wrong — but any restart (crash, OOM,docker restart) then fails exactly like a reboot does (§1b)./etc/systemd/logind.conf.d/10-keep-deploy-shm.confnow setsRemoveIPC=no(live on prod and ininfra/cloud-init/prod.user-data), so the dirs last until the next reboot. The re-stage rule below still holds for boxes built before that and costs nothing; after a reboot everything in §1b still applies.
- Docker copies each secret into a container at CREATE time. Containers made during a deploy
keep their keys — which is why
compose-api-1stays healthy and everything looks fine. - Any container created later — a fresh
docker compose run pipeline-llm, the D5 probe, a reprocess job — finds/dev/shm/podcast-secretsgone, mounts nothing, and starts with an empty/run/secrets/litellm_api_key. EmptyBearer→ 401 (or "Deepgram key required", etc.). "prod is deployed" and "prod has credentials" are two different facts.
THE RULE: immediately before any step that creates a container, re-stage with
.github/actions/stage-prod-secrets (or the inline podcast-secrets.staged equivalent). Every real
pipeline workflow already does this (backfill-audio-prod, reprocess-prod, reenrich-prod,
gi-repair-prod, inspect-prod-corpus). deploy-prod.yml's D5 gateway probe re-stages right before
the probe for exactly this reason — do not remove that step.
The gate — and its blind spot. scripts/tools/check_prod_secret_staging.py (in ci-fast) fails
any prod workflow that creates a container but never stages/mounts secrets. It is file-level: it
only checks the tokens exist somewhere in the file. So a workflow that stages once for step A can
still be broken at step B if B creates a second container after the reap — which is exactly how D5
stayed broken while the file "passed." The gate is necessary, not sufficient: you must ensure a
re-stage precedes each container-creation that runs after a session boundary. (Improving the gate
to be proximity-aware is tracked separately.)
1b. A REBOOT wipes all three secret dirs — run restage-prod-secrets FIRST¶
§1 is about the logout reap, which takes /dev/shm/podcast-secrets mid-session. A reboot
is worse: it clears all of /dev/shm, including the two public-surface dirs that §1 does not
mention — operator-secrets and player-secrets. Those are mounted at container create time,
so the containers do not degrade, they refuse to start at all:
player-api-1 Exited(127) open /dev/shm/player-secrets/app_oauth_google_client_secret
operator-api-1 Exited(127) open /dev/shm/operator-secrets/app_oauth_google_client_secret
Exit 127 with NO application logs — runc fails the mount before any process starts. Do not
go looking for an app bug; there is nothing in docker logs.
The damage presents one layer downstream, which is what makes it confusing. On 2026-09-15 the
visible symptom was two nginx containers crashlooping 527 times each
(player-learning-app-1, operator-viewer-1) with host not found in upstream "api" — because
nginx resolves proxy_pass upstreams at startup and each project's own api was the thing that was
dead. Public visitors saw the coming-soon page the whole time, so nothing alerted for 9.5 hours.
Detection (added 2026-09-16). That "nothing alerted" is no longer true, but understand why it was, because the obvious alerts still cannot see this failure:
candidate fires? why Scrape target down (up==0)no player-api-1/operator-api-1are not scraped at all —base.alloy'sdiscovery.relabel "api"keeps only/compose-api-1*logs darkdead-man rulesno the box is alive and still shipping journal + caddy logs, so log volume looks normal a public HTTP check no the coming-soon page returns 200 Container crashlooping on prod (restart storm)YES cadvisor scrapes these containers by name; container_start_time_secondschanges on every restartThe crashloop rule (
changes(...[15m]) > 3, severity critical) fires within ~15 minutes on both the Exited(127) api containers and the nginx storm downstream. It is deploy-safe: a normal deploy produces exactly one restart per container.Still open: widening the api scrape to cover
player-api-1/operator-api-1for trueup==0coverage — needs their/metricsendpoints verified first, or the scrape fails permanently and the target-down rule false-fires forever.
THE RULE after any reboot: run restage-prod-secrets.yml (surfaces: all,
recreate: true). It stages all three dirs — control plane via the canonical action, the two
surfaces via scripts/ops/restage_prod_secrets.sh — and recreates only containers that are
actually down, each at the image tag that container was already on. It never resolves "newest
from main", because shipping untested code during an incident is its own outage.
Before that workflow existed, recovery meant two full public-surface deploys, each with its own
typed confirm and prod gate, and remembering to pin override_image_sha.
The control plane after a reboot is UP, healthy, and has no keys. The earlier claim here — that
compose-api-1 survives a reboot on its create-time copy — was false (2026-09-30).
podcast-scraper.service runs docker compose up at boot without
docker-compose.secrets.yml, so it RECREATES compose-api-1 with no /run/secrets at all. It
passes its health check. scripts/ops/restage_prod_recreate.sh therefore treats an api container
with no /run/secrets/* mounts as down. Do not check keys with docker exec … env: the secrets shim
exports them inside the server process only, so exec shows them empty even when they are there —
list /run/secrets instead.
Three bugs the first real run of this workflow hit (2026-09-30), all fixed: printf %q passed the
surface list with escaped spaces, so podcast and operator were skipped as unknown while the run
went green; the tag was read off the project's first running container, which moved player-api to
an untested image because an obs-only deploy had left player-obs newer; and the control plane was
never recreated because it looked healthy. tests/unit/scripts/ops/test_restage_prod_recreate.py
covers each.
2. A gateway 401 = missing secret, not a bad key — re-stage, NEVER re-mint¶
The single most expensive mistake in this repo's history: a LiteLLM 401 read as "the key is
wrong," a live production key deleted and re-minted, multiple deploys burned. The key was always
correct (2026-08-18; repeated in shape 2026-08-21, where a bogus "reload-race retry" was added on
the same false premise and reverted).
When D5 / the gateway returns 401, in this order — all read-only:
- Is the key even in the container? From a container that has it:
docker exec compose-api-1 sh -c 'wc -c </run/secrets/litellm_api_key'— 0 bytes = missing = you skipped the re-stage (see §1). Fix the staging, not the key. - Does the running stack authenticate right now?
docker exec compose-api-1 sh -c 'K=$(cat /run/secrets/litellm_api_key); curl -s -o /dev/null -w "%{http_code}" -H "Authorization: Bearer $K" "$LITELLM_API_BASE/models"'→ 200 means the key is fine and only the new/probe container lacked it. - Does the mounted key match the expected one?
… | sha256sum | cut -c1-12should be3b88f1c6ee41(proj-podcast-prod). Matches + still 401 only from a fresh container ⇒ §1. - Only if the key is genuinely mounted, correct-hash, and still rejected by the gateway do you have a real key problem — and even then, ask the operator before touching a live key.
Do NOT reach for a retry loop, a re-mint, or "the gateway must be reloading." Those are the wrong tails. The right question is always "does the container that failed actually have the secret?"
Tool: scripts/tools/rehearse_gateway_key_gate.sh runs the deploy gate's logic against a real
gateway. Use it before changing the D5 gate.
3. Secret plumbing details worth knowing¶
- GH Secrets are the source of truth; the on-box RAM files are re-staged from them every time. A stale/empty GH secret would stage a bad value — but the usual cause of a 401 is §1 (not staged at all), not a wrong GH secret. Confirm "mounted + correct hash" before suspecting GH drift.
- A probe that bypasses the entrypoint (
--entrypoint python, etc.) skipsdocker/secrets-shim.sh, so$LITELLM_API_KEYis never exported — read the secret FILE directly (/run/secrets/...), which is what D5 does. - The overlay must be joined or nothing mounts even when the dir is present:
compose/docker-compose.secrets.yml, joined by the presence of/dev/shm/podcast-secrets(deploy.sh:38). Re-staging (§1) is what makes that presence check true at probe time.
4. The prod LiteLLM gateway is the VPS :4001, not the homelab¶
prod (the pipeline running on the VPS) authenticates against the VPS LiteLLM gateway
(vars.PROD_LITELLM_API_BASE, reaching the container as $LITELLM_API_BASE). The homelab gateway is reached only from the operator's
laptop. Config: vars.PROD_LITELLM_API_BASE + the profile pin (D4 / #1676). If a check points the
pipeline at the homelab gateway, that's the bug (#1676) — not the key.
5. Deploy the published 7-char image sha, not a workflow-only sha¶
Images are tagged :sha-<7> (7-char short of the commit that the Stack test publish job built).
deploy-all / deploy-* accept image_sha with or without the sha- prefix but pin sha-<7>. If
you hotfix a workflow (no image rebuild), there is no image at that new sha — deploy the last
published sha (from the Stack test run summary), not the workflow-hotfix sha.
6. Reusable-workflow gotchas (deploy-all calls deploy-player/operator/prod)¶
deploy-all runs the three surface deploys as reusable-workflow jobs (secrets: inherit), so:
- The caller's
permissions:must cover what the called workflows request (packages: readfor GHCR pulls,actions: readfor deploy-prod) — otherwise the run startup_failures at load with jobs:0 ("workflow file issue"). deploy-all grantscontents/packages/actions: read. - A called workflow inherits the TOP-LEVEL dispatch's inputs. When deploy-all (dispatched with
confirm=DEPLOY_ALL) calls a surface, that surface seesinputs.confirm == "DEPLOY_ALL"(NOT empty) andgithub.event_name == "workflow_dispatch"(the caller's). So a "typed-confirm" gate cannot useevent_nameto detect a call — the surfaces acceptconfirm ∈ {SURFACE_DEPLOY, DEPLOY_ALL}. - deploy-all runs the surfaces in PARALLEL against the SAME
/srv/podcast-scraper/.git. They allgit reset --hardand collided on git's.git/shallow.lock(2026-08-21, player failed with "Unable to create shallow.lock: File exists"). The refresh is now serialized with a sharedflock(/tmp/podcast-scraper-git-refresh.lock) in all three (deploy-player.yml,deploy-operator.yml,deploy.sh). Keep the lock if you touch the refresh. - The per-surface "Emit deploy event to VictoriaLogs" step is
if: always()and must staycontinue-on-error: true— an unreachable telemetry sink (e.g. when an earlier step failed before the tailnet join) must never turn the deploy red.
7. Before you blame code, check the machine — and never say "pre-existing"¶
- A saturated box returns transient 502/504 from a healthy api. Check
uptime(load vs 8 vCPU) and container state (docker inspect <c> --format '{{.RestartCount}} {{.State.OOMKilled}}') before theorizing a code bug. - The branch was green before this deploy; red after = this deploy caused it. Do not label a
regression "pre-existing" — diff against the pre-deploy state (
docker psshas, the prior good sha) and own it. (This page exists because that framing wasted trust and time.)
8. The deploy checklist (the happy path)¶
- main green; the Stack test publish job pushed all images at ONE
sha-<7>— record it from the run summary (do NOT rely on "newest from main"). - Dispatch
deploy-all:confirm=DEPLOY_ALL,image_sha=sha-<7>. Oneprod-environment approval releases all three surfaces. - Watch it. If a surface fails, open the step log, then SSH and verify the actual state —
401→ the container is missing the secret; re-stage (§1/§2), do NOT re-mint. - Post-deploy: verify per
docs/wip/OBS-MCP-DEPLOY-DAY-RUNBOOK.md(sha alignment, container health, the exposed surfaces, gateway auth 200).