26 — When it breaks
Read this first: this chapter is a set of real incidents from this system, each written as a symptom first. Read the symptom, work out what you would check, then read on. Every one of these actually happened, and every one left a guard behind in the code that you have already met.
Time: about 35 minutes. Earned by the whole of Track A.
How to use this
For each case: cover the diagnosis, read the symptom, write down your first three checks. Then compare. The point is not to memorise five incidents — it is to notice that your triage instinct is now producing the right questions.
The general order, from lesson 12: what is running, what did it say, what holds the port, is the disk full.
Case 1 — the site is down and the edge is crash-looping
Symptom. A routine deploy. Now nothing responds on any hostname — not the app, not staging, not the docs. The application containers are healthy. The Caddy container is restarting in a loop.
What would you check?
Diagnosis. docker logs motorph_caddy shows a config parse error naming a nonsensical address
like stage.https://example.com.
Someone set DOMAIN=https://example.com instead of example.com. Caddy derives every site block by
prefixing that value, so a scheme in the middle produces an address it refuses to parse — and it
refuses at startup, so the container never comes up, and the one container that publishes ports
is gone. Everything behind it becomes unreachable, despite being perfectly healthy.
This took production down three times in one night.
The guard. ensure_edge() in deploy.sh now validates the shape before compose touches the
running edge (lesson 23):
case "$domain" in
*://*|*/*|*:*) die "DOMAIN must be a bare hostname" ;;
esac
The lesson. A single container that owns all ingress is a single point of failure, so config that reaches it must be validated before it is applied — not after, when the validation failure is the outage.
Case 2 — everything is fine and nothing works
Symptom. Certificate renewal is failing. Image pulls hang. Outbound email stopped. Every
dashboard is green, SSH is responsive, curl from the host works perfectly, and DNS lookups from
inside a container resolve correctly.
What would you check?
Diagnosis. You reproduced this in lesson 13. A firewall rule was added to
DOCKER-USER without an interface match. That chain sees forwarded traffic in both directions,
so a rule meant to block inbound traffic also blocked every container's outbound connections.
DNS kept working — the red herring — because Docker's embedded resolver forwards queries to the daemon, which resolves them on the host, where nothing was blocked.
The guard. bootstrap.sh inserts the interface-qualified RETURN first, and preflight.sh
tests container egress over TCP from inside a container, with a comment saying not to use DNS for
the test.
The lesson. Test the thing, not a proxy for the thing. A check that passes while the capability is broken is worse than no check, because it actively redirects you.
Case 3 — the config change that did not take
Symptom. A fix is committed, the deploy runs green, git log on the server shows the new commit,
caddy reload reported success — and the running config is still the old one.
What would you check?
Diagnosis. The Caddyfile is bind-mounted as a single file, which resolves to an inode when
the container starts. git pull does not edit the file in place; it replaces it, creating a new
inode. The running container is still reading the old one, and reloading /etc/caddy/Caddyfile from
inside re-applies the config it already had — reporting success.
The first documentation deploy failed this way: the container was healthy, the site block was never loaded, and the CDN served 525 errors.
The guard. Both deploy.sh and docs-deploy.sh copy the file in and reload from the copy:
docker cp "$REPO_ROOT/deploy/Caddyfile" motorph_caddy:/tmp/Caddyfile
docker exec motorph_caddy caddy reload --config /tmp/Caddyfile --adapter caddyfile
The lesson. A successful-looking command is not evidence of effect. Verify the outcome, not the exit code — a habit worth having wherever a reload, a cache flush, or a restart is involved.
Case 4 — the secret that was three characters long
Symptom. Nothing. The system worked. Logins succeeded, tokens validated, no error anywhere.
What would you check? — Nothing, which is the problem. This was found by reading, not by alerting.
Diagnosis. A JWT_SECRET in an env file contained a $. Compose interpolates env-file values,
so everything from the $ onwards was read as a variable name, found nothing, and was replaced with
empty. A 64-character signing key became three characters, and the application signed authentication
tokens with it.
You measured this in lesson 09, including the part that makes it dangerous: double quotes do not help. Only single quotes do.
The same bug hit a bcrypt hash, where a prefix-only length check passed the truncated value through into production.
The guard. check_env_interpolation() in deploy.sh inspects every variable in the env file
— not a list of known-sensitive ones — and refuses to deploy if any value contains an
un-single-quoted $.
The lesson. The dangerous failures are the quiet ones. A crash gets fixed in ten minutes; a silently weakened credential survives until someone reads the file. When a class of mistake is silent, write a check rather than a convention.
Case 5 — the deploy that could not find the server
Symptom. The deploy job fails with Error: missing server host. The secret is definitely set —
you are looking at it in the settings page.
What would you check?
Diagnosis. The secret was attached to a GitHub Environment, and the job did not declare
environment:. Environment-scoped secrets are readable only by jobs that declare them, and an
unreadable secret is an empty string, not an error
(lesson 17).
The follow-up was worse: after a partial fix, every secret read as empty, and the failures moved somewhere less obvious.
The guard. The preflight step that prints ${#VAR} character counts and fails loudly on any
empty one — with a comment explaining why it uses five direct expansions instead of a clever loop:
"A diagnostic that can be wrong in the same direction as the fault it reports is worse than none."
The lesson. Absent and empty look identical in a shell. Anywhere a value arrives from outside, check it is non-empty at the boundary, and make the message name the variable.
Case 6 — the rollback that rolled nothing back
Symptom. A bad deploy is detected, the rollback runs, the log says rollback succeeded, the site is serving the old version — and it never actually rolled back.
What would you check?
Diagnosis. Two ways this happens, and both are in Track A.
First: the rollback function's steps were not &&-chained, and it was called from an if, which
suspends set -e for the whole function body (lesson 03). The image pull
failed, execution fell through to wait_healthy, which found the still-running old container
healthy, and returned success. The site looked right because nothing had changed — which is also
exactly what a successful rollback looks like.
Second: PREVIOUS_TAG had been overwritten by a re-deploy of the current tag, so the rollback target
was the broken version.
The guard. The && chain in do_rollback(), and the rule that re-deploying the current tag must
not rotate state.
The lesson. A rollback is the code path you exercise least and need most. That is why lesson 19 makes you deploy a deliberately broken version — the failure path deserves the same testing as the success path, and it almost never gets it.
The pattern
Read those six together and the same shape appears in all of them:
| The trap | Appears as |
|---|---|
| A check that can fail the same way as the thing it checks | Cases 2, 5 |
| A command that reports success without having an effect | Cases 3, 6 |
| A silent truncation or empty value | Cases 4, 5 |
| Config validated after it is applied rather than before | Case 1 |
None of these are exotic. They are all "the system told me something reassuring and it was wrong". Building the instinct to distrust a reassuring signal — and to ask what would be true if it were lying — is most of what senior operational judgement actually is.
Where to go when it is your turn
| Read | For |
|---|---|
| troubleshooting.md | The symptom index — start here at 2am |
| manual-deploy.md | Driving a deploy by hand when the pipeline cannot |
| vps-guide.md §14 | Host-level troubleshooting |
| monitoring.md | Dashboards and alerts |
Recap
- Validate config before applying it, especially for the one component that owns all ingress.
- Test the capability, not a proxy for it.
- A successful-looking command is not evidence of effect.
- Quiet failures are the expensive ones — write a check, not a convention.
- Exercise the rollback path deliberately, because production will not give you the practice.