Skip to main content

26 — When it breaks

Read this first: this chapter is a set of real incidents from this system, each written as a symptom first. Read the symptom, work out what you would check, then read on. Every one of these actually happened, and every one left a guard behind in the code that you have already met.

Time: about 35 minutes. Earned by the whole of Track A.

How to use this​

For each case: cover the diagnosis, read the symptom, write down your first three checks. Then compare. The point is not to memorise five incidents — it is to notice that your triage instinct is now producing the right questions.

The general order, from lesson 12: what is running, what did it say, what holds the port, is the disk full.


Case 1 — the site is down and the edge is crash-looping​

Symptom. A routine deploy. Now nothing responds on any hostname — not the app, not staging, not the docs. The application containers are healthy. The Caddy container is restarting in a loop.

What would you check?

Diagnosis. docker logs motorph_caddy shows a config parse error naming a nonsensical address like stage.https://example.com.

Someone set DOMAIN=https://example.com instead of example.com. Caddy derives every site block by prefixing that value, so a scheme in the middle produces an address it refuses to parse — and it refuses at startup, so the container never comes up, and the one container that publishes ports is gone. Everything behind it becomes unreachable, despite being perfectly healthy.

This took production down three times in one night.

The guard. ensure_edge() in deploy.sh now validates the shape before compose touches the running edge (lesson 23):

case "$domain" in
*://*|*/*|*:*) die "DOMAIN must be a bare hostname" ;;
esac

The lesson. A single container that owns all ingress is a single point of failure, so config that reaches it must be validated before it is applied — not after, when the validation failure is the outage.


Case 2 — everything is fine and nothing works​

Symptom. Certificate renewal is failing. Image pulls hang. Outbound email stopped. Every dashboard is green, SSH is responsive, curl from the host works perfectly, and DNS lookups from inside a container resolve correctly.

What would you check?

Diagnosis. You reproduced this in lesson 13. A firewall rule was added to DOCKER-USER without an interface match. That chain sees forwarded traffic in both directions, so a rule meant to block inbound traffic also blocked every container's outbound connections.

DNS kept working — the red herring — because Docker's embedded resolver forwards queries to the daemon, which resolves them on the host, where nothing was blocked.

The guard. bootstrap.sh inserts the interface-qualified RETURN first, and preflight.sh tests container egress over TCP from inside a container, with a comment saying not to use DNS for the test.

The lesson. Test the thing, not a proxy for the thing. A check that passes while the capability is broken is worse than no check, because it actively redirects you.


Case 3 — the config change that did not take​

Symptom. A fix is committed, the deploy runs green, git log on the server shows the new commit, caddy reload reported success — and the running config is still the old one.

What would you check?

Diagnosis. The Caddyfile is bind-mounted as a single file, which resolves to an inode when the container starts. git pull does not edit the file in place; it replaces it, creating a new inode. The running container is still reading the old one, and reloading /etc/caddy/Caddyfile from inside re-applies the config it already had — reporting success.

The first documentation deploy failed this way: the container was healthy, the site block was never loaded, and the CDN served 525 errors.

The guard. Both deploy.sh and docs-deploy.sh copy the file in and reload from the copy:

docker cp "$REPO_ROOT/deploy/Caddyfile" motorph_caddy:/tmp/Caddyfile
docker exec motorph_caddy caddy reload --config /tmp/Caddyfile --adapter caddyfile

The lesson. A successful-looking command is not evidence of effect. Verify the outcome, not the exit code — a habit worth having wherever a reload, a cache flush, or a restart is involved.


Case 4 — the secret that was three characters long​

Symptom. Nothing. The system worked. Logins succeeded, tokens validated, no error anywhere.

What would you check? — Nothing, which is the problem. This was found by reading, not by alerting.

Diagnosis. A JWT_SECRET in an env file contained a $. Compose interpolates env-file values, so everything from the $ onwards was read as a variable name, found nothing, and was replaced with empty. A 64-character signing key became three characters, and the application signed authentication tokens with it.

You measured this in lesson 09, including the part that makes it dangerous: double quotes do not help. Only single quotes do.

The same bug hit a bcrypt hash, where a prefix-only length check passed the truncated value through into production.

The guard. check_env_interpolation() in deploy.sh inspects every variable in the env file — not a list of known-sensitive ones — and refuses to deploy if any value contains an un-single-quoted $.

The lesson. The dangerous failures are the quiet ones. A crash gets fixed in ten minutes; a silently weakened credential survives until someone reads the file. When a class of mistake is silent, write a check rather than a convention.


Case 5 — the deploy that could not find the server​

Symptom. The deploy job fails with Error: missing server host. The secret is definitely set — you are looking at it in the settings page.

What would you check?

Diagnosis. The secret was attached to a GitHub Environment, and the job did not declare environment:. Environment-scoped secrets are readable only by jobs that declare them, and an unreadable secret is an empty string, not an error (lesson 17).

The follow-up was worse: after a partial fix, every secret read as empty, and the failures moved somewhere less obvious.

The guard. The preflight step that prints ${#VAR} character counts and fails loudly on any empty one — with a comment explaining why it uses five direct expansions instead of a clever loop: "A diagnostic that can be wrong in the same direction as the fault it reports is worse than none."

The lesson. Absent and empty look identical in a shell. Anywhere a value arrives from outside, check it is non-empty at the boundary, and make the message name the variable.


Case 6 — the rollback that rolled nothing back​

Symptom. A bad deploy is detected, the rollback runs, the log says rollback succeeded, the site is serving the old version — and it never actually rolled back.

What would you check?

Diagnosis. Two ways this happens, and both are in Track A.

First: the rollback function's steps were not &&-chained, and it was called from an if, which suspends set -e for the whole function body (lesson 03). The image pull failed, execution fell through to wait_healthy, which found the still-running old container healthy, and returned success. The site looked right because nothing had changed — which is also exactly what a successful rollback looks like.

Second: PREVIOUS_TAG had been overwritten by a re-deploy of the current tag, so the rollback target was the broken version.

The guard. The && chain in do_rollback(), and the rule that re-deploying the current tag must not rotate state.

The lesson. A rollback is the code path you exercise least and need most. That is why lesson 19 makes you deploy a deliberately broken version — the failure path deserves the same testing as the success path, and it almost never gets it.


The pattern​

Read those six together and the same shape appears in all of them:

The trapAppears as
A check that can fail the same way as the thing it checksCases 2, 5
A command that reports success without having an effectCases 3, 6
A silent truncation or empty valueCases 4, 5
Config validated after it is applied rather than beforeCase 1

None of these are exotic. They are all "the system told me something reassuring and it was wrong". Building the instinct to distrust a reassuring signal — and to ask what would be true if it were lying — is most of what senior operational judgement actually is.

Where to go when it is your turn​

ReadFor
troubleshooting.mdThe symptom index — start here at 2am
manual-deploy.mdDriving a deploy by hand when the pipeline cannot
vps-guide.md §14Host-level troubleshooting
monitoring.mdDashboards and alerts

Recap​

  • Validate config before applying it, especially for the one component that owns all ingress.
  • Test the capability, not a proxy for it.
  • A successful-looking command is not evidence of effect.
  • Quiet failures are the expensive ones — write a check, not a convention.
  • Exercise the rollback path deliberately, because production will not give you the practice.