23 — deploy.sh decoded
Read this first: this chapter reads the real deploy runner — 439 lines of bash that execute on the production server every time anything ships. You wrote a 90-line version of it in lesson 19; this is what the full one adds and why. The code is quoted inline, so you can read this without the repository open.
Time: about 45 minutes. Earned by lessons 03, 04 and 19.
Its job, in one sentence
Make failure cheap. Every design decision below comes from that, and the order of operations is the safety — not the individual commands.
The interface:
deploy/deploy.sh stage deploy <tag> # every merge to main
deploy/deploy.sh promote [tag] # prod gets staging's current tag
deploy/deploy.sh prod deploy <tag>
deploy/deploy.sh stage|prod rollback [tag] # default: PREVIOUS_TAG
deploy/deploy.sh stage|prod backup # nightly cron
deploy/deploy.sh stage|prod preflight [tag]
The lock
exec 200>"$REPO_ROOT/.deploy.lock"
flock -w 900 200 || die "another deploy/rollback is still running (lock held >15 min)"
Two unfamiliar things in two lines.
exec 200>file opens file descriptor 200 onto the lock file without running a command. Used
this way, exec redirects the shell's own descriptors, and 200 is just a high number unlikely to
collide. The descriptor stays open for the life of the script — which is the point, because the lock
is held by the descriptor.
flock -w 900 200 takes an exclusive lock on that descriptor, waiting up to 15 minutes. When the
script exits, by any route including a kill, the descriptor closes and the lock releases. There is no
cleanup to forget.
Why not rely on concurrency: in the workflow? The comment says it:
GitHub keeps one running + at most ONE pending run (a newly queued run cancels the pending one) — it is not a queue, and a canceled workflow can leave its SSH-spawned script running.
That is lesson 17's concurrency semantics, plus a fact GitHub cannot know: cancelling a workflow does not kill the process it started over SSH. The lock is the only guarantee that lives on the machine being changed.
Validating input
validate_tag() {
[[ "$1" =~ ^[A-Za-z0-9_][A-Za-z0-9._-]{0,127}$ ]] || die "invalid image tag: '$1'"
}
[[ =~ ]] from lesson 04, used where regex is genuinely needed. The comment
explains the threat:
Tags reach this script from a GitHub workflow_dispatch text field; accept only what a Docker tag may contain so nothing shell-relevant gets through.
The tag is typed by a human into a web form and then interpolated into a compose variable. Anything arriving from outside is input, not configuration.
The state file
current_tag() { { grep -s '^CURRENT_TAG=' "$STATE_FILE" || true; } | cut -d= -f2; }
previous_tag() { { grep -s '^PREVIOUS_TAG=' "$STATE_FILE" || true; } | cut -d= -f2; }
write_state() { printf 'CURRENT_TAG=%s\nPREVIOUS_TAG=%s\n' "$1" "$2" > "$STATE_FILE"; }
Rollback needs to know what to roll back to, and a plain two-line file is the right amount of
machinery. The { ... || true; } is lesson 04: grep exits 1 when the file
has no such line, which is normal on a first deploy, and under set -e that would kill the script.
The subtle part is when not to write:
Re-deploying the tag that is already current must not rotate state, or a routine re-run would set
PREVIOUS=CURRENTand destroy the rollback target.
You met this in the capstone. It is the kind of bug that only appears on the day you need the rollback.
The health gate
wait_healthy() {
local deadline s
deadline=$(( $(date +%s) + HEALTH_TIMEOUT ))
while :; do
s=$(docker inspect -f '{{.State.Health.Status}}' "$BACKEND" 2>/dev/null || echo missing)
[ "$s" = healthy ] && { echo "$BACKEND healthy"; return 0; }
if [ "$(date +%s)" -ge "$deadline" ]; then
echo "$BACKEND not healthy after ${HEALTH_TIMEOUT}s (last status: $s); recent logs:" >&2
docker logs --tail 100 "$BACKEND" >&2 || true
return 1
fi
sleep 5
done
}
Worth reading line by line:
- A deadline, not a counter.
$(( $(date +%s) + HEALTH_TIMEOUT ))computes an absolute time, so the total wait is correct regardless of how long eachdocker inspecttakes. || echo missingturns "no such container" into a value rather than an error — otherwiseset -ewould kill the script instead of retrying.- Logs go to stderr before returning failure, from lesson 12: the container may not survive long enough for a human to look.
|| trueon the log dump, because failing to fetch logs must not mask the real failure.
The status it reads comes from the container healthcheck in the compose file — the one you wrote in lesson 08. This is where that pays off.
The && chain
do_rollback() {
set_env_tag "$1" \
&& IMAGE_TAG="$1" compose_rb pull -q \
&& IMAGE_TAG="$1" compose_rb up -d \
&& wait_healthy
}
The comment above it is the best four lines of teaching in the repository:
Every step is &&-chained: callers run this inside
if, which suspendsset -efor the whole function body — without the chain, a failed pull or up would fall through to wait_healthy, which would then bless the [still-running old container as a successful rollback].
That is lesson 03's trap, in production, with real consequences: a rollback that reports success while doing nothing, during an incident, when you are least able to check.
Note also IMAGE_TAG="$1" compose_rb pull — the prefix assignment from
lesson 09. The OS environment beats --env-file, so the tag is pinned
explicitly for that one command rather than trusted to the file.
The deploy order
The sequence, and the reason for each position:
| Step | Why here |
|---|---|
validate_tag | Reject bad input before touching anything |
ensure_edge | The proxy must exist before anything can be reached |
Raise HEALTH_TIMEOUT to 600 if no CURRENT_TAG | A first boot runs every migration |
compose pull -q | Before stopping anything. A typo'd tag or an auth failure dies with the old version still serving |
backup_db "pre-$new" | Before the change, not after |
set_env_tag then compose up -d | Now the swap |
wait_healthy | The gate |
Success → rotate_state; failure → roll back and exit 1 regardless | A deploy that rolled back did not succeed |
That last row is the one people get wrong. The site is fine after an auto-rollback — and the run must still be red, because "the new version is live" is false. Green has to mean one thing.
ensure_edge: three outages in one function
case "$domain" in
*://*|*/*|*:*) die "DOMAIN must be a bare hostname" ;;
www.*) die "DOMAIN must not start with www." ;;
esac
case globbing from lesson 04. If DOMAIN contains a scheme, Caddy builds a
nonsense address like stage.https://example.com, refuses its config, and crash-loops the edge —
taking the live site down. The comment records that this happened three times in one night. The
check runs before compose touches the running edge, so a bad value fails without disturbing
anything.
The same function validates that the Grafana basic-auth hash is a complete bcrypt string rather than a prefix, because a compose-truncated hash once reached production — the truncation trap from lesson 09, in a second form.
And then:
if docker cp "$REPO_ROOT/deploy/Caddyfile" motorph_caddy:/tmp/Caddyfile 2>/dev/null; then
docker exec motorph_caddy caddy reload --config /tmp/Caddyfile --adapter caddyfile 2>/dev/null \
|| echo "note: caddy reload skipped (fresh start loads the config at boot)"
fi
Why copy the file in rather than reload the mounted path? A single-file bind mount resolves to an
inode when the container starts. git pull replaces the file rather than editing it in place,
so the running container keeps reading the version it booted with — and a reload against that path
is a no-op that reports success. That is exactly how the first documentation deploy left the CDN
serving 525 errors: container healthy, config never loaded.
The optional-flags array
compose() {
local extra; mapfile -t extra < <(monitoring_args)
docker compose --project-directory "$REPO_ROOT" --env-file "$ENV_FILE" -f "$APP_YML" "${extra[@]}" "$@"
}
Lesson 04, in production: monitoring_args prints two lines or nothing,
mapfile -t reads them into an array, and "${extra[@]}" expands to zero or two words. A plain
string would expand to one empty argument when monitoring is off, and compose would reject it.
--project-directory "$REPO_ROOT" is not decoration either: without it, compose resolves relative
mount paths against deploy/, which breaks the Caddyfile mount and the monitoring overlay's
./infra/monitoring/... paths.
What to steal
If you take four things from this file into your own work:
flockon a descriptor, because concurrency controls upstream do not bind the machine.- Pull before you stop anything, so the common failure is free.
&&-chain any function called fromif, or errexit is not protecting it.- Roll back and still exit non-zero. Green must mean one thing.
Recap
- The script's job is to make failure cheap, and the ordering is the safety.
exec 200>file+flockis a lock that cannot be left behind.&&-chaining defeats the errexit-inside-iftrap, which would otherwise report a rollback that never happened.- Several guards exist because of specific outages: the
DOMAINshape check, the bcrypt length check, and the copy-then-reload Caddy dance.
Next: 24 — preflight and bootstrap.