Skip to main content

23 — deploy.sh decoded

Read this first: this chapter reads the real deploy runner — 439 lines of bash that execute on the production server every time anything ships. You wrote a 90-line version of it in lesson 19; this is what the full one adds and why. The code is quoted inline, so you can read this without the repository open.

Time: about 45 minutes. Earned by lessons 03, 04 and 19.

Its job, in one sentence​

Make failure cheap. Every design decision below comes from that, and the order of operations is the safety — not the individual commands.

The interface:

deploy/deploy.sh stage deploy <tag> # every merge to main
deploy/deploy.sh promote [tag] # prod gets staging's current tag
deploy/deploy.sh prod deploy <tag>
deploy/deploy.sh stage|prod rollback [tag] # default: PREVIOUS_TAG
deploy/deploy.sh stage|prod backup # nightly cron
deploy/deploy.sh stage|prod preflight [tag]

The lock​

exec 200>"$REPO_ROOT/.deploy.lock"
flock -w 900 200 || die "another deploy/rollback is still running (lock held >15 min)"

Two unfamiliar things in two lines.

exec 200>file opens file descriptor 200 onto the lock file without running a command. Used this way, exec redirects the shell's own descriptors, and 200 is just a high number unlikely to collide. The descriptor stays open for the life of the script — which is the point, because the lock is held by the descriptor.

flock -w 900 200 takes an exclusive lock on that descriptor, waiting up to 15 minutes. When the script exits, by any route including a kill, the descriptor closes and the lock releases. There is no cleanup to forget.

Why not rely on concurrency: in the workflow? The comment says it:

GitHub keeps one running + at most ONE pending run (a newly queued run cancels the pending one) — it is not a queue, and a canceled workflow can leave its SSH-spawned script running.

That is lesson 17's concurrency semantics, plus a fact GitHub cannot know: cancelling a workflow does not kill the process it started over SSH. The lock is the only guarantee that lives on the machine being changed.

Validating input​

validate_tag() {
[[ "$1" =~ ^[A-Za-z0-9_][A-Za-z0-9._-]{0,127}$ ]] || die "invalid image tag: '$1'"
}

[[ =~ ]] from lesson 04, used where regex is genuinely needed. The comment explains the threat:

Tags reach this script from a GitHub workflow_dispatch text field; accept only what a Docker tag may contain so nothing shell-relevant gets through.

The tag is typed by a human into a web form and then interpolated into a compose variable. Anything arriving from outside is input, not configuration.

The state file​

current_tag() { { grep -s '^CURRENT_TAG=' "$STATE_FILE" || true; } | cut -d= -f2; }
previous_tag() { { grep -s '^PREVIOUS_TAG=' "$STATE_FILE" || true; } | cut -d= -f2; }
write_state() { printf 'CURRENT_TAG=%s\nPREVIOUS_TAG=%s\n' "$1" "$2" > "$STATE_FILE"; }

Rollback needs to know what to roll back to, and a plain two-line file is the right amount of machinery. The { ... || true; } is lesson 04: grep exits 1 when the file has no such line, which is normal on a first deploy, and under set -e that would kill the script.

The subtle part is when not to write:

Re-deploying the tag that is already current must not rotate state, or a routine re-run would set PREVIOUS=CURRENT and destroy the rollback target.

You met this in the capstone. It is the kind of bug that only appears on the day you need the rollback.

The health gate​

wait_healthy() {
local deadline s
deadline=$(( $(date +%s) + HEALTH_TIMEOUT ))
while :; do
s=$(docker inspect -f '{{.State.Health.Status}}' "$BACKEND" 2>/dev/null || echo missing)
[ "$s" = healthy ] && { echo "$BACKEND healthy"; return 0; }
if [ "$(date +%s)" -ge "$deadline" ]; then
echo "$BACKEND not healthy after ${HEALTH_TIMEOUT}s (last status: $s); recent logs:" >&2
docker logs --tail 100 "$BACKEND" >&2 || true
return 1
fi
sleep 5
done
}

Worth reading line by line:

  • A deadline, not a counter. $(( $(date +%s) + HEALTH_TIMEOUT )) computes an absolute time, so the total wait is correct regardless of how long each docker inspect takes.
  • || echo missing turns "no such container" into a value rather than an error — otherwise set -e would kill the script instead of retrying.
  • Logs go to stderr before returning failure, from lesson 12: the container may not survive long enough for a human to look.
  • || true on the log dump, because failing to fetch logs must not mask the real failure.

The status it reads comes from the container healthcheck in the compose file — the one you wrote in lesson 08. This is where that pays off.

The && chain​

do_rollback() {
set_env_tag "$1" \
&& IMAGE_TAG="$1" compose_rb pull -q \
&& IMAGE_TAG="$1" compose_rb up -d \
&& wait_healthy
}

The comment above it is the best four lines of teaching in the repository:

Every step is &&-chained: callers run this inside if, which suspends set -e for the whole function body — without the chain, a failed pull or up would fall through to wait_healthy, which would then bless the [still-running old container as a successful rollback].

That is lesson 03's trap, in production, with real consequences: a rollback that reports success while doing nothing, during an incident, when you are least able to check.

Note also IMAGE_TAG="$1" compose_rb pull — the prefix assignment from lesson 09. The OS environment beats --env-file, so the tag is pinned explicitly for that one command rather than trusted to the file.

The deploy order​

The sequence, and the reason for each position:

StepWhy here
validate_tagReject bad input before touching anything
ensure_edgeThe proxy must exist before anything can be reached
Raise HEALTH_TIMEOUT to 600 if no CURRENT_TAGA first boot runs every migration
compose pull -qBefore stopping anything. A typo'd tag or an auth failure dies with the old version still serving
backup_db "pre-$new"Before the change, not after
set_env_tag then compose up -dNow the swap
wait_healthyThe gate
Success → rotate_state; failure → roll back and exit 1 regardlessA deploy that rolled back did not succeed

That last row is the one people get wrong. The site is fine after an auto-rollback — and the run must still be red, because "the new version is live" is false. Green has to mean one thing.

ensure_edge: three outages in one function​

case "$domain" in
*://*|*/*|*:*) die "DOMAIN must be a bare hostname" ;;
www.*) die "DOMAIN must not start with www." ;;
esac

case globbing from lesson 04. If DOMAIN contains a scheme, Caddy builds a nonsense address like stage.https://example.com, refuses its config, and crash-loops the edge — taking the live site down. The comment records that this happened three times in one night. The check runs before compose touches the running edge, so a bad value fails without disturbing anything.

The same function validates that the Grafana basic-auth hash is a complete bcrypt string rather than a prefix, because a compose-truncated hash once reached production — the truncation trap from lesson 09, in a second form.

And then:

if docker cp "$REPO_ROOT/deploy/Caddyfile" motorph_caddy:/tmp/Caddyfile 2>/dev/null; then
docker exec motorph_caddy caddy reload --config /tmp/Caddyfile --adapter caddyfile 2>/dev/null \
|| echo "note: caddy reload skipped (fresh start loads the config at boot)"
fi

Why copy the file in rather than reload the mounted path? A single-file bind mount resolves to an inode when the container starts. git pull replaces the file rather than editing it in place, so the running container keeps reading the version it booted with — and a reload against that path is a no-op that reports success. That is exactly how the first documentation deploy left the CDN serving 525 errors: container healthy, config never loaded.

The optional-flags array​

compose() {
local extra; mapfile -t extra < <(monitoring_args)
docker compose --project-directory "$REPO_ROOT" --env-file "$ENV_FILE" -f "$APP_YML" "${extra[@]}" "$@"
}

Lesson 04, in production: monitoring_args prints two lines or nothing, mapfile -t reads them into an array, and "${extra[@]}" expands to zero or two words. A plain string would expand to one empty argument when monitoring is off, and compose would reject it.

--project-directory "$REPO_ROOT" is not decoration either: without it, compose resolves relative mount paths against deploy/, which breaks the Caddyfile mount and the monitoring overlay's ./infra/monitoring/... paths.

What to steal​

If you take four things from this file into your own work:

  1. flock on a descriptor, because concurrency controls upstream do not bind the machine.
  2. Pull before you stop anything, so the common failure is free.
  3. &&-chain any function called from if, or errexit is not protecting it.
  4. Roll back and still exit non-zero. Green must mean one thing.

Recap​

  • The script's job is to make failure cheap, and the ordering is the safety.
  • exec 200>file + flock is a lock that cannot be left behind.
  • &&-chaining defeats the errexit-inside-if trap, which would otherwise report a rollback that never happened.
  • Several guards exist because of specific outages: the DOMAIN shape check, the bcrypt length check, and the copy-then-reload Caddy dance.

Next: 24 — preflight and bootstrap.