Skip to main content

24 — preflight.sh and bootstrap.sh

Read this first: these are the two scripts a human runs by hand — one checks whether a deploy will work, the other turns a bare Ubuntu server into one that can be deployed to. They are shaped very differently from deploy.sh, and the differences are the lesson.

Time: about 35 minutes. Earned by lessons 03, 11 and 13.

preflight.sh — a checker, not a doer​

It opens differently from every other script in the repository:

set -uo pipefail

No -e. That is not an oversight, it is the shape of the tool. A checker that stops at the first problem makes you fix, re-run, fix, re-run — five round trips to a server to learn five things it could have told you at once. So it counts instead:

fails=0
warns=0
ok() { echo " ok $*"; }
bad() { echo " FAIL $*"; fails=$((fails + 1)); }
warn() { echo " warn $*"; warns=$((warns + 1)); }

and decides at the end. It also distinguishes fail from warn, which matters: a dirty git checkout is a warning (the deploy's git pull --ff-only will refuse, but you may know that); a missing required variable is a failure.

The closing line uses an expansion worth knowing:

${warns:+ ($warns warning(s))}

${var:+text} substitutes text only if var is set and non-empty — so the summary says "(3 warning(s))" when there are some and stays silent when there are none. It is the mirror of ${var:-default} from lesson 03.

Requirements it discovers rather than restates​

for v in $(grep -ohE '\$\{[A-Z_]+:\?' "$1" 2>/dev/null | sed 's/\${//;s/:?//' | sort -u); do
[ "$v" = IMAGE_TAG ] && continue # deploy.sh writes this one

It reads the compose files for the ${VAR:?} form from lesson 09 and reports every one that is empty.

This is the good idea in the file. There is no second list of required variables to fall out of date. The compose file is the specification, and the checker parses it. Add a ${NEW_THING:?} to a compose file and preflight starts checking it with no further work — which is also why docker-compose.docs.yml deliberately contains no ${VAR:?} at all: a required variable introduced there would start failing application deploys.

Note the asymmetry it handles: the app compose is checked against the target's env file, but the edge compose is always checked against .env, because there is one shared Caddy serving both environments.

Testing the right thing, from the right place​

docker run --rm --network ... curlimages/curl:latest -sS -m 15 -o /dev/null https://acme-v02.api.letsencrypt.org/directory

Container egress, tested from inside a container, over TCP. The comment says explicitly not to use DNS for this — and you measured why in lesson 13: Docker's embedded resolver forwards to the daemon, which resolves on the host, so lookups succeed with zero container connectivity.

On failure it prints the rescue command. A check that tells you what is wrong but not what to do next is only half a check.

Its other checks are lesson 12 applied: env file mode is 600 via stat -c %a, ports 80/443 are free or held by the expected container, free disk is above 2 GB, and both images actually exist in the registry (docker manifest inspect) — that last one catching a typo'd tag before it can take the stack down.

It also checks for container-name collisions across compose projects by inspecting each container's com.docker.compose.project.working_dir label. On a host running four projects, "is this container mine?" is a real question, and the name alone does not answer it.

bootstrap.sh — bare Ubuntu to deploy-ready​

Run once per server. Its most interesting property is that every mutation goes through one function:

run() {
if [ "$DRY_RUN" = 1 ]; then echo " [dry-run] $*"; else "$@"; fi
}
run_sh() {
if [ "$DRY_RUN" = 1 ]; then echo " [dry-run] sh -c '$*'"; else sh -c "$*"; fi
}

The comment: "there is no second path that could still touch the system."

That is what makes --dry-run trustworthy. A dry-run mode implemented with scattered if statements is a promise; one implemented as the single gate through which all changes pass is a guarantee. If you write one tool that has a dry-run flag, write it this way.

Three details with scars on them​

Preseeding debconf. Installing iptables-persistent normally asks a question. In an unattended run that hangs forever with no output. The script preseeds the answers and sets DEBIAN_FRONTEND=noninteractive, so the install completes rather than waiting for a human who is not there.

Refusing to accept a private key.

case "$SSH_KEY" in ssh-*|ecdsa-*) ;; *) die "that does not look like a PUBLIC key" ;; esac

case globbing again. The guard exists because "install my SSH key" is ambiguous, and the wrong answer means pasting a private key onto a server.

Hardening sshd only after a key is installed. PermitRootLogin no and PasswordAuthentication no are applied only if the key step actually ran. Reversing that order is precisely how you lock yourself out of a machine that has no console.

The firewall block​

This is lesson 13 in its production form:

if ! iptables -C DOCKER-USER ! -i "$EXT_IF" -j RETURN 2>/dev/null; then
run iptables -I DOCKER-USER 1 ! -i "$EXT_IF" -j RETURN

with EXT_IF derived as ip route show default | awk '{print $5; exit}'. Then the DROPs for 80/443 on the external interface, then ACCEPT rules for the CDN's published ranges inserted at position 2 so they precede the DROPs, then netfilter-persistent save.

Every rule is guarded with iptables -C ... ||, which checks whether the rule already exists. That is what makes the script idempotent — run it twenty times, get one rule. Bootstrap scripts get re-run, usually while something is going wrong, and one that accumulates duplicate firewall rules each time is its own incident.

The closing heredoc​

The script ends by printing what it could not do: create DNS records, write the env files, configure the GitHub secrets.

That list is the most honest thing in it. A bootstrap script that exits silently implies the server is ready. This one tells you the three manual steps between here and a working deploy — and a tool that lies about being finished is worse than one that does less.

What to steal​

  1. Checkers do not use set -e. Count failures, report them all, decide at the end.
  2. Discover requirements from the artefact, do not restate them in a second list that can drift.
  3. Route every mutation through one run() wrapper so --dry-run is a guarantee.
  4. Guard every change with an existence check so re-running is safe.
  5. Print what you did not do.

Recap​

  • preflight.sh uses set -uo pipefail without -e so it reports every problem at once, and it greps the compose files to discover what must be configured.
  • It tests container egress over TCP from inside a container, because DNS is a false positive.
  • bootstrap.sh's single run() wrapper is what makes --dry-run trustworthy, and iptables -C ... || is what makes it idempotent.
  • Harden sshd after installing a key, never before.

Next: 25 — the compose stacks and the edge.