03 — Scripts that fail safely
Read this first: this is the most important lesson in the course. It explains every letter of
set -euo pipefail — the line that opens every script in deploy/ — by showing you
exactly what goes wrong without it. If you only ever read one page here, read this one.
Time: about 45 minutes. Assumes lesson 02.
The problem, stated plainly
A shell script does not stop when a command fails. It runs the next line.
That sentence sounds harmless until you picture a deploy script whose line 4 is "change into the release directory" and whose line 5 is "delete everything here".
Set up
mkdir -p ~/devops-course/l03
cd ~/devops-course/l03
-e — stop at the first failure
Create bad.sh:
#!/usr/bin/env bash
cd /nope/does/not/exist
echo "still running, and about to delete things"
Predict: cd into a directory that does not exist will fail. Does line 3 run? What is the
script's final exit code?
chmod +x bad.sh
./bad.sh
echo "exit = $?"
./bad.sh: line 2: cd: /nope/does/not/exist: No such file or directory
still running, and about to delete things
exit = 0
Read that carefully, because all three lines are bad news:
cdfailed and said so.- The script kept going anyway. In a real script, line 3 would now be operating on whatever directory you happened to be in — most likely your home directory.
- The script reported success. Exit code 0. A CI job running this goes green.
That is not a hypothetical failure mode. It is the classic one. Now add one line:
#!/usr/bin/env bash
set -e
cd /nope/does/not/exist
echo "you should NOT see this"
./good.sh
echo "exit = $?"
./good.sh: line 3: cd: /nope/does/not/exist: No such file or directory
exit = 1
The script stopped, and it reported failure. set -e — also spelled set -o errexit — means
"exit immediately if any command returns non-zero".
-u — refuse to use a variable that was never set
Recall from lesson 02 that a missing argument is silently empty. Here is why
that matters. Create u1.sh:
#!/usr/bin/env bash
echo "deploying to: $TARGET_DIR"
rm -rf "$TARGET_DIR/old"
echo done
Predict: TARGET_DIR is not set. What does line 3 actually delete?
./u1.sh
echo "exit = $?"
deploying to:
done
exit = 0
$TARGET_DIR expanded to nothing, so line 3 ran rm -rf "/old". On this machine there was no
/old, so nothing happened and the script cheerfully reported success. On a machine where /old
exists — or with a slightly different path — it would have deleted something at the filesystem root.
Now with -u:
#!/usr/bin/env bash
set -u
echo "deploying to: $TARGET_DIR"
./u2.sh
echo "exit = $?"
./u2.sh: line 3: TARGET_DIR: unbound variable
exit = 1
set -u (nounset) makes referring to an unset variable a fatal error. It converts "silently
deleted the wrong thing" into "refused to start". That is the trade you want every time.
When you deliberately want a variable that might be unset, say so explicitly:
| Form | Means |
|---|---|
${VAR:-default} | Use VAR, or default if it is unset or empty |
${VAR:-} | Use VAR, or empty — "I know this may be unset, and that's fine" |
${VAR:?some message} | Fail with some message if unset or empty |
That third form is not just a shell trick. You will meet it again in
lesson 08 as ${IMAGE_TAG:?} inside a compose file, where it turns a required
setting into a machine-readable contract.
-o pipefail — stop the pipeline from lying
This is the one you already met in lesson 01. Recall:
grep delta words.txt | wc -l
echo "pipeline exit = $?"
reported 0 even though grep failed, because a pipeline's exit code is its last command's
exit code. Combine that with set -e and you get the worst outcome available: set -e is watching
for a non-zero exit, the pipeline hands it a zero, and the script sails past a step that did nothing.
set -o pipefail changes the rule to: the pipeline fails if any stage fails.
Putting the three together gives you the line at the top of every script in this repository:
set -euo pipefail
| Letter | Long name | Without it |
|---|---|---|
-e | errexit | The script continues past a failed command |
-u | nounset | An unset variable expands to nothing, silently |
-o pipefail | pipefail | A failure inside a pipeline is invisible |
Break it on purpose: the three places -e does not apply
Here is the part that trips up people who think they already know this. set -e is suspended
inside a condition. Specifically: inside if, inside && and || chains, and after !.
That is deliberate and correct — if grep -q foo file; then has to be allowed to have grep fail,
or if would be useless. But it has a consequence that is genuinely surprising.
Predict: this script has set -e. The function's first command fails. Does echo run? What is
printed?
#!/usr/bin/env bash
set -e
check() {
false
echo "still here"
}
if check; then
echo "check passed"
else
echo "check failed"
fi
Most people say the function dies at false. Run it:
still here
check passed
Both wrong answers at once. Because check was called from an if, errexit was suspended for
the entire body of the function — not just for the call. So false did not stop it, echo ran,
and the function's exit code became the exit code of its last command, which succeeded. The if
concluded that the check passed.
Now imagine check is really do_rollback, its first command is "pull the previous image", and
its last command is "wait for the container to become healthy". The pull fails. The old container
is still running and still healthy. The function returns success. You have just reported a
successful rollback that never happened.
Fix it
Chain the steps with && so that any failure short-circuits the rest:
do_rollback() {
set_env_tag "$1" \
&& pull_image "$1" \
&& start_container \
&& wait_healthy
}
Now if pull_image fails, nothing after it runs, and the function returns non-zero. This is exactly
what the real script does. From deploy/deploy.sh:
do_rollback() {
set_env_tag "$1" \
&& IMAGE_TAG="$1" compose_rb pull -q \
&& IMAGE_TAG="$1" compose_rb up -d \
&& wait_healthy
}
and it carries a comment explaining precisely the trap you just fell into.
The other escape hatch: || true
The mirror-image problem: sometimes a non-zero exit is expected and should not kill the script.
grep returning 1 because a value is legitimately absent is the common case.
current_tag() { { grep -s '^CURRENT_TAG=' "$STATE_FILE" || true; } | cut -d= -f2; }
That is a real line from deploy.sh. The || true says "a no-match here is normal, not a failure".
The braces group the command so the || true applies to grep and not to the whole pipeline.
Use || true deliberately and rarely. Every one of them is a place where you have switched the
safety off, so each deserves to be obviously intentional.
Where this shows up in MotorPH
- Every script in deploy/ and scripts/ starts with
set -euo pipefail. - Except one, on purpose: deploy/preflight.sh uses
set -uo pipefailwith no-e, because it is a checker. A checker that stops at the first problem makes you fix and re-run five times; this one counts every failure and reports them all at once. - The
&&-chaineddo_rollbackabove is the real function, and the comment above it in the source is worth reading now that you know what it is protecting against.
Recap
set -estops on the first failure,set -urefuses unset variables,set -o pipefailstops a pipeline from hiding a failed stage. Together:set -euo pipefail.set -eis suspended insideif,&&,||and!— including for the whole body of a function called from anif. Chain steps with&&when a function must fail as a unit.|| truedeliberately tolerates an expected non-zero exit. Every use is a switched-off safety and should look like one.
Next: 05 — Containers 101.