Skip to main content

03 — Scripts that fail safely

Read this first: this is the most important lesson in the course. It explains every letter of set -euo pipefail — the line that opens every script in deploy/ — by showing you exactly what goes wrong without it. If you only ever read one page here, read this one.

Time: about 45 minutes. Assumes lesson 02.

The problem, stated plainly​

A shell script does not stop when a command fails. It runs the next line.

That sentence sounds harmless until you picture a deploy script whose line 4 is "change into the release directory" and whose line 5 is "delete everything here".

Set up​

mkdir -p ~/devops-course/l03
cd ~/devops-course/l03

-e — stop at the first failure​

Create bad.sh:

#!/usr/bin/env bash
cd /nope/does/not/exist
echo "still running, and about to delete things"

Predict: cd into a directory that does not exist will fail. Does line 3 run? What is the script's final exit code?

chmod +x bad.sh
./bad.sh
echo "exit = $?"
./bad.sh: line 2: cd: /nope/does/not/exist: No such file or directory
still running, and about to delete things
exit = 0

Read that carefully, because all three lines are bad news:

  1. cd failed and said so.
  2. The script kept going anyway. In a real script, line 3 would now be operating on whatever directory you happened to be in — most likely your home directory.
  3. The script reported success. Exit code 0. A CI job running this goes green.

That is not a hypothetical failure mode. It is the classic one. Now add one line:

#!/usr/bin/env bash
set -e
cd /nope/does/not/exist
echo "you should NOT see this"
./good.sh
echo "exit = $?"
./good.sh: line 3: cd: /nope/does/not/exist: No such file or directory
exit = 1

The script stopped, and it reported failure. set -e — also spelled set -o errexit — means "exit immediately if any command returns non-zero".

-u — refuse to use a variable that was never set​

Recall from lesson 02 that a missing argument is silently empty. Here is why that matters. Create u1.sh:

#!/usr/bin/env bash
echo "deploying to: $TARGET_DIR"
rm -rf "$TARGET_DIR/old"
echo done

Predict: TARGET_DIR is not set. What does line 3 actually delete?

./u1.sh
echo "exit = $?"
deploying to:
done
exit = 0

$TARGET_DIR expanded to nothing, so line 3 ran rm -rf "/old". On this machine there was no /old, so nothing happened and the script cheerfully reported success. On a machine where /old exists — or with a slightly different path — it would have deleted something at the filesystem root.

Now with -u:

#!/usr/bin/env bash
set -u
echo "deploying to: $TARGET_DIR"
./u2.sh
echo "exit = $?"
./u2.sh: line 3: TARGET_DIR: unbound variable
exit = 1

set -u (nounset) makes referring to an unset variable a fatal error. It converts "silently deleted the wrong thing" into "refused to start". That is the trade you want every time.

When you deliberately want a variable that might be unset, say so explicitly:

FormMeans
${VAR:-default}Use VAR, or default if it is unset or empty
${VAR:-}Use VAR, or empty — "I know this may be unset, and that's fine"
${VAR:?some message}Fail with some message if unset or empty

That third form is not just a shell trick. You will meet it again in lesson 08 as ${IMAGE_TAG:?} inside a compose file, where it turns a required setting into a machine-readable contract.

-o pipefail — stop the pipeline from lying​

This is the one you already met in lesson 01. Recall:

grep delta words.txt | wc -l
echo "pipeline exit = $?"

reported 0 even though grep failed, because a pipeline's exit code is its last command's exit code. Combine that with set -e and you get the worst outcome available: set -e is watching for a non-zero exit, the pipeline hands it a zero, and the script sails past a step that did nothing.

set -o pipefail changes the rule to: the pipeline fails if any stage fails.

Putting the three together gives you the line at the top of every script in this repository:

set -euo pipefail
LetterLong nameWithout it
-eerrexitThe script continues past a failed command
-unounsetAn unset variable expands to nothing, silently
-o pipefailpipefailA failure inside a pipeline is invisible

Break it on purpose: the three places -e does not apply​

Here is the part that trips up people who think they already know this. set -e is suspended inside a condition. Specifically: inside if, inside && and || chains, and after !.

That is deliberate and correct — if grep -q foo file; then has to be allowed to have grep fail, or if would be useless. But it has a consequence that is genuinely surprising.

Predict: this script has set -e. The function's first command fails. Does echo run? What is printed?

#!/usr/bin/env bash
set -e

check() {
false
echo "still here"
}

if check; then
echo "check passed"
else
echo "check failed"
fi

Most people say the function dies at false. Run it:

still here
check passed

Both wrong answers at once. Because check was called from an if, errexit was suspended for the entire body of the function — not just for the call. So false did not stop it, echo ran, and the function's exit code became the exit code of its last command, which succeeded. The if concluded that the check passed.

Now imagine check is really do_rollback, its first command is "pull the previous image", and its last command is "wait for the container to become healthy". The pull fails. The old container is still running and still healthy. The function returns success. You have just reported a successful rollback that never happened.

Fix it​

Chain the steps with && so that any failure short-circuits the rest:

do_rollback() {
set_env_tag "$1" \
&& pull_image "$1" \
&& start_container \
&& wait_healthy
}

Now if pull_image fails, nothing after it runs, and the function returns non-zero. This is exactly what the real script does. From deploy/deploy.sh:

do_rollback() {
set_env_tag "$1" \
&& IMAGE_TAG="$1" compose_rb pull -q \
&& IMAGE_TAG="$1" compose_rb up -d \
&& wait_healthy
}

and it carries a comment explaining precisely the trap you just fell into.

The other escape hatch: || true​

The mirror-image problem: sometimes a non-zero exit is expected and should not kill the script. grep returning 1 because a value is legitimately absent is the common case.

current_tag() { { grep -s '^CURRENT_TAG=' "$STATE_FILE" || true; } | cut -d= -f2; }

That is a real line from deploy.sh. The || true says "a no-match here is normal, not a failure". The braces group the command so the || true applies to grep and not to the whole pipeline.

Use || true deliberately and rarely. Every one of them is a place where you have switched the safety off, so each deserves to be obviously intentional.

Where this shows up in MotorPH​

  • Every script in deploy/ and scripts/ starts with set -euo pipefail.
  • Except one, on purpose: deploy/preflight.sh uses set -uo pipefail with no -e, because it is a checker. A checker that stops at the first problem makes you fix and re-run five times; this one counts every failure and reports them all at once.
  • The &&-chained do_rollback above is the real function, and the comment above it in the source is worth reading now that you know what it is protecting against.

Recap​

  • set -e stops on the first failure, set -u refuses unset variables, set -o pipefail stops a pipeline from hiding a failed stage. Together: set -euo pipefail.
  • set -e is suspended inside if, &&, || and ! — including for the whole body of a function called from an if. Chain steps with && when a function must fail as a unit.
  • || true deliberately tolerates an expected non-zero exit. Every use is a switched-off safety and should look like one.

Next: 05 — Containers 101.