19 — Capstone: ship it, break it, roll it back
Read this first: this is the exam. You build a complete pipeline from an empty repository — Dockerfile, compose file, deploy script, workflow — deploy a working version, deploy a broken one, and watch your own health gate catch it and roll back. Write every line yourself. Reading a finished pipeline teaches you nothing that typing one does.
Time: three to four hours, and it is worth doing in one sitting. Assumes lessons 01 through 18.
What "finished" means
You can do all of this without help:
- Build an image from a Dockerfile you wrote, with a working layer cache
- Run a multi-service stack with a health gate that means something
- Push an image to a registry from CI, tagged by commit
- Deploy it to a server over SSH, with configuration from secrets
- Detect a bad deploy automatically and roll back
- Explain, out loud, why each of those steps is in that order
The last one is the real test. If you can build it but not explain it, you have copied it.
Stage 1 — the application
Start a new public repository called greeter. Public gets you free Actions minutes and pullable
packages; it is also a second portfolio artifact.
Copy in the toy app from learn-devops/greeter-api and learn-devops/greeter-web, but delete
every Dockerfile, compose file and workflow. Those are the answers.
Confirm it runs locally before containerising anything. A pipeline debugged on top of a broken app is two problems wearing one coat.
Stage 2 — Dockerfiles
Write both from scratch. Requirements, not solutions:
greeter-api/Dockerfile — multi-stage, Maven build → JRE runtime, pom.xml copied before src,
Boot layer extraction, non-root user, MaxRAMPercentage, exec-form ENTRYPOINT, APP_VERSION as a
build arg.
greeter-web/Dockerfile — multi-stage, Node build → nginx runtime, manifest copied before
source, an nginx config that falls back to index.html and proxies /api/.
Verify by measuring, not by looking:
time docker build -t greeter-api:dev ./greeter-api # note the time
# edit one line of Java
time docker build -t greeter-api:dev ./greeter-api # should be much faster
If the second build re-resolves dependencies, your COPY order is wrong. Go back to
lesson 06.
Stage 3 — two compose files
docker-compose.yml (local): build:, published ports, throwaway credentials, healthchecks,
condition: service_healthy.
deploy/docker-compose.prod.yml (server): the same services with
image: ghcr.io/<you>/greeter-api:${IMAGE_TAG:?set by deploy.sh — do not run compose up by hand}
and no published ports except the frontend's.
Predict, then check: run docker compose -f deploy/docker-compose.prod.yml config with no
IMAGE_TAG set. Write down the message you expect before you see it.
Stage 4 — deploy.sh
This is the heart, roughly 90 lines. It runs on the server:
./deploy.sh <tag> deploy that tag
./deploy.sh rollback go back to the previous tag
Requirements:
| Must | Because |
|---|---|
set -euo pipefail | Lesson 03 |
flock on a file descriptor | Two deploys must never interleave — and a cancelled workflow can leave its script running |
A .deploy-state file with CURRENT_TAG / PREVIOUS_TAG | Rollback needs a target |
| Pull before stopping anything | A bad tag must fail with the old version still serving |
IMAGE_TAG="$tag" docker compose up -d | OS env beats --env-file (lesson 09) |
wait_healthy polling docker inspect against a deadline | Lesson 18 |
| Roll back on failure and exit non-zero | A deploy that rolled back did not succeed |
| Dump logs before giving up | The container may not survive (lesson 12) |
Two traps worth designing for deliberately:
The && chain. Your rollback function will be called from an if. Errexit is suspended for the
whole body of a function called that way (lesson 03), so chain the steps:
do_rollback() {
set_tag "$1" \
&& docker compose pull -q \
&& docker compose up -d \
&& wait_healthy
}
Without the chain, a failed pull falls through to wait_healthy, which finds the old container
still healthy and reports a successful rollback that never happened.
State rotation. Re-deploying the tag that is already live must not rotate state:
if [ "$new" != "$cur" ]; then prev="$cur"; fi
Otherwise a routine re-run sets PREVIOUS = CURRENT and destroys your only rollback target.
Predict: deploy v1 twice in a row. What is in .deploy-state afterwards? Write it down, then
look. If your answer was wrong, you just found the bug while it was free.
Stage 5 — the workflow
.github/workflows/deploy.yml, with the job graph from lesson 17:
test ──┬── build-api ──┐
└── build-web ──┴── deploy
testrunsmvn -f greeter-api/pom.xml test- both build jobs compute
sha-${GITHUB_SHA:0:12}and push to GHCR deployneeds both, reads the tag fromneeds.build-api.outputs.tag, SSHes in, passes the env file viaenvs:underumask 077, and runs./deploy.sh "$TAG"concurrency: { group: deploy, cancel-in-progress: false }permissions: { contents: read, packages: write }on the build jobs
For the server, use the in-runner lab VPS from lesson 18.
Stage 6 — the drills
The pipeline working is half the exercise. These three are the other half, and they are what separate this from a tutorial.
Drill 1 — a deploy that must be caught
The toy app has a switch that makes it start, serve traffic, and never report healthy — the shape of a real bad deploy.
Deploy v2 with FAIL_HEALTH=true.
Predict all four before you push:
- What does the health gate print?
- What does the auto-rollback do?
- What colour is the workflow run?
- What does
curl /api/return afterwards, and which version?
The answers you should get: the gate times out and dumps logs; the rollback restores v1; the run is
red; and the site serves v1 — working. Green site, red pipeline. If your run went green,
your script is not exiting non-zero after a rollback, and you have built something that will lie to
you.
Drill 2 — data survives, code does not
Before deploying, hit /api/count a few times. Note the number.
Deploy a new version. Check the count. Roll back. Check again.
The count keeps climbing throughout. A rollback swaps images; it does not touch the volume.
Now the uncomfortable half. Add a column to the visits table in v2, deploy it, then roll back to
v1. The application is v1 again — the column is still there.
Write down, in your own words, why rollback.yml in this repository says migrations are never
reverted. That sentence should now feel obvious rather than arbitrary.
Drill 3 — the interleaved deploy
Two terminals. Run deploy.sh in both at once.
First with flock removed: two sets of compose commands racing, and a .deploy-state file whose
contents depend on timing. Then with flock restored: the second waits.
Stage 7 — the honest check
Leave it a week. Come back and, without opening any notes:
- Add a
promoteworkflow that deploys an existing tag with no rebuild. - Explain why production should promote rather than rebuild from the same commit.
- Explain why
deploy.shpulls before it stops anything. - Explain why a rolled-back deploy still fails the pipeline.
If those come easily, you are done, and the answer to "can you set up CI/CD?" is yes with a repository attached.
What you have not learned here
Say this plainly to yourself, because a course that oversells is worse than one that is short:
- Real TLS and DNS. The lab uses no certificates. See vps-guide.md §5 and §9.
- Real hostile traffic. Nothing scanned your lab box.
- systemd, and the reboot behaviour that comes with it.
- Multi-server anything. scaling.md is honest about where this shape stops working.
- Databases under load, backup verification, and restore drills.
You now have the vocabulary to read all of those. That was the goal.
Where to go next
| Read | For |
|---|---|
| 20 — The big picture | The whole MotorPH pipeline, now that you have built its miniature |
| ci.md | What each real workflow gates, and what is deliberately not gated |
| vps-guide.md | Standing up a real server, end to end |
| troubleshooting.md | The symptom index, for 2am |
Recap
- You built a pipeline: image → registry → SSH → health gate → rollback, every line your own.
- You exercised the failure path, which is the part most people never test until it matters.
- Green means "the new version is live", not "the site is up". A rolled-back deploy is red.
- Rollback restores code, not data — which is why migrations must always go forwards.