25 — The compose stacks and the edge
Read this first: this chapter reads the production compose files and the Caddy configuration — how four independent stacks share one server without colliding, and how a request from a browser finds the right container. It also covers the DNS and TLS material the lab box cannot teach.
Time: about 40 minutes. Earned by lessons 08 and 09.
Four projects, one network
| Project | File | Contains |
|---|---|---|
motorph-edge | docker-compose.edge.yml | Caddy — the only container publishing ports |
motorph | docker-compose.prod.yml | Production database, backend, frontend |
motorph-stage | docker-compose.stage.yml | Staging database, backend, frontend, Mailpit |
motorph-docs | docker-compose.docs.yml | This documentation site |
They are separate compose projects, which means separate lifecycles: deploying the docs site cannot restart the payroll application. They share one Docker network so the edge can reach all of them by container name — lesson 08's DNS, at production scale:
networks:
motorph_payroll_network:
external: true
external: true means "this network already exists; do not create it". Nobody owns it, and
docs-deploy.sh refuses to run if it is missing rather than creating an empty one — because a
network created by the wrong project would leave the site unreachable in a way that looks fine.
And the negative space matters most: the application services have no ports: at all. Only Caddy
publishes anything. docker ps on that box shows port bindings on one container and nothing on the
rest, which is exactly the arrangement you saw in miniature in lesson 08.
${VAR:?} as an executable contract
image: ghcr.io/jomariabejo/motorph-payroll-backend:${IMAGE_TAG:?managed by deploy/deploy.sh — do not run compose up by hand}
The message is aimed at a human who typed the wrong command, and it arrives at exactly the moment they need it. Compose refuses to start and prints it.
The second-order effect, from lesson 24: preflight.sh greps for
that :? pattern to discover requirements. The compose file is both the configuration and the
specification, and there is no second list to drift.
Which is why docker-compose.docs.yml deliberately contains no ${VAR:?} — a required variable
introduced there would start failing application deploys, because preflight checks the shared edge
config against .env regardless of which environment is being deployed.
Local, staging, production
| Local | Staging | Production | |
|---|---|---|---|
| Images | build: | image: ...:${IMAGE_TAG:?} | same |
| Published ports | 5173, 8081, 5434, 5050, 8025 | Mailpit UI on loopback only | none |
| Network | created by the project | external: true | external: true |
| Swagger | on by default | hardcoded false | hardcoded false |
| CAPTCHA | off by default | hardcoded true | hardcoded true |
| Mailpit container | Mailpit container | a real relay, MAIL_HOST:? required | |
start_period | 90s | 180s | 180s |
| Tracing | present, off | omitted entirely | present, off by default |
Two rows repay attention.
start_period: 180s in production, versus 90 on a laptop, with the comment: a first boot runs
every Flyway migration on VPS-grade disk, and the value that is plenty locally is not enough there.
A health gate tuned on a fast machine will fail its first real deploy.
Tracing is omitted from staging entirely rather than set to off, because there is no collector on that side and an OTLP exporter with nowhere to send retries every span. "Configured off" and "not configured" are different states.
Healthchecks with scars
healthcheck:
test: ["CMD-SHELL", "wget -qS -O /dev/null http://localhost:8080/ 2>&1 | grep -q 'HTTP/' || exit 1"]
The docs container's is more interesting, because it probes 127.0.0.1 and not localhost, with a
comment explaining that Alpine resolves localhost to the IPv6 ::1 first, its nginx listens on
IPv4 only, and the probe therefore fails forever while the site serves perfectly. The container is
marked unhealthy, the deploy gate never passes, and nothing is actually wrong.
You copied that detail into the toy app's healthcheck in lesson 08.
The network alias
networks:
motorph_payroll_network:
aliases: [motorph-payroll-backend]
The container is named motorph_payroll_backend, with underscores. Tomcat rejects underscores in the
Host header per RFC 9110 — so Prometheus scraping the container by its own name gets a bare 400
before Spring ever sees the request. The alias provides a hyphenated name that works.
This is the kind of thing nobody guesses. It is worth reading precisely because it is the shape of problem you will hit: two correct components disagreeing about a standard.
The edge
Caddy owns 80 and 443. Six site blocks, all derived from one {$DOMAIN} variable: the app,
stage., docs., grafana. (behind basic auth), a legacy client host, and a www. redirect.
Within the main block, one matcher decides everything:
@api path /api/* /ws /ws/*
handle @api {
reverse_proxy motorph_payroll_backend:8080 {
header_up X-Forwarded-For {header.CF-Connecting-IP}
}
}
/api/* and the websocket paths go straight to the backend; everything else goes to the frontend's
nginx. /api deliberately bypasses the frontend's nginx, and the reason is the next section.
The header line that looks wrong
header_up X-Forwarded-For {header.CF-Connecting-IP}
header_up overwrites X-Forwarded-For rather than appending to it. Overwriting a forwarding
header is normally exactly wrong.
It is right here because the backend takes the client IP from the last entry, and that is only sound with exactly one trusted hop. Add a second proxy to the chain and the last entry becomes a Docker-internal address — putting every visitor on the site into a single rate-limit bucket, so one user's failed logins lock out everyone.
And it is only trustworthy because ports 80 and 443 are firewalled to the CDN's ranges (lesson 13), so nothing else can reach Caddy to forge the header. On a request that somehow arrives without the CDN header, it degrades to the socket address — degraded, not spoofable.
That is the shape of a real system: a line that looks wrong, is right, and is only right because of something enforced somewhere else entirely. You cannot review it correctly by reading it alone.
The Grafana block
basic_auth {
{$GRAFANA_BASIC_AUTH_USER} {$GRAFANA_BASIC_AUTH_HASH}
}
reverse_proxy motorph_grafana:3000 {
header_up -Authorization
}
Two independent gates: Caddy's basic auth, then Grafana's own login. header_up -Authorization
strips the proxy credential before forwarding — without it every API call arrives at Grafana as
a login attempt for a user it has never heard of, and its auth log fills with what looks like an
attack.
The absence
Prometheus (9090) and Alertmanager (9093) have no site block at all, and the file says so: "If you ever find yourself adding a site block for them, don't." They are reached through an SSH tunnel (lesson 12).
Documenting a deliberate absence is rare and valuable. Without that line, a future engineer would reasonably assume it was an oversight.
What the lab box cannot teach you
Honest tour of the four things from lesson 10:
DNS. Records for @, www, stage, docs, grafana point at the server, proxied through the
CDN. VPS_HOST in the deploy secrets must be the IP, because the name now resolves to the CDN
and a CDN does not proxy SSH.
TLS. Caddy obtains certificates from Let's Encrypt automatically — which requires that the ACME server can reach your server from the internet. The certificate volume is worth backing up thoughtfully: losing it forces re-issuance, and issuance is rate-limited.
The edge also sets dns: [1.1.1.1, 8.8.8.8], because the host's systemd-resolved stub at
127.0.0.53 is unreachable from inside a container namespace, and ACME lookups would fail forever.
The CDN. TLS mode is Full (strict), and "Always Use HTTPS" stays off until the first certificate issues — turning it on early creates a redirect loop before there is anything to redirect to.
Reboots. restart: unless-stopped brings containers back; netfilter-persistent brings the
firewall back. There are no systemd units in this repository — scheduling is two crontab lines,
documented rather than committed.
Recap
- Four compose projects, one external network, one publishing container. Separate projects mean separate lifecycles.
${VAR:?}is a contract thatpreflight.shreads, so requirements never drift into a second list.- Production differs from local in ways that matter —
start_period: 180s, no published ports, and tracing omitted rather than disabled. header_upoverwrites, which is correct only because the firewall guarantees a single trusted hop.- Prometheus and Alertmanager are deliberately unrouted, and the file says so.
Next: 26 — when it breaks.