Skip to main content

25 — The compose stacks and the edge

Read this first: this chapter reads the production compose files and the Caddy configuration — how four independent stacks share one server without colliding, and how a request from a browser finds the right container. It also covers the DNS and TLS material the lab box cannot teach.

Time: about 40 minutes. Earned by lessons 08 and 09.

Four projects, one network​

ProjectFileContains
motorph-edgedocker-compose.edge.ymlCaddy — the only container publishing ports
motorphdocker-compose.prod.ymlProduction database, backend, frontend
motorph-stagedocker-compose.stage.ymlStaging database, backend, frontend, Mailpit
motorph-docsdocker-compose.docs.ymlThis documentation site

They are separate compose projects, which means separate lifecycles: deploying the docs site cannot restart the payroll application. They share one Docker network so the edge can reach all of them by container name — lesson 08's DNS, at production scale:

networks:
motorph_payroll_network:
external: true

external: true means "this network already exists; do not create it". Nobody owns it, and docs-deploy.sh refuses to run if it is missing rather than creating an empty one — because a network created by the wrong project would leave the site unreachable in a way that looks fine.

And the negative space matters most: the application services have no ports: at all. Only Caddy publishes anything. docker ps on that box shows port bindings on one container and nothing on the rest, which is exactly the arrangement you saw in miniature in lesson 08.

${VAR:?} as an executable contract​

image: ghcr.io/jomariabejo/motorph-payroll-backend:${IMAGE_TAG:?managed by deploy/deploy.sh — do not run compose up by hand}

The message is aimed at a human who typed the wrong command, and it arrives at exactly the moment they need it. Compose refuses to start and prints it.

The second-order effect, from lesson 24: preflight.sh greps for that :? pattern to discover requirements. The compose file is both the configuration and the specification, and there is no second list to drift.

Which is why docker-compose.docs.yml deliberately contains no ${VAR:?} — a required variable introduced there would start failing application deploys, because preflight checks the shared edge config against .env regardless of which environment is being deployed.

Local, staging, production​

LocalStagingProduction
Imagesbuild:image: ...:${IMAGE_TAG:?}same
Published ports5173, 8081, 5434, 5050, 8025Mailpit UI on loopback onlynone
Networkcreated by the projectexternal: trueexternal: true
Swaggeron by defaulthardcoded falsehardcoded false
CAPTCHAoff by defaulthardcoded truehardcoded true
MailMailpit containerMailpit containera real relay, MAIL_HOST:? required
start_period90s180s180s
Tracingpresent, offomitted entirelypresent, off by default

Two rows repay attention.

start_period: 180s in production, versus 90 on a laptop, with the comment: a first boot runs every Flyway migration on VPS-grade disk, and the value that is plenty locally is not enough there. A health gate tuned on a fast machine will fail its first real deploy.

Tracing is omitted from staging entirely rather than set to off, because there is no collector on that side and an OTLP exporter with nowhere to send retries every span. "Configured off" and "not configured" are different states.

Healthchecks with scars​

healthcheck:
test: ["CMD-SHELL", "wget -qS -O /dev/null http://localhost:8080/ 2>&1 | grep -q 'HTTP/' || exit 1"]

The docs container's is more interesting, because it probes 127.0.0.1 and not localhost, with a comment explaining that Alpine resolves localhost to the IPv6 ::1 first, its nginx listens on IPv4 only, and the probe therefore fails forever while the site serves perfectly. The container is marked unhealthy, the deploy gate never passes, and nothing is actually wrong.

You copied that detail into the toy app's healthcheck in lesson 08.

The network alias​

networks:
motorph_payroll_network:
aliases: [motorph-payroll-backend]

The container is named motorph_payroll_backend, with underscores. Tomcat rejects underscores in the Host header per RFC 9110 — so Prometheus scraping the container by its own name gets a bare 400 before Spring ever sees the request. The alias provides a hyphenated name that works.

This is the kind of thing nobody guesses. It is worth reading precisely because it is the shape of problem you will hit: two correct components disagreeing about a standard.

The edge​

Caddy owns 80 and 443. Six site blocks, all derived from one {$DOMAIN} variable: the app, stage., docs., grafana. (behind basic auth), a legacy client host, and a www. redirect.

Within the main block, one matcher decides everything:

@api path /api/* /ws /ws/*
handle @api {
reverse_proxy motorph_payroll_backend:8080 {
header_up X-Forwarded-For {header.CF-Connecting-IP}
}
}

/api/* and the websocket paths go straight to the backend; everything else goes to the frontend's nginx. /api deliberately bypasses the frontend's nginx, and the reason is the next section.

The header line that looks wrong​

header_up X-Forwarded-For {header.CF-Connecting-IP}

header_up overwrites X-Forwarded-For rather than appending to it. Overwriting a forwarding header is normally exactly wrong.

It is right here because the backend takes the client IP from the last entry, and that is only sound with exactly one trusted hop. Add a second proxy to the chain and the last entry becomes a Docker-internal address — putting every visitor on the site into a single rate-limit bucket, so one user's failed logins lock out everyone.

And it is only trustworthy because ports 80 and 443 are firewalled to the CDN's ranges (lesson 13), so nothing else can reach Caddy to forge the header. On a request that somehow arrives without the CDN header, it degrades to the socket address — degraded, not spoofable.

That is the shape of a real system: a line that looks wrong, is right, and is only right because of something enforced somewhere else entirely. You cannot review it correctly by reading it alone.

The Grafana block​

basic_auth {
{$GRAFANA_BASIC_AUTH_USER} {$GRAFANA_BASIC_AUTH_HASH}
}

reverse_proxy motorph_grafana:3000 {
header_up -Authorization
}

Two independent gates: Caddy's basic auth, then Grafana's own login. header_up -Authorization strips the proxy credential before forwarding — without it every API call arrives at Grafana as a login attempt for a user it has never heard of, and its auth log fills with what looks like an attack.

The absence​

Prometheus (9090) and Alertmanager (9093) have no site block at all, and the file says so: "If you ever find yourself adding a site block for them, don't." They are reached through an SSH tunnel (lesson 12).

Documenting a deliberate absence is rare and valuable. Without that line, a future engineer would reasonably assume it was an oversight.

What the lab box cannot teach you​

Honest tour of the four things from lesson 10:

DNS. Records for @, www, stage, docs, grafana point at the server, proxied through the CDN. VPS_HOST in the deploy secrets must be the IP, because the name now resolves to the CDN and a CDN does not proxy SSH.

TLS. Caddy obtains certificates from Let's Encrypt automatically — which requires that the ACME server can reach your server from the internet. The certificate volume is worth backing up thoughtfully: losing it forces re-issuance, and issuance is rate-limited.

The edge also sets dns: [1.1.1.1, 8.8.8.8], because the host's systemd-resolved stub at 127.0.0.53 is unreachable from inside a container namespace, and ACME lookups would fail forever.

The CDN. TLS mode is Full (strict), and "Always Use HTTPS" stays off until the first certificate issues — turning it on early creates a redirect loop before there is anything to redirect to.

Reboots. restart: unless-stopped brings containers back; netfilter-persistent brings the firewall back. There are no systemd units in this repository — scheduling is two crontab lines, documented rather than committed.

Recap​

  • Four compose projects, one external network, one publishing container. Separate projects mean separate lifecycles.
  • ${VAR:?} is a contract that preflight.sh reads, so requirements never drift into a second list.
  • Production differs from local in ways that matter — start_period: 180s, no published ports, and tracing omitted rather than disabled.
  • header_up overwrites, which is correct only because the firewall guarantees a single trusted hop.
  • Prometheus and Alertmanager are deliberately unrouted, and the file says so.

Next: 26 — when it breaks.