30 — Traefik and Caddy
Read this first: two reverse proxies that do the same job in opposite ways. Caddy is configured; Traefik is discovered. That one difference explains every other difference between them, including which one this repository uses and why it changed its mind. You will run the same application behind both.
Time: about 45 minutes. Assumes lesson 25 and lesson 29.
The same job, two philosophies
Both answer the same question: a request arrives on port 443 — which container gets it?
Caddy answers from a file. You write the hostnames and their upstreams. If something moves, you edit the file and reload.
Traefik answers from labels. Containers advertise their own routes, Traefik watches the Docker socket, and routes appear and disappear as containers do. Nothing central lists them.
Neither is correct in general. They are correct for different questions, and the question is: does your set of routes change on its own?
Caddy: the file
From the lab:
:80 {
reverse_proxy 10.9.0.20:8080 {
header_up X-Forwarded-For {remote_host}
}
}
cd learn-devops/lab-cluster
docker compose --profile caddy up -d
curl -s localhost:8095/api/
{"version":"dev","greeting":"Hello from the app server"}
The file names its upstream. Nothing is discovered. If the app server moves to a new address, that request fails until a human edits this file — which is either a liability or a feature depending on whether you wanted addresses to change without anyone noticing.
Traefik: the labels
Now the same application, reached the same way, with nothing anywhere naming it. The route lives on the application container:
labels:
traefik.enable: 'true'
traefik.http.routers.greeter.rule: 'PathPrefix(`/`)'
traefik.http.routers.greeter.entrypoints: 'web'
traefik.http.services.greeter.loadbalancer.server.port: '8080'
docker compose --profile caddy down
docker compose --profile traefik up -d
curl -s localhost:8095/api/
{"version":"dev","greeting":"Hello from the app server"}
Identical response, and the edge configuration contains no mention of the application at all.
Traefik's vocabulary, which is worth learning because it is the part that looks like jargon:
| Term | Is |
|---|---|
| Entrypoint | A port Traefik listens on (web = :80) |
| Router | A rule deciding which requests match (PathPrefix, Host) |
| Service | Where matching requests go — one or more backends |
| Middleware | Something applied in between: headers, auth, rate limits, redirects |
| Provider | Where the configuration comes from — here, the Docker socket |
--providers.docker.exposedbydefault=false makes routing opt-in: a container without
traefik.enable=true is ignored. Always set it. The default publishes every container you start.
Ask Traefik what it found:
curl -s localhost:8099/api/http/routers
curl -s localhost:8099/api/http/services
greeter@docker PathPrefix(`/`) -> greeter [enabled]
greeter@docker -> ['http://10.9.0.20:8080']
It resolved the container's address by itself. Nobody typed 10.9.0.20. That is the whole
pitch — and also, as you are about to see, the whole risk.
Break it on purpose: two containers, one router name
A configuration file cannot have this failure. Discovery can.
Predict: you start an unrelated container that happens to carry the same
traefik.http.routers.greeter labels. What happens to your route? An error? The new one wins? The
old one wins?
docker run -d --name lab-impostor --network labcluster_dc \
--label traefik.enable=true \
--label 'traefik.http.routers.greeter.rule=PathPrefix(`/`)' \
--label traefik.http.routers.greeter.entrypoints=web \
--label traefik.http.services.greeter.loadbalancer.server.port=80 \
nginx:alpine
curl -s localhost:8099/api/http/routers
greeter@docker status=enabled err=None
Enabled. No error. Now look at where it sends traffic:
greeter@docker -> ['http://10.9.0.20:8080', 'http://10.9.0.2:80'] err= None
Two backends. Traefik merged them into one load-balanced service. Five requests:
for i in 1 2 3 4 5; do curl -s -o /dev/null -w "%{http_code} " localhost:8095/api/; done
200 404 200 404 200
Half your traffic is going to the wrong container, and every dashboard says enabled, err=None.
This is the shape of discovery-based failure: it does not error, it includes things.
Clean up:
docker rm -f lab-impostor
The defence is naming discipline and exposedbydefault=false — and being aware that anything able to
start a labelled container on that host can attach itself to your routing. Note what the edge needs
to do its job: the Docker socket. From lesson 11, access to that
socket is equivalent to root on the host. Caddy needs no such thing.
Break it on purpose: the version that 404s everything
While building this lab, traefik:v3.3 produced a 404 for every request with the application
perfectly healthy. The logs:
ERR Failed to retrieve information of the docker client and server host
error="Error response from daemon: client version 1.24 is too old.
Minimum supported API version is 1.40, please upgrade your client to a newer version"
providerName=docker
The provider could not talk to the daemon, so it discovered nothing, so no router existed, so every request fell through to a 404. The failure was in the discovery layer and the symptom was in the routing layer — which is a general property of this design, and the thing to remember when debugging it.
Worth knowing for this repo specifically: infra/traefik/docker-compose.yml
pins traefik:v3.3 and sets DOCKER_API_VERSION=1.44 to address exactly this. That environment
variable does not fix it — tested, and v3.3 still negotiates the old version. Only a newer Traefik
image did. If that legacy stack is ever revived on a current Docker, this is what it will do first.
Why this repository chose Caddy
Traefik was here first. ADR-0008 chose it; ADR-0014 superseded that. The reasons are specific, and none of them is "Traefik is a worse proxy":
1. It was deployed HTTP-only. One entrypoint on :80, no certificate resolver, TLS terminated at
the CDN in Flexible mode. ADR-0008 named this against itself:
The Cloudflare→origin hop is unencrypted HTTP. Anyone who can observe traffic between Cloudflare and the VPS sees payroll data in the clear.
That is the same plaintext-wire problem as the database in lesson 29, and it is a deployment choice — Traefik can do ACME perfectly well. The stated objection was to the machinery: a wildcard certificate needs a DNS challenge, which needs a provider API token stored on the server.
2. It forced an extra proxy hop, and that broke security. The single label-defined route sent
everything to the client's nginx, which proxied /api onward — two appending hops.
ADR-0014:
The backend derives the client IP from the last
X-Forwarded-Forentry and keys rate limiting, account lockout, and audit logging on it. That is sound with exactly one trusted proxy hop; every additional appending hop replaces the real client with a container IP and collapses all users into one rate-limit bucket.
This is the deepest of the four, and it is not really about Traefik. Any extra hop does it —
including a load balancer in front of several servers, which is exactly what
lesson 29's second topology adds. If you put anything new in that chain, the
X-Forwarded-For handling has to be re-derived, or one user's failed logins start locking out
everybody.
3. Its main advantage stopped applying. Traefik was chosen so new clients became routable without touching the edge — one stack per client, onboarding by script. Then ADR-0013 made tenants rows in one database:
Onboarding becomes a row rather than a deployment.
Dynamic route discovery buys nothing when the hostname set is five fixed names that change once a year. This is the reason that reverses for multi-server — see below.
4. It held port 80. Both want it. deploy/preflight.sh now treats a running Traefik as a
deploy-blocking fault, and the Caddyfile records that Traefik's routes moved into it.
When Traefik is the right answer
Stated fairly, because the above is a list of reasons it lost here:
- Services come and go. Preview environments per pull request, many short-lived services, a container fleet that changes daily. Editing a file for each is untenable.
- Many services, one edge. Dozens of routes, each owned by the team that owns the service, is exactly what labels are for.
- Per-service middleware. Rate limits, auth, header rewriting, circuit breaking, retries — Traefik has a richer built-in library than Caddy.
- Multi-host discovery. Providers for Swarm, Kubernetes, Consul and Nomad mean it can discover backends across machines — which is precisely the problem lesson 29 solved with hard-coded IPs.
That last point is the honest case for reconsidering it here. Reason 3 above says discovery stopped mattering on one box. On three boxes it starts mattering again.
What a return would have to answer
If you proposed going back, these are the questions — and they come from the repo's own history, not from general advice:
- A
:443entrypoint with a certificate resolver — and an answer to ADR-0008's objection about storing a provider API token on the server if you go the DNS-challenge route. HTTP-01 on real hostnames avoids it, as Caddy does today. - A separate router for
/apiand/wsstraight to the backend, plus a middleware overwritingX-Forwarded-Forfrom the CDN's header — reproducingheader_up X-Forwarded-For {header.CF-Connecting-IP}exactly, or reason 2 comes straight back. - The
basic_authgate ongrafana., including stripping theAuthorizationheader before forwarding. No Traefik middleware in this repo does this today. - What replaces the
DOMAIN-derivation guard. Caddy derives five hostnames from one variable, andensure_edge()validates its shape because a malformed value crash-looped the edge and "took production down three times in one night" (lesson 26).
That list is the actual work. "Traefik is more flexible" is not an argument until it is answered.
Choosing, in one table
| If | Use |
|---|---|
| A fixed set of hostnames that changes rarely | Caddy. The file is the documentation. |
| Automatic TLS with the least machinery | Caddy. It is the default behaviour, not a feature to configure. |
| Routes that appear and disappear on their own | Traefik |
| Backends spread across several hosts or an orchestrator | Traefik |
| You do not want the edge to hold the Docker socket | Caddy |
| Rich per-route middleware out of the box | Traefik |
For MotorPH today — six hostnames, one box — Caddy is right, and the repo's reasoning holds up. If lesson 29's split happened and servers started coming and going, that changes.
Clean up
cd learn-devops/lab-cluster
docker compose --profile traefik --profile caddy down -v
Recap
- Caddy is configured, Traefik is discovered — every other difference follows from that.
- Traefik's vocabulary: entrypoint → router → service, with middleware in between and a
provider supplying it all. Always set
exposedbydefault=false. - Discovery fails by including, not by erroring: two containers claiming one router name silently round-robin, and the dashboard says everything is fine.
- This repo dropped Traefik for four specific reasons, and the extra-proxy-hop one applies to any load balancer you add, not just to Traefik.
That is the end of the course. Back to the index, or on to ci.md and vps-guide.md — which should now read like reference material rather than a wall.