Skip to main content

30 — Traefik and Caddy

Read this first: two reverse proxies that do the same job in opposite ways. Caddy is configured; Traefik is discovered. That one difference explains every other difference between them, including which one this repository uses and why it changed its mind. You will run the same application behind both.

Time: about 45 minutes. Assumes lesson 25 and lesson 29.

The same job, two philosophies​

Both answer the same question: a request arrives on port 443 — which container gets it?

Caddy answers from a file. You write the hostnames and their upstreams. If something moves, you edit the file and reload.

Traefik answers from labels. Containers advertise their own routes, Traefik watches the Docker socket, and routes appear and disappear as containers do. Nothing central lists them.

Neither is correct in general. They are correct for different questions, and the question is: does your set of routes change on its own?

Caddy: the file​

From the lab:

:80 {
reverse_proxy 10.9.0.20:8080 {
header_up X-Forwarded-For {remote_host}
}
}
cd learn-devops/lab-cluster
docker compose --profile caddy up -d
curl -s localhost:8095/api/
{"version":"dev","greeting":"Hello from the app server"}

The file names its upstream. Nothing is discovered. If the app server moves to a new address, that request fails until a human edits this file — which is either a liability or a feature depending on whether you wanted addresses to change without anyone noticing.

Traefik: the labels​

Now the same application, reached the same way, with nothing anywhere naming it. The route lives on the application container:

labels:
traefik.enable: 'true'
traefik.http.routers.greeter.rule: 'PathPrefix(`/`)'
traefik.http.routers.greeter.entrypoints: 'web'
traefik.http.services.greeter.loadbalancer.server.port: '8080'
docker compose --profile caddy down
docker compose --profile traefik up -d
curl -s localhost:8095/api/
{"version":"dev","greeting":"Hello from the app server"}

Identical response, and the edge configuration contains no mention of the application at all.

Traefik's vocabulary, which is worth learning because it is the part that looks like jargon:

TermIs
EntrypointA port Traefik listens on (web = :80)
RouterA rule deciding which requests match (PathPrefix, Host)
ServiceWhere matching requests go — one or more backends
MiddlewareSomething applied in between: headers, auth, rate limits, redirects
ProviderWhere the configuration comes from — here, the Docker socket

--providers.docker.exposedbydefault=false makes routing opt-in: a container without traefik.enable=true is ignored. Always set it. The default publishes every container you start.

Ask Traefik what it found:

curl -s localhost:8099/api/http/routers
curl -s localhost:8099/api/http/services
greeter@docker PathPrefix(`/`) -> greeter [enabled]
greeter@docker -> ['http://10.9.0.20:8080']

It resolved the container's address by itself. Nobody typed 10.9.0.20. That is the whole pitch — and also, as you are about to see, the whole risk.

Break it on purpose: two containers, one router name​

A configuration file cannot have this failure. Discovery can.

Predict: you start an unrelated container that happens to carry the same traefik.http.routers.greeter labels. What happens to your route? An error? The new one wins? The old one wins?

docker run -d --name lab-impostor --network labcluster_dc \
--label traefik.enable=true \
--label 'traefik.http.routers.greeter.rule=PathPrefix(`/`)' \
--label traefik.http.routers.greeter.entrypoints=web \
--label traefik.http.services.greeter.loadbalancer.server.port=80 \
nginx:alpine
curl -s localhost:8099/api/http/routers
greeter@docker status=enabled err=None

Enabled. No error. Now look at where it sends traffic:

greeter@docker -> ['http://10.9.0.20:8080', 'http://10.9.0.2:80'] err= None

Two backends. Traefik merged them into one load-balanced service. Five requests:

for i in 1 2 3 4 5; do curl -s -o /dev/null -w "%{http_code} " localhost:8095/api/; done
200 404 200 404 200

Half your traffic is going to the wrong container, and every dashboard says enabled, err=None. This is the shape of discovery-based failure: it does not error, it includes things.

Clean up:

docker rm -f lab-impostor

The defence is naming discipline and exposedbydefault=false — and being aware that anything able to start a labelled container on that host can attach itself to your routing. Note what the edge needs to do its job: the Docker socket. From lesson 11, access to that socket is equivalent to root on the host. Caddy needs no such thing.

Break it on purpose: the version that 404s everything​

While building this lab, traefik:v3.3 produced a 404 for every request with the application perfectly healthy. The logs:

ERR Failed to retrieve information of the docker client and server host
error="Error response from daemon: client version 1.24 is too old.
Minimum supported API version is 1.40, please upgrade your client to a newer version"
providerName=docker

The provider could not talk to the daemon, so it discovered nothing, so no router existed, so every request fell through to a 404. The failure was in the discovery layer and the symptom was in the routing layer — which is a general property of this design, and the thing to remember when debugging it.

Worth knowing for this repo specifically: infra/traefik/docker-compose.yml pins traefik:v3.3 and sets DOCKER_API_VERSION=1.44 to address exactly this. That environment variable does not fix it — tested, and v3.3 still negotiates the old version. Only a newer Traefik image did. If that legacy stack is ever revived on a current Docker, this is what it will do first.

Why this repository chose Caddy​

Traefik was here first. ADR-0008 chose it; ADR-0014 superseded that. The reasons are specific, and none of them is "Traefik is a worse proxy":

1. It was deployed HTTP-only. One entrypoint on :80, no certificate resolver, TLS terminated at the CDN in Flexible mode. ADR-0008 named this against itself:

The Cloudflare→origin hop is unencrypted HTTP. Anyone who can observe traffic between Cloudflare and the VPS sees payroll data in the clear.

That is the same plaintext-wire problem as the database in lesson 29, and it is a deployment choice — Traefik can do ACME perfectly well. The stated objection was to the machinery: a wildcard certificate needs a DNS challenge, which needs a provider API token stored on the server.

2. It forced an extra proxy hop, and that broke security. The single label-defined route sent everything to the client's nginx, which proxied /api onward — two appending hops. ADR-0014:

The backend derives the client IP from the last X-Forwarded-For entry and keys rate limiting, account lockout, and audit logging on it. That is sound with exactly one trusted proxy hop; every additional appending hop replaces the real client with a container IP and collapses all users into one rate-limit bucket.

This is the deepest of the four, and it is not really about Traefik. Any extra hop does it — including a load balancer in front of several servers, which is exactly what lesson 29's second topology adds. If you put anything new in that chain, the X-Forwarded-For handling has to be re-derived, or one user's failed logins start locking out everybody.

3. Its main advantage stopped applying. Traefik was chosen so new clients became routable without touching the edge — one stack per client, onboarding by script. Then ADR-0013 made tenants rows in one database:

Onboarding becomes a row rather than a deployment.

Dynamic route discovery buys nothing when the hostname set is five fixed names that change once a year. This is the reason that reverses for multi-server — see below.

4. It held port 80. Both want it. deploy/preflight.sh now treats a running Traefik as a deploy-blocking fault, and the Caddyfile records that Traefik's routes moved into it.

When Traefik is the right answer​

Stated fairly, because the above is a list of reasons it lost here:

  • Services come and go. Preview environments per pull request, many short-lived services, a container fleet that changes daily. Editing a file for each is untenable.
  • Many services, one edge. Dozens of routes, each owned by the team that owns the service, is exactly what labels are for.
  • Per-service middleware. Rate limits, auth, header rewriting, circuit breaking, retries — Traefik has a richer built-in library than Caddy.
  • Multi-host discovery. Providers for Swarm, Kubernetes, Consul and Nomad mean it can discover backends across machines — which is precisely the problem lesson 29 solved with hard-coded IPs.

That last point is the honest case for reconsidering it here. Reason 3 above says discovery stopped mattering on one box. On three boxes it starts mattering again.

What a return would have to answer​

If you proposed going back, these are the questions — and they come from the repo's own history, not from general advice:

  1. A :443 entrypoint with a certificate resolver — and an answer to ADR-0008's objection about storing a provider API token on the server if you go the DNS-challenge route. HTTP-01 on real hostnames avoids it, as Caddy does today.
  2. A separate router for /api and /ws straight to the backend, plus a middleware overwriting X-Forwarded-For from the CDN's header — reproducing header_up X-Forwarded-For {header.CF-Connecting-IP} exactly, or reason 2 comes straight back.
  3. The basic_auth gate on grafana., including stripping the Authorization header before forwarding. No Traefik middleware in this repo does this today.
  4. What replaces the DOMAIN-derivation guard. Caddy derives five hostnames from one variable, and ensure_edge() validates its shape because a malformed value crash-looped the edge and "took production down three times in one night" (lesson 26).

That list is the actual work. "Traefik is more flexible" is not an argument until it is answered.

Choosing, in one table​

IfUse
A fixed set of hostnames that changes rarelyCaddy. The file is the documentation.
Automatic TLS with the least machineryCaddy. It is the default behaviour, not a feature to configure.
Routes that appear and disappear on their ownTraefik
Backends spread across several hosts or an orchestratorTraefik
You do not want the edge to hold the Docker socketCaddy
Rich per-route middleware out of the boxTraefik

For MotorPH today — six hostnames, one box — Caddy is right, and the repo's reasoning holds up. If lesson 29's split happened and servers started coming and going, that changes.

Clean up​

cd learn-devops/lab-cluster
docker compose --profile traefik --profile caddy down -v

Recap​

  • Caddy is configured, Traefik is discovered — every other difference follows from that.
  • Traefik's vocabulary: entrypoint → router → service, with middleware in between and a provider supplying it all. Always set exposedbydefault=false.
  • Discovery fails by including, not by erroring: two containers claiming one router name silently round-robin, and the dashboard says everything is fine.
  • This repo dropped Traefik for four specific reasons, and the extra-proxy-hop one applies to any load balancer you add, not just to Traefik.

That is the end of the course. Back to the index, or on to ci.md and vps-guide.md — which should now read like reference material rather than a wall.