29 — Splitting the stack across servers
Read this first: this lesson is about what actually changes when your stack stops living on one machine. The short version, and the thing most tutorials skip: moving a service to another server costs you Docker's DNS, and everything else is a consequence of that. It is optional — the capstone at 19 is the end of the required path — and it is the natural next question once one server is not enough.
Time: about 50 minutes. Assumes lessons 08, 12 and 13.
First: "three servers" means four different things
Almost every confused conversation about scaling is two people meaning different rows of this table.
| Shape | What it is | For | Ready here today? |
|---|---|---|---|
| Split by role | Edge on one box, app on another, database on a third | Isolation, dedicated resources | Yes — no application changes at all |
| Replicas behind a load balancer | Three copies of the same app, traffic shared | Throughput, redundancy | No — see below |
| Isolated stacks per client | A full copy of everything per customer | Strong tenant isolation | Yes, but it is the older model |
| Active/passive failover | One live, others waiting | Uptime | No — needs database failover |
This lesson does the first one, because it is the only one that works today without touching application code, and because it is usually what people actually need. The reason the second is marked "No" is worth reading in full at scaling.md — six pieces of state live in one process's memory, and splitting the app into three copies makes each of them quietly wrong. Not slow. Wrong, with no error.
What splitting by role buys you — and what it does not
It buys: the database gets its own disk and memory instead of competing with a JVM. A memory leak in the app cannot take the database with it. You can move the documentation site off the box that runs payroll — which scaling.md explicitly flags, because today a popular link and a payroll run compete for the same CPU.
It does not buy throughput. One backend before, one backend after. If your app is CPU-bound, this changes nothing at all.
Say that out loud before you spend money on it. "Three servers" sounds like three times the capacity and is not.
Set up the lab
Three containers standing in for three machines, on a network with fixed addresses:
cd learn-devops/lab-cluster
docker compose up -d --build
NAME STATUS
lab-app Up 55 seconds (healthy)
lab-data Up About a minute (healthy)
| "Server" | Address | Runs |
|---|---|---|
data | 10.9.0.30 | PostgreSQL |
app | 10.9.0.20 | The Spring Boot API |
edge | 10.9.0.10 | A proxy — started later |
The honest simplification: these are three containers on one Docker daemon, so Docker's DNS
would still resolve data from app. On genuinely separate machines it would not. Every address
below is written as an IP, and you should resist using the names — the whole point is what happens
when names stop working. Separate kernels, provider private networking and real latency are not
simulated here; they get a tour at the end.
Predict: what breaks first?
You move PostgreSQL to its own machine and change nothing else. The database is running. The network is fine. Credentials are correct.
What is the first error, and what is it about? Write it down.
Most people say "connection refused" or "authentication failed". Run it — this points the app at the database's old name, which is what a lift-and-shift leaves behind:
docker run --rm --network labcluster_dc \
-e APP_SECRET=your-lab-cluster-secret-value \
-e DATABASE_URL=jdbc:postgresql://motorph_payroll_db:5432/greeter \
-e DATABASE_USER=greeter -e DATABASE_PASSWORD=changeme-local-only \
labcluster-greeter-api
Caused by: org.postgresql.util.PSQLException: The connection attempt failed.
Caused by: java.net.UnknownHostException: motorph_payroll_db
UnknownHostException. Not a network error, not an auth error — the name does not exist. You
never got as far as making a connection.
That is the whole lesson in one line. On one host, motorph_payroll_db works because Docker runs an
embedded DNS server for each network and registers every container in it. That DNS server belongs
to one Docker daemon. A second machine has its own daemon, its own DNS, and no knowledge of your
containers. And motorph_payroll_network is a plain bridge network — bridges do not span hosts.
Look at what the real system would have to change. Every upstream in
deploy/Caddyfile is a container name — all seven of them. So is the
backend's database address (POSTGRES_DB_SERVER_ADDRESS: motorph_payroll_db). Split the machines and
every one of those resolves to nothing.
The four things that break, in order
1. Names → addresses
Something has to tell each machine where the others are. In increasing order of effort: hard-coded IPs in env files (fine for three fixed servers), real DNS records on a private zone, or a service discovery system (which is what lesson 30 is partly about).
The lab uses IPs. That is genuinely what a three-server split usually does.
2. The database has to listen on the network
On one host, the compose file says # no ports: and the database is reachable only over the local
bridge. That is not "firewalled" — it is unreachable, which is much stronger.
Move it to another machine and it must accept connections over a real interface:
command: [postgres, -c, listen_addresses=*]
The moment you type that, three things become your problem that were not before. Prove it:
docker run --rm --network labcluster_dc postgres:16-alpine pg_isready -h 10.9.0.30 -p 5432
10.9.0.30:5432 - accepting connections
That is an unrelated container — nothing to do with your app — reaching your database.
docker run --rm --network labcluster_dc -e PGPASSWORD=changeme-local-only postgres:16-alpine \
psql -h 10.9.0.30 -U greeter -d greeter -tAc "select 'connected as '||current_user"
connected as greeter
It authenticated. The password is the only thing standing between the network and your payroll data.
3. The wire is now plaintext
Predict: is that connection encrypted?
docker run --rm --network labcluster_dc -e PGPASSWORD=changeme-local-only postgres:16-alpine \
psql -h 10.9.0.30 -U greeter -d greeter -tAc \
"select coalesce((select 'TLS: '||version from pg_stat_ssl where pid=pg_backend_pid() and ssl),'TLS: NO - plaintext over the wire')"
TLS: NO - plaintext over the wire
Every query, every result, every payroll figure, in the clear. On one host that did not matter — the traffic never left the machine. Now it crosses a wire, and whoever can observe that wire can read it.
This is not a hypothetical for this project. It is the exact mistake ADR-0008 recorded against itself for the CDN-to-origin hop, and it is why ADR-0014 replaced it. The document said plainly that restricting who can connect limits nothing about who can observe.
So a split needs: ssl = on in PostgreSQL with a real certificate, and sslmode=verify-full on the
client. Not require — require encrypts but does not check who you are talking to, which leaves
you open to an impostor.
4. The firewall becomes load-bearing between your own machines
On one host, "the database is not published" was your access control. Now the only thing between port 5432 and the internet is a firewall rule — and from lesson 13 you know that a Docker host's firewall is not simply "UFW".
The rule you want allows the app server and nothing else:
# on the database server — allow only the app server, using documentation addresses
sudo ufw allow from 198.51.100.20 to any port 5432 proto tcp
sudo ufw deny 5432/tcp
And from lesson 13: if PostgreSQL runs in a container, the rule belongs in
DOCKER-USER, interface-qualified, or you will block your own outbound traffic too.
Also tighten pg_hba.conf, which is PostgreSQL's own access list — it can require TLS and restrict
by source range independently of the firewall. Two independent gates, like the Grafana example in
lesson 25.
What else quietly breaks
Not obvious until it bites:
- Backups. The real script runs
docker exec <db-container> pg_dump. There is no local container any more. It becomespg_dump -h <db-host>with credentials and network access — a different security posture for the backup job. - Health gates. A deploy script checking
docker inspecton the app host cannot see the database container. "Is the database up" becomes a network question. - Monitoring.
cadvisorandnode-exportermeasure a host. Three hosts means three of each, not one moved. Prometheus targets that were container names need addresses. - Latency. Every query now crosses a network. Usually sub-millisecond on a provider's private network, but it is no longer zero, and chatty code notices.
How the machines should talk
Three options, honestly compared:
| Option | Good | Bad |
|---|---|---|
| Provider private network | Simplest. Traffic never touches the public internet. Usually free. | You trust the provider's segment. Does not span providers. |
| WireGuard / Tailscale | Portable, encrypted by definition, works across providers | One more thing to run and to debug when it breaks |
| Public internet + TLS | Nothing to set up | Your database has a public address. Avoid. |
Start with the provider's private network and put TLS on the database connection anyway. Defence in depth costs you one config line here.
What MotorPH would actually need
Concretely, if you did this for real:
- Parameterise the seven container-name upstreams in deploy/Caddyfile
- Change
POSTGRES_DB_SERVER_ADDRESSto an address, and add TLS withverify-full - Set
APP_DB_USER=app_runtimeon the new database host — per ADR-0013 that single variable is the gap between "row-level security deployed" and "row-level security enforced" - Rework the pre-deploy
pg_dumpanddeploy/backup.shfor a remote database - Split
deploy/preflight.sh's checks — port 80/443 belongs to the edge host, disk belongs to each host separately - Firewall each host to accept only what the next tier needs
Do the documentation site first. It has no shared state, it already has its own compose project
and its own deploy script, and scaling.md already names its co-tenancy with payroll as a risk. It
is the cheapest possible rehearsal of every problem above.
Recap
- "Three servers" means four different things. Splitting by role works today; three app replicas does not, and the failures are silent.
- Splitting buys isolation, not throughput.
- The first thing that breaks is name resolution —
UnknownHostException, before any network or auth error, because Docker's DNS belongs to one daemon and bridges do not span hosts. - A database on the network needs TLS,
pg_hba.conf, and a firewall — three things that "unpublished port" was doing for free.
Next: 30 — Traefik and Caddy, on what should sit in front of all this.