12 — Processes, ports, logs, disk
Read this first: this lesson is the 2am toolkit. Four questions resolve most incidents — what is running, what is listening, what do the logs say, and is the disk full — and this teaches you to answer all four in under a minute without looking anything up.
Time: about 40 minutes. Runs on the lab box from lesson 10.
The triage order
When something is broken, ask these in order. The order matters: each one narrows what the next one has to consider.
| # | Question | Command |
|---|---|---|
| 1 | What is running, and is it healthy? | docker ps |
| 2 | Why did it stop or fail? | docker logs --tail 100 <name> |
| 3 | Is something already on the port? | sudo ss -tlnp |
| 4 | Is the disk full? | df -h / |
Almost every incident in this system is one of those four. The troubleshooting runbook opens with the same triage flow for exactly that reason.
1. What is running
docker ps
The columns that matter are STATUS and PORTS. Up 8 seconds (healthy) means a healthcheck is
defined and passing. Up 8 seconds alone means no healthcheck is defined — the container is running
and nothing has verified it works. Restarting means it is crash-looping.
Add -a to include stopped containers, which is where you find Exited (1) from
lesson 05.
Health status directly, which is what a deploy script polls:
docker inspect -f '{{.State.Health.Status}}' <name>
healthy
2. Logs
docker logs --tail 100 <name> # the last 100 lines
docker logs -f <name> # follow
docker compose logs -f greeter-api # one service, via compose
Predict: a container was started with --rm and has exited. Can you read its logs?
No. --rm deletes the container, and the logs go with it. This is why a deploy script must dump
logs before it gives up rather than after — and why
deploy/deploy.sh's health wait ends with docker logs --tail 100 inside
the timeout branch. By the time a human looks, the evidence would otherwise be gone.
3. Ports
sudo ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 0 128 0.0.0.0:22 0.0.0.0:* users:(("sshd",pid=1,fd=3))
LISTEN 0 4096 127.0.0.11:32885 0.0.0.0:*
LISTEN 0 128 [::]:22 [::]:* users:(("sshd",pid=1,fd=4))
| Flag | Means |
|---|---|
-t | TCP only |
-l | Listening sockets only |
-n | Numeric — do not resolve names, which is much faster |
-p | Show the process (needs sudo to see other users') |
Read the Local Address column carefully, because it is the whole security story:
0.0.0.0:22— listening on every interface. Reachable from outside.127.0.0.1:8025— listening on loopback only. Reachable only from the machine itself.
That distinction is how the production monitoring stack is protected. Grafana, Prometheus and
Alertmanager all bind to 127.0.0.1, and you reach them through an SSH tunnel:
ssh -L 3000:localhost:3000 deploy@<VPS_IP>
The tunnel forwards your local port 3000 to the server's loopback. Nothing was ever exposed to the internet.
Break it on purpose: the port conflict
Predict: two containers, both publishing host port 9000. What does the second one say?
docker run -d --name p1 -p 9000:80 nginx:alpine
docker run -d --name p2 -p 9000:80 nginx:alpine
Bind for 0.0.0.0:9000 failed: port is already allocated
Clear and actionable. Find the owner with sudo ss -tlnp | grep 9000, or docker ps and read the
PORTS column.
This is common enough that the course's own compose file makes its host ports configurable, and the
main docker-compose.yml does the same with
${FRONTEND_PORT:-5173}. The fix is never "kill whatever is using it" — it is "find out what it is
first", because on a shared machine that may be something you need.
4. Disk
df -h /
Filesystem Size Used Avail Use% Mounted on
overlay 234G 218G 3.5G 99% /
99% used. That box is minutes away from failing in ways that will look like anything but a disk
problem: docker pull fails mid-layer, Postgres refuses writes, a build dies with a confusing error.
Docker is usually the culprit, and it has a dedicated accounting command:
docker system df
Reclaim space — carefully:
docker image prune # dangling images only. Safe.
docker system prune # + stopped containers, unused networks, build cache
docker system prune -a # + every image not used by a running container
docker system prune -a --volumes deletes volumes. On a production database host that is the
end of the database. The --volumes flag deserves the same reflex as down -v from
lesson 08.
deploy/preflight.sh checks free space before a deploy and warns below 2 GB, precisely because a deploy that runs out of disk halfway through is a bad time.
Processes and memory
ps aux --sort=-%mem | head -5
Sorts by memory, largest first. The container equivalent:
docker stats --no-stream
Relevant to a JVM: an Exited (137) container was killed by signal 9, which usually means the kernel
ran it out of memory. That is what -XX:MaxRAMPercentage=75 in
lesson 06 exists to prevent — without it, a JVM sizes its heap against the
host's memory rather than the container's limit.
Write yourself a status script
Everything from lesson 02 and lesson 03, applied:
#!/usr/bin/env bash
set -euo pipefail
echo "=== containers ==="
docker ps --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'
echo
echo "=== listening ==="
sudo ss -tlnp | grep LISTEN
echo
echo "=== disk ==="
df -h / | tail -1
echo
echo "=== memory ==="
free -h | head -2
Save it as status.sh, chmod +x, and run it as the first thing you do when something is wrong.
Having one command that answers all four questions is worth more than remembering four commands
under pressure.
Where this shows up in MotorPH
- docs/troubleshooting.md opens with a three-command triage flow and then a symptom index — it is this lesson at production scale, and it is the page to read at 2am.
deploy.sh's health wait pollsdocker inspect -f '{{.State.Health.Status}}'every 5 seconds against a deadline, and dumps logs on timeout.- The monitoring overlay binds everything to
127.0.0.1and is reached over an SSH tunnel, for the reason in the ports section above.
Recap
- Four questions: what is running (
docker ps), what did it say (docker logs), what holds the port (ss -tlnp), is the disk full (df -h). 0.0.0.0versus127.0.0.1in the listen address is the whole exposure story.- Logs die with a
--rmcontainer, so capture them before giving up. prunereclaims space;--volumesdeletes data.
Next: 13 — Firewalls, and the rule that killed the containers.