Skip to main content

12 — Processes, ports, logs, disk

Read this first: this lesson is the 2am toolkit. Four questions resolve most incidents — what is running, what is listening, what do the logs say, and is the disk full — and this teaches you to answer all four in under a minute without looking anything up.

Time: about 40 minutes. Runs on the lab box from lesson 10.

The triage order​

When something is broken, ask these in order. The order matters: each one narrows what the next one has to consider.

#QuestionCommand
1What is running, and is it healthy?docker ps
2Why did it stop or fail?docker logs --tail 100 <name>
3Is something already on the port?sudo ss -tlnp
4Is the disk full?df -h /

Almost every incident in this system is one of those four. The troubleshooting runbook opens with the same triage flow for exactly that reason.

1. What is running​

docker ps

The columns that matter are STATUS and PORTS. Up 8 seconds (healthy) means a healthcheck is defined and passing. Up 8 seconds alone means no healthcheck is defined — the container is running and nothing has verified it works. Restarting means it is crash-looping.

Add -a to include stopped containers, which is where you find Exited (1) from lesson 05.

Health status directly, which is what a deploy script polls:

docker inspect -f '{{.State.Health.Status}}' <name>
healthy

2. Logs​

docker logs --tail 100 <name> # the last 100 lines
docker logs -f <name> # follow
docker compose logs -f greeter-api # one service, via compose

Predict: a container was started with --rm and has exited. Can you read its logs?

No. --rm deletes the container, and the logs go with it. This is why a deploy script must dump logs before it gives up rather than after — and why deploy/deploy.sh's health wait ends with docker logs --tail 100 inside the timeout branch. By the time a human looks, the evidence would otherwise be gone.

3. Ports​

sudo ss -tlnp
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
LISTEN 0 128 0.0.0.0:22 0.0.0.0:* users:(("sshd",pid=1,fd=3))
LISTEN 0 4096 127.0.0.11:32885 0.0.0.0:*
LISTEN 0 128 [::]:22 [::]:* users:(("sshd",pid=1,fd=4))
FlagMeans
-tTCP only
-lListening sockets only
-nNumeric — do not resolve names, which is much faster
-pShow the process (needs sudo to see other users')

Read the Local Address column carefully, because it is the whole security story:

  • 0.0.0.0:22 — listening on every interface. Reachable from outside.
  • 127.0.0.1:8025 — listening on loopback only. Reachable only from the machine itself.

That distinction is how the production monitoring stack is protected. Grafana, Prometheus and Alertmanager all bind to 127.0.0.1, and you reach them through an SSH tunnel:

ssh -L 3000:localhost:3000 deploy@<VPS_IP>

The tunnel forwards your local port 3000 to the server's loopback. Nothing was ever exposed to the internet.

Break it on purpose: the port conflict​

Predict: two containers, both publishing host port 9000. What does the second one say?

docker run -d --name p1 -p 9000:80 nginx:alpine
docker run -d --name p2 -p 9000:80 nginx:alpine
Bind for 0.0.0.0:9000 failed: port is already allocated

Clear and actionable. Find the owner with sudo ss -tlnp | grep 9000, or docker ps and read the PORTS column.

This is common enough that the course's own compose file makes its host ports configurable, and the main docker-compose.yml does the same with ${FRONTEND_PORT:-5173}. The fix is never "kill whatever is using it" — it is "find out what it is first", because on a shared machine that may be something you need.

4. Disk​

df -h /
Filesystem Size Used Avail Use% Mounted on
overlay 234G 218G 3.5G 99% /

99% used. That box is minutes away from failing in ways that will look like anything but a disk problem: docker pull fails mid-layer, Postgres refuses writes, a build dies with a confusing error.

Docker is usually the culprit, and it has a dedicated accounting command:

docker system df

Reclaim space — carefully:

docker image prune # dangling images only. Safe.
docker system prune # + stopped containers, unused networks, build cache
docker system prune -a # + every image not used by a running container

docker system prune -a --volumes deletes volumes. On a production database host that is the end of the database. The --volumes flag deserves the same reflex as down -v from lesson 08.

deploy/preflight.sh checks free space before a deploy and warns below 2 GB, precisely because a deploy that runs out of disk halfway through is a bad time.

Processes and memory​

ps aux --sort=-%mem | head -5

Sorts by memory, largest first. The container equivalent:

docker stats --no-stream

Relevant to a JVM: an Exited (137) container was killed by signal 9, which usually means the kernel ran it out of memory. That is what -XX:MaxRAMPercentage=75 in lesson 06 exists to prevent — without it, a JVM sizes its heap against the host's memory rather than the container's limit.

Write yourself a status script​

Everything from lesson 02 and lesson 03, applied:

#!/usr/bin/env bash
set -euo pipefail

echo "=== containers ==="
docker ps --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'

echo
echo "=== listening ==="
sudo ss -tlnp | grep LISTEN

echo
echo "=== disk ==="
df -h / | tail -1

echo
echo "=== memory ==="
free -h | head -2

Save it as status.sh, chmod +x, and run it as the first thing you do when something is wrong. Having one command that answers all four questions is worth more than remembering four commands under pressure.

Where this shows up in MotorPH​

  • docs/troubleshooting.md opens with a three-command triage flow and then a symptom index — it is this lesson at production scale, and it is the page to read at 2am.
  • deploy.sh's health wait polls docker inspect -f '{{.State.Health.Status}}' every 5 seconds against a deadline, and dumps logs on timeout.
  • The monitoring overlay binds everything to 127.0.0.1 and is reached over an SSH tunnel, for the reason in the ports section above.

Recap​

  • Four questions: what is running (docker ps), what did it say (docker logs), what holds the port (ss -tlnp), is the disk full (df -h).
  • 0.0.0.0 versus 127.0.0.1 in the listen address is the whole exposure story.
  • Logs die with a --rm container, so capture them before giving up.
  • prune reclaims space; --volumes deletes data.

Next: 13 — Firewalls, and the rule that killed the containers.