Skip to main content

22 — The gates decoded

Read this first: deploy.yml is only allowed to run because other workflows said yes. This chapter reads the seven that gate and guard — what each actually blocks, what it merely reports, and what is deliberately not gated at all. That last category is the most interesting.

Time: about 35 minutes. Earned by lessons 15 and 17.

The seven​

WorkflowRuns onBlocks a merge?
tests.ymlPull requests, and called by deploy.ymlYes
playwright.ymlPush and pull request (smoke); manual (full)Yes, smoke only
codeql.ymlPush, pull request, weekly cronNo — reports
gitleaks.ymlPush and pull requestYes
deploy.ymlPush to main / stagingIt is the deploy
docs.ymlChanges under docs/, docs-site/, and some backend pathsYes, for docs
promote.yml / rollback.ymlManual onlyOperator tools

tests.yml — and the trigger that is missing​

Two jobs: test-backend (JUnit, 30-minute timeout) and test-frontend (typecheck, lint, Vitest, 20 minutes). It declares on: pull_request and on: workflow_call — and no push:.

That omission is deliberate. With a push: trigger, merging to main would run the whole suite twice: once for the merge commit, once because deploy.yml calls it. Reusable workflows only save you duplication if you also remove the duplicate trigger.

The retry loop worth stealing​

- name: Pre-pull Testcontainers images
run: |
for attempt in 1 2 3; do
if docker pull --quiet postgres:16-alpine; then
exit 0
fi
echo "::warning::docker pull failed (attempt $attempt/3); retrying in $((attempt * 15))s"
sleep $((attempt * 15))
done
echo "::error::could not pull postgres:16-alpine after 3 attempts — the tenancy tests cannot run"
exit 1

The backend's multi-tenancy tests start a real Postgres through Testcontainers. When a registry rate limit blocked that pull, the failure surfaced as ExceptionInInitializerError and NoClassDefFoundError — Java errors that look exactly like a code bug and are nothing of the sort. Engineers went looking in the wrong place.

The fix does not make the pull more reliable so much as make it legible: pull explicitly, retry with $((attempt * 15)) backoff, and if it still fails say so in a sentence that names the real cause. ::warning:: and ::error:: are workflow commands (lesson 15) that surface in the run summary.

The transferable idea: when a dependency fails in a way that gets misattributed, add a step whose only job is to fail first, clearly.

Documented debt​

- run: npm run lint
continue-on-error: true

Lint runs and cannot fail the build. That is normally a smell — here it is accompanied by a comment stating the exact debt (58 problems, 27 errors, 31 warnings) and the condition under which the flag comes off.

The distinction matters: continue-on-error with a number attached and an exit criterion is a decision. continue-on-error alone is someone who got tired.

playwright.yml — two jobs, one event switch​

smoke:
if: github.event_name != 'workflow_dispatch'
full:
if: github.event_name == 'workflow_dispatch'

Mutually exclusive by trigger: pushes get the fast smoke test, the full 180-minute suite runs only when someone asks. Both start with:

run: echo "JWT_SECRET=$(openssl rand -hex 64)" >> .env

That exists because compose defaults JWT_SECRET to a placeholder, the backend's JwtProperties.validateSecret rejects it at startup, and the only symptom compose reports is "container is unhealthy". The comment says exactly that. It is the same class of problem as the Testcontainers pull: a precise cause producing a vague message.

full also overrides the config's CI worker count to 2, because the runner has two cores.

codeql.yml — a gate that cannot gate​

- uses: github/codeql-action/analyze@v3
with:
category: "/language:java-payroll-backend"
upload: never
output: sarif-results

upload: never — results are uploaded as artifacts instead of to GitHub's code-scanning UI, because that feature needs Advanced Security, which a personal private repository does not have. The analysis still runs; only the destination changes.

Its permissions: block explicitly includes actions: read, with a comment: the analyze action reads its own workflow run through the API and 403s without it, because an explicit permissions block zeroes every unlisted scope (lesson 16).

The category: is preserved from a since-removed four-way matrix so that alert history carries over — a small thing that saves re-triaging every finding.

docs.yml — a gate that reads the output​

The documentation workflow does something unusual and worth copying: rather than trusting its own denylist of pages that must not be published, it greps the built HTML for things that must never appear — credential values, complete bcrypt hashes, firewall rules naming routable addresses.

Verifying the output rather than the intent is the whole idea. A future rename or a new page quoting a password would slip past a path list; it cannot slip past a scan of what was actually generated.

One line in it is a small masterclass:

done < <(grep -oE '^\| Password \| `[^`]+`' ../docs/reference/demo-users.md \
| sed -E 's/.*`([^`]+)`.*/\1/' | sort -u)

That is process substitution (lesson 04) and it is load-bearing. The loop body sets fail=1; piping into while instead would run the loop in a subshell, the assignment would evaporate, and the gate would pass while finding problems. A security check that silently always passes is worse than no check.

Note too that the file sets set -uo pipefail with no -e, so it can accumulate every failure and exit once at the end — the same choice preflight.sh makes, for the same reason.

What is deliberately not gated​

The honest list, which the repository states plainly rather than leaving you to discover:

  • Frontend lint — continue-on-error, with the debt written down.
  • The full end-to-end suite — too slow for every push; smoke gates instead.
  • CodeQL findings — reported, not blocking.
  • Performance and load — no load testing has been done, and scaling.md says so rather than implying otherwise.

A pipeline that claims to gate everything is either lying or unusable. Naming what you do not check is part of the design.

Recap​

  • Reusable workflows only help if you also remove the duplicate trigger — hence no push: in tests.yml.
  • When a dependency failure gets misattributed, add a step that fails first and clearly.
  • continue-on-error with a number and an exit criterion is a decision; without one it is neglect.
  • Verify the output, not the intent — and beware the subshell that swallows your failure flag.

Next: 23 — deploy.sh decoded.