02 — One application.yml, zero profiles
Read this first: the backend ships exactly one configuration file, and the running
application never selects a Spring profile — every difference between your laptop, CI, and a
client's production stack is an environment variable with a deliberately chosen default. This
lesson reads application.yml and logback-spring.xml and extracts the rule they both follow:
the default is the off switch. Code is quoted inline so you can read this without the
repository open.
Time: about 25 minutes. Assumes lesson 01.
One file, and the grep that proves it
Look inside backend/src/main/resources/: one application.yml. No application-dev.yml, no
application-prod.yml, no application.properties. Now grep the whole backend for @Profile
— zero matches, main and test source alike. The only profile in the entire tree is
application-tenancy-test.yml under src/test/resources, activated by the tenancy
integration tests and nothing else. The application you deploy has no profiles to be in.
What replaces them is one pattern, repeated on nearly every line:
spring:
datasource:
url: jdbc:postgresql://${POSTGRES_DB_SERVER_ADDRESS:localhost}:${POSTGRES_DB_SERVER_PORT:5434}/${POSTGRES_DB:motorph}
${ENV_VAR:default} — read the environment variable, fall back to the default. The reason to
prefer this over profiles is drift: profile files fork the configuration into parallel
documents that stop agreeing the week after someone edits only one of them. Here the file you
are reading is the configuration in every environment, and the difference between
environments is a short list of variables, not a second document.
Off by default is a doctrine
The defaults are not arbitrary. Each one is chosen so that a machine with nothing else running still boots. The file says so itself, in almost the same sentence, at every integration point:
tracing:
# Off by default, like altcha, billing.provider=stub and motorph.mail.enabled: with no
# collector listening on 4318 the exporter would retry on every span. The monitoring
# overlay turns it on by setting TRACING_ENABLED=true for the backend.
enabled: ${TRACING_ENABLED:false}
mail:
# Off by default, like altcha and billing.provider=stub: a fresh clone, CI and `mvn test`
# boot with nothing listening on 1025 and must not spend five seconds failing to connect on
# every send. docker-compose.yml turns it on, because that stack brings up a mail catcher.
#
# Enabled with no reachable host is a worse state than disabled -- Boot's autoconfiguration
# accepts an empty spring.mail.host and builds a sender that fails at connect time -- so
# SmtpMailService refuses to start in that combination rather than dropping mail silently.
enabled: ${MAIL_ENABLED:false}
billing:
# 'stub' keeps fresh clones and CI booting with no payment secrets configured.
# Set to 'polar' once POLAR_* below are populated. See docs/backend/billing.md.
provider: ${BILLING_PROVIDER:stub}
Notice the second paragraph of the mail comment. The doctrine is not "off is safer" — it is that enabled-but-unreachable is the worst state, because it fails at send time, quietly, in production. So the off state is honest, and turning a subsystem on is an explicit act.
Predict: you clone the repository, set not a single environment variable, start the dev database from lesson 01, and boot the backend. No mail server exists, no trace collector, no ALTCHA key, no payment provider, no Chatwoot token. Write down which of those five absences breaks the boot.
None of them. Mail is disabled, so nothing dials port 1025. Tracing is disabled, so the OTLP exporter never retries against a collector that is not there. ALTCHA is off, so login needs no key and no widget. Billing runs the stub provider. The Chatwoot endpoint answers 204 and the frontend renders no widget. Even the platform operator account follows the doctrine — its password defaults to blank, and "a blank one means no account is created and a warning is logged, which is safer than shipping a known credential." The only hard dependency a fresh clone has is the database itself.
The switchboard
The capability this lesson promised, on one screen — what a fresh clone boots with, and the variable that turns each subsystem on:
Subsystem Fresh clone (no env vars set) The switch
------------------- --------------------------------------------- --------------------------------
Outbound mail off; nothing dials port 1025 MAIL_ENABLED
Tracing off; no exporter, no retry storm TRACING_ENABLED
ALTCHA login proof off; login posts credentials directly ALTCHA_ENABLED (+ ALTCHA_HMAC_KEY)
Billing 'stub' provider; no payment secrets BILLING_PROVIDER (+ POLAR_* values)
Chatwoot widget off; /api/public/chatwoot answers 204 CHATWOOT_WEBSITE_TOKEN
Platform operator not seeded; warning logged instead PLATFORM_ADMIN_PASSWORD
Row-level security inert; app connects as the owning role APP_DB_USER + APP_DB_PASSWORD
JSON log format console format LOG_FORMAT
Swagger UI ON -- the one that defaults on SWAGGER_ENABLED turns it off
One row runs against the grain: Swagger defaults on, because a fresh clone is a development
machine and the API explorer is part of developing. Production stacks set SWAGGER_ENABLED
to false — an on-by-default is acceptable only where the off state costs a developer and the
on state costs nothing without a routed port.
Flyway brings its own credentials
The datasource section carries a comment about who the application connects as:
# Set APP_DB_USER=app_runtime (and APP_DB_PASSWORD to match) to put the application behind
# row-level security. The role and policies are created by V101 and are inert until then,
# because the owning role bypasses RLS -- which is what makes the switch a reversible config
# change rather than a migration. See docs/adr/0013-shared-db-multitenancy.md.
username: ${APP_DB_USER:${POSTGRES_USER:motorph}}
# …
Note the nested fallback: APP_DB_USER if set, else POSTGRES_USER, else the dev default.
Then, a few lines down, Flyway refuses to share:
flyway:
enabled: true
# …
# Migrations run as the owning role, which bypasses row-level security. That is deliberate:
# a migration has to be able to see and backfill every tenant's rows. The application itself
# connects as app_runtime (see spring.datasource above), which does not bypass it.
#
# Giving Flyway its own credentials makes Boot build it a separate DataSource, which then needs
# its own url as well -- it cannot borrow the application's.
url: jdbc:postgresql://${POSTGRES_DB_SERVER_ADDRESS:localhost}:${POSTGRES_DB_SERVER_PORT:5434}/${POSTGRES_DB:motorph}
user: ${POSTGRES_USER:motorph}
# …
Before reading on, answer this one yourself: why would a migration need to see every
tenant's rows, when a request never should? Because migrations change shape, and changing
shape means backfilling — an UPDATE that computes a new column has to touch every tenant's
data, and under row-level security a filtered role would silently update only the rows its
policy shows it. Postgres exempts the table's owning role from RLS, so migrations run as the
owner and requests do not. The mechanical consequence is the part that surprises people: the
moment Flyway has its own user, Boot builds it a separate DataSource, and a separate
DataSource needs its own url — "it cannot borrow the application's". How the request side of
that split is enforced is lesson 13.
The actuator two-file rule
# Only health and prometheus are exposed over HTTP. Neither is routed by the frontend
# nginx (it only proxies /api and /ws), and production client stacks publish no backend
# host ports, so these stay reachable only from inside the compose network -- see
# docker-compose.monitoring.yml for the Prometheus scraper that consumes them.
#
# Adding an endpoint here is only half the change: SecurityConfiguration lists the two
# permitted actuator paths individually (never /actuator/**) so that enabling a new one
# cannot silently expose it. Both files have to agree.
management:
endpoints:
web:
exposure:
include: health,prometheus
Read the second paragraph twice. The security layer never writes /actuator/**; it permits
exactly /actuator/health and /actuator/prometheus, by name. So the two files can only fail
closed: expose a new endpoint here and forget the other file, and the new endpoint answers
401 until you make the change in both places. Had the filter chain used a wildcard, this yml
edit alone would have published the new endpoint — one file quietly deciding what two files
should. The filter chain side of this agreement is lesson 10.
The observability traps
This file has collected several scars, and each one is written down where it happened.
The Boot 3 to 4 property rename. The trace exporter endpoint moved, and the old name fails in the worst possible way:
opentelemetry:
tracing:
export:
otlp:
# Boot 4's property path. The Boot 3 name (management.otlp.tracing.endpoint)
# still appears in the configuration metadata but no longer binds: setting it
# leaves the exporter on its built-in default of localhost:4318, which inside
# a container is the container itself. Spans then vanish with no error --
# this is worth knowing before debugging an empty Tempo for an afternoon.
endpoint: ${OTLP_TRACES_ENDPOINT:http://otel-collector:4318/v1/traces}
Your IDE autocompletes the dead name, nothing errors, and every span posts to the container's own localhost. The comment exists because someone paid the afternoon.
The mail health indicator is off. Boot would otherwise fold SMTP reachability into
/actuator/health:
health:
# The mail indicator opens an SMTP connection on every health read and reports DOWN when the
# relay is unreachable. Email is not on the critical path for serving payroll, and several
# runbooks (docs/deployment/vps-guide.md, docs/security/hardening.md) treat /actuator/health
# as authoritative -- a stopped mail catcher must not make a healthy deployment look broken.
mail:
enabled: false
A health endpoint is a contract about what "down" means. Payroll serves fine with mail down, so mail does not get a vote.
Histograms are rationed. Percentile buckets are what make "what is p95" answerable, and they are enabled for exactly two timers:
# Restricted to the two timers worth the cardinality: every histogram is ~15 extra
# series per tag combination, and http.server.requests already carries uri x method
# x status. Turning this on globally is the classic way to fell a Prometheus.
percentiles-histogram:
http.server.requests: true
motorph.payroll.payslip.generation: true
The same restraint shapes the bucket ceilings: "Payslip generation loops over every active
employee, so minutes are normal here and a 10s ceiling would put every real run in the
overflow bucket." And it shapes the global tags — the env tag is kept to "a small fixed set
of values" because a per-tenant value there would multiply every series in the process.
Never spring.mail.test-connection. "It opens SMTP during startup, which would make a
fresh clone, CI and the e2e suite depend on a reachable mail server just to boot." It is the
off-by-default doctrine restated as a prohibition: a convenience check that adds a boot-time
network dependency undoes the whole design.
Logs pick a format the same way
logback-spring.xml is configuration too, and it follows the same one-file-plus-variable
pattern. Its header explains the trick:
Logback has no if/else without adding Janino, so the switch is a variable in the appender-ref: the appenders are literally named
consoleandjson, and ${LOG_FORMAT} picks one. Anything other than those two values fails loudly at startup, which is the right outcome for a typo in a deployment variable.
<property name="LOG_APPENDER" value="${LOG_FORMAT:-console}"/>
<!-- … -->
<root level="INFO">
<appender-ref ref="${LOG_APPENDER}"/>
</root>
No conditional logic — the variable is the appender name, and a misspelled value points at an appender that does not exist. Both appenders emit the trace and span ids, and the header says why that is the entire point:
Both formats carry the trace and span id. That is the whole point: the three signals only correlate if the same id appears in the metric exemplar, the log line and the span, and this file is where the log half of that happens.
Two details connect this back to the rest of the lesson. First, the ids come from Micrometer
Tracing's MDC keys, written "once management.tracing.enabled is true. With tracing off they
are simply absent and both formats degrade to blank, which is why neither appender guards on
it" — the off-by-default tracing switch degrades logs gracefully instead of breaking them.
Second, the JSON field names are another two-file agreement, just like the actuator rule:
promtail "reads timestamp, level, logger_name, message, trace_id, span_id,
stack_trace by name, and a rename here silently breaks log labelling there."
Where this shows up in MotorPH
application.ymlandlogback-spring.xml— the two files this lesson read; the comments quoted above live there.SecurityConfiguration.java— the other half of the actuator two-file rule.SmtpMailService.java,AltchaProperties.java,PlatformAdminSeeder.java— the fail-fast guards the comments point at: refuse to start misconfigured rather than fail silently later.docker-compose.yml— the stack that flips the switches for containers: mail on (it runs the catcher), JSON logs on.docker-compose.monitoring.ymland the monitoring guide — whereTRACING_ENABLEDbecomes true and the exposed actuator endpoints get their consumer.promtail.yml— the pipeline that reads the JSON log fields by name.- ADR 0013 — shared-database multi-tenancy — the decision behind the two database roles.
- Billing guide, backend docs, deployment docs — turning the stubbed and disabled subsystems on for real.
Recap
- One file, zero profiles. Every environment reads the same
application.yml; the difference between environments is a list of${ENV_VAR:default}values, so configuration cannot fork and drift. - The default is the off switch. Mail, tracing, ALTCHA, billing and Chatwoot all boot disabled or stubbed, because enabled-but-unreachable fails later and quieter than disabled.
- Flyway is the owning role. Migrations must see and backfill every tenant's rows, so they bypass RLS deliberately — and their own credentials force their own DataSource and url.
- Exposure is only half. Actuator endpoints are named in two files that must agree, and
the security side never uses
/actuator/**, so a mismatch fails closed. - Correlation is one id in three places. Both log formats carry trace and span ids, and the JSON field names are a contract with promtail — rename one and labelling breaks silently.