Skip to main content

15 — The payroll spine: runs, payslips, and a ledger that never computes

Read this first: Part 4 is taught at architecture level, which is a split of ownership rather than a gap. Three things own the payroll engine, each completely: ../business-rules.md holds every statutory rule and number, ../api/payroll.md holds the endpoints, the source holds the algorithms. What none of them holds is the shape — which tables one run touches, which may still change afterwards, and which are finished the moment they are written. That is this lesson, and you want it first: numbers and endpoints both read better once you know which rows are permanent.

Time: about 25 minutes. Assumes lesson 14.

Seven tables, one run​

Everything a run touches lives in model/; hold them not by name but by how long they stay true.

EntityTableWritten byAfterwards
Payrollpayrollthe create endpointmutable header — status moves through a small machine
PayrollApprovalpayroll_approvaleach approve/reject decisionappend-only; rows are never edited
Payslippayslipthe generation passfrozen — derived once, then permanent
PayslipDeductionItempayslip_deduction_itemsthe same passfrozen — snapshot lines under a payslip
PayrollTransactionpayroll_transactionsa human, after the facta ledger alongside; never an input
PayrollChangepayroll_changesevery run-level transitionappend-only audit
PayslipHistorypayslip_historyevery payslip-level changeappend-only audit

Read three properties off that table before anything else. Only one row in the set is expected to change after it is written — the Payroll header; everything else is appended or frozen. No entity here declares a @OneToMany — every child holds a foreign key pointing up, no parent holds a collection, which is why a run cannot be deleted by deleting one row and why the cleanup in ../backend/payroll-regeneration.md is a hand-ordered sequence rather than one cascade. And every one extends TenantOwned (lesson 13), so nothing below has to mention tenancy again — the database is already enforcing it.

The header, and a status that is a plain String​

The run entity is flat: a run date, a period start and end, an audit quartet, a pay-frequency snapshot (lesson 16 is entirely about that column) — and this:

@Column(name = "status", nullable = false)
private String status;

Six values pass through that column: Draft, Pending, Approved, Rejected, Processed, Cancelled. Not an enum — while the same entity carries payFrequency, which is an enum, mapped @Enumerated(EnumType.STRING). One class, both choices, deliberately.

The contrast gives you the rule. payFrequency is a closed set the engine branches on, so the compiler should be checking the branches. status is a workflow label that services compare, reports filter, and the UI renders — an enum would buy compiler-checked comparisons at the price of a code change plus a migration per new state and a conversion in every DTO. What pays for the String is where the values get pinned instead: @Pattern on the request DTO (lesson 05) means an unknown status cannot arrive from outside, and every legal transition is checked in one service class. The cost is real — a typo in a comparison is a runtime bug, not a compile error — and it is bounded by keeping the comparisons in one file. That file is lesson 16's subject.

Approval rows are added, never edited​

PayrollApproval is five fields: an id, the run, the approver, a timestamp, a status string. The decide path constructs a new one and saves it; nothing ever loads an approval to modify it. The repository has exactly one query, and its name is the whole design:

List<PayrollApproval> findAllByPayroll_PayrollIdOrderByApprovalDateDesc(Integer payrollId);

A List, newest first. A decision is therefore recorded twice, on purpose: the run header holds the current state, the approval table the sequence of decisions that produced it. Both writes happen in one transaction, and neither is derivable from the other afterwards — the point of writing both.

The payslip is a snapshot, not a view​

Payslip is about fifty columns wide because of one design decision repeated many times: it copies what it could have joined to. The employee's identifying details, the position, the department, the company, the period — all duplicated onto the row rather than reached through the foreign key sitting right there.

A join would give you a view: what this person would be paid if we recomputed against today's data. A payslip answers a different question: what this person was paid. People get married, change department, get promoted, and leave; rate tables get corrected. A row that joins its way to those facts quietly rewrites history every time one of them changes.

The same discipline shows up in how the entity grows — new columns arrive nullable and stay that way, with a comment saying why:

// Employer statutory shares (informational, not deducted); null on payslips
// generated before the feature existed.

That pattern repeats for every field added after the first version. Backfilling would mean recomputing old payslips, the one operation this system does not have. NULL here is not missing data; it accurately says no snapshot was taken, because the thing snapshotted did not exist yet.

The itemized lines snapshot too​

PayslipDeductionItem is the same idea one level down, stated in its own javadoc:

/**
* Itemized custom-deduction line generated for a payslip. label/amount are
* snapshots taken at generation time so historical payslips stay accurate
* even if the deduction type is later renamed or the assignment is deleted.
*/

It keeps a nullable foreign key back to the employee-deduction assignment — useful for tracing — but the label and amount it displays are copies. Rename the deduction type next quarter, delete the assignment when the employee leaves: last year's line still reads as it did on the day it was issued. The link is for provenance; the copy is for truth.

What a rate fix does to a Processed run​

Predict: a statutory contribution rate is corrected — a bracket table is fixed and redeployed. A payroll run from three months ago is already Processed. What happens to that run's payslips? Write your answer down before reading on.

Nothing happens to them. The regeneration runbook says so in bold, then restates it as an absolute a paragraph later:

Existing payslips do not recompute themselves.

There is no in-place recalculation anywhere in this app.

If you predicted "they update" you were reading the frozen columns as a cache. They are not a cache. A payslip is a legal document: it was issued, someone was paid that amount, and the period's BIR filings were produced from it. A row that silently changed to match a later fix would make every filed form disagree with the database it came from. The fix applies only to payslips generated after it.

Which sets up the trap the runbook exists for: cancelling a bad run changes one string on the header and removes no payslips, while every report reads payslip rows by year and month, never by run status.

So if you cancel a bad run and generate a fresh one for the same period, you now have two sets of payslips for that period and every report double-counts.

Correcting history is therefore a deliberate delete-then-regenerate in foreign-key-safe order, not a status change. Freezing is what makes that awkward, and the awkwardness is the feature.

The runbook is behind the source — and that is the normal condition​

Read that runbook and you find two internals it names that no longer exist under those names:

  • PayrollServiceImpl.calculateWithholdingTax — no such method today; withholding tax goes through a WithholdingTaxCalculator collaborator, constructed per run.
  • ...Repository.findAll() for the contribution lookups — those methods now call findBracketAsOf(...) with an earliest-cohort fallback, because the rate tables became versioned cohorts rather than one live table.

Notice what did not change: rates are still read at generation time, payslips are still frozen, the double-count trap is still exactly as described. The reasoning survived; two identifiers did not.

Treat that as the general rule, not an embarrassment — lesson 14 flagged the same about the scheduling doc. When a document and the source disagree, the source wins, and this course will drift the same way. Read prose for reasoning; check names against the tree.

The ledger that never computes​

PayrollTransaction looks like it belongs to the engine and does not. It holds foreign keys to both the run and the payslip, an amount, a date, a description, and a type from a fixed list:

private static final List<String> TRANSACTION_TYPES = List.of(
"Deduction", "Bonus", "Adjustment", "Payment", "Refund", "Other");

Two facts, both checkable by grep, define what it is. Only two files touch its repository — the repository itself and its own CRUD service — and the string PayrollTransaction appears nowhere in PayrollServiceImpl: the generation pass never reads it.

So this is a human-entered record of money movements about a run, posted after the fact and never an input. That separation does real work: if generation read the ledger, regenerating a run would have to decide whether an already-posted adjustment applies again, a question with no correct answer and two wrong ones. Kept apart, the payslip says what the engine computed and the ledger says what was subsequently moved, and neither can quietly rewrite the other.

Two audit trails, one shape​

PayrollChange and PayslipHistory are the same row twice, at two levels: who changed which field, from what, to what, when. Their service interfaces are nearly identical:

void logChange(Integer payrollId, String modifiedBy, String fieldChanged,
String oldValue, String newValue);

Run-level transitions write a PayrollChange; generation writes a PayslipHistory row per payslip. Both stamp their own timestamp in @PrePersist if the caller did not, and neither is ever updated. PayslipHistory carries two extra columns — a department id and a free-text detail string — so the row stands on its own. They are also why the demo-data reset can fail on a foreign-key violation in an environment with real history: audit rows outlive what they describe, which is the point of one.

The settings singleton, in one paragraph​

One more table sits behind all of this without being part of a run. PayrollSettings is:

/**
* Single-row table of configurable payroll computation parameters.
* Field initializers double as the statutory defaults used when no row exists yet.
*/

Roughly seventy fields, exactly one row — per tenant, since it extends TenantOwned like everything else here. It is live configuration, not a snapshot: change it and the next generation pass sees the new value immediately, which is precisely why the run header needs its own frequency column and why lesson 16 opens on it. The values, and which of them the law fixes, live in ../business-rules.md.

Where this shows up in MotorPH​

Recap​

  • One row per run is mutable; the rest are appended or frozen. The Payroll header carries current state; every other table in the spine is an append-only trail or a permanent snapshot.
  • A payslip copies what it could join to, because it answers "what was paid", not "what would be paid". New columns arrive nullable forever: backfilling would mean recomputing, and nothing here recomputes a payslip in place.
  • Cancelling is not deleting. Reports read payslip rows by period, never by the parent run's status, so a cancelled run's payslips still count — fixing history is a deliberate delete-then-regenerate.
  • The ledger is alongside the engine, never inside it. Generation never reads PayrollTransaction, so regeneration never has to guess whether an adjustment was already applied.
  • When a doc and the source disagree, the source wins. The regeneration runbook's reasoning is intact while two of its identifiers are not — and this course will age the same way.

Next: 16 — The run lifecycle: gates, snapshots, and a clock that keeps counting.