Most absence platforms can produce a PDF of a leave case when an auditor asks. Far fewer can hand over a structured export that a compliance analyst can independently verify — one where timestamps are machine-readable, payroll amounts tie back to a specific journal entry, and nobody can quietly edit a record after the fact without leaving a trace.
That gap is where month-end reconciliations fall apart and where litigation support turns into a three-week scramble. If you've already set up your audit-ready leave documentation checklist and folder structure, this is the technical layer underneath it — the actual schemas, checksums, and sampling rules that make the folder structure defensible rather than just tidy.
This isn't a governance philosophy piece. It's a spec. Developers can implement the schemas directly, and compliance teams can use the verification commands to confirm an export is actually sound before certifying anything.
## 1. Canonical JSON schema for leave case metadata
Everything downstream — manifests, checksums, reconciliation — depends on a stable leave case record. The field names and types below should be treated as a contract. Once you ship them, renaming a field breaks every tie-out that references it.
{ "caseid": "a3f8c2e1-4b9d-4f21-8e77-9c0d1a2b3c4d", "employeeid": "b7e9d4a2-1c33-4a88-9f12-7d6e5c4b3a21", "leavetype": "FMLACONTINUOUS", "status": "APPROVED", "requestedat": "2026-03-02T14:21:07Z", "decisionat": "2026-03-04T09:12:44Z", "leavestartdate": "2026-03-10", "leaveenddate": "2026-04-21", "totalhoursapproved": 240.0, "intermittent": false, "jurisdiction": "US-CA", "createdby": "c9f1a3b5-6d2e-4710-8c99-2a4b6d8e0f13", "createdat": "2026-03-02T14:21:07Z", "lastmodifiedat": "2026-03-04T09:12:44Z", "schema_version": "1.2.0" }
Field rules that matter in practice:
| Field | Type | Format / Allowed values |
|---|---|---|
caseid, employeeid, created_by | string | UUID v4 (lowercase, hyphenated) |
leave_type | string (enum) | FMLACONTINUOUS, FMLAINTERMITTENT, STD, LTD, PTO, SICK, PARENTAL, BEREAVEMENT, JURY, UNPAID_PERSONAL |
status | string (enum) | DRAFT, SUBMITTED, UNDER_REVIEW, APPROVED, DENIED, CANCELLED, CLOSED |
requestedat, decisionat, createdat, lastmodified_at | string | ISO 8601 with UTC Z suffix, seconds precision |
leavestartdate, leaveenddate | string | ISO 8601 date only (YYYY-MM-DD) |
totalhoursapproved | number | decimal, never a string, never null for approved cases |
intermittent | boolean | true/false, never "yes"/"no" |
jurisdiction | string | ISO 3166-2 code |
The most common mistake is mixing types across exports — totalhoursapproved arriving as "240" in one file and 240.0 in another because two services serialized it differently.
Minimal implementation checklist:
-
[ ] Reject any record missing
caseid,leavetype,status, orcreated_at -
[ ] Enforce enum membership server-side, not just in the UI
-
[ ] Store
schema_versionon every record so old exports stay interpretable -
[ ] Never reuse a
case_id, even for a re-opened case — link instead
Your reconciliation scripts will silently cast the string, produce a result that looks right, and then fail on the one record where someone typed "240 hrs". Lock the type in the schema and validate on export, not after.
## 2. Export package manifest
An export is not just a folder of files. It's a folder of files plus a manifest that says exactly what should be there and what each file's fingerprint is. Without the manifest, there's no reliable way to tell whether a file was dropped, swapped, or truncated in transit.
Stop managing absences manually.
Absencely simplifies leave requests, approvals, and absence monitoring for your entire workforce.
- Automated leave tracking
- Manager approval workflows
- Compliance & reporting tools
No credit card required
{ "manifestversion": "1.0", "exportid": "d4b2f6a8-3c71-4e90-91aa-5f7c9b1d2e34", "createdby": "c9f1a3b5-6d2e-4710-8c99-2a4b6d8e0f13", "createdat": "2026-05-01T06:00:12Z", "exportreason": "MONTHENDRECONCILIATION", "files": [ { "filepath": "cases/2026-04/leavecases.jsonl", "recordids": ["a3f8c2e1-...", "e5d7c9b1-..."], "recordcount": 412, "checksumsha256": "9f2c8a1b4d6e0f3a7c5b9d2e1f4a6c8b0d3e5f7a9c1b3d5e7f9a1c3b5d7e9f1a", "filesizebytes": 184233, "redactionstatus": "REDACTED", "linkedpayrollids": ["pr202604run07", "pr202604run08"] } ], "manifestchecksumsha256": "computedoverfilesarrayexcludingthis_field" }
manifestchecksumsha256 should be computed over the canonicalized files array (sorted keys, no whitespace) so the manifest can vouch for itself. redactionstatus is an enum: NONE, REDACTED, PENDINGREVIEW. Never ship a file marked PENDING_REVIEW to an external party — that status exists precisely to block it in CI.
## 3. Cryptographic integrity controls and verification
The point of checksums here isn't security theater. It's answering one question an opposing counsel or external auditor will eventually ask: can you prove this export is the same one you generated on the first of the month?
-
Record snapshot hash — hash each case record at export time, stored inside the record itself as
snapshot_sha256. This freezes the record's state. -
File checksum — SHA-256 over each output file, recorded in the manifest.
-
Manifest checksum — SHA-256 over the manifest's file list, making the manifest tamper-evident.
-
Export signature — sign the manifest checksum with an organizational key (GPG or a cloud KMS key), so authenticity is provable, not just integrity.
Verify a single file against its manifest entry sha256sum cases/2026-04/leavecases.jsonl # Verify the manifest signature (GPG example) gpg --verify manifest.json.sig manifest.json # Verify every file listed in the manifest in one pass jq -r '.files[] | "\(.checksumsha256) \(.file_path)"' manifest.json \ | sha256sum -c -
That last command is the one to put in your runbook. It reads the manifest, reconstructs the expected checksum line for each file, and lets sha256sum -c confirm all of them at once. If it prints anything other than OK for every file, the export is not certifiable — full stop.
A simple verification workflow:
One practical note worth flagging: generate the snapshot hash before redaction, store it, then generate a separate post-redaction file checksum. Otherwise you can't prove the original existed in a specific state while still shipping only the redacted version.
## 4. Immutable audit log spec
Append-only is the whole game. A leave audit log that can be edited is worth roughly nothing when you need to show who approved a case and when. The format below is JSONL (one event per line), which streams cleanly and never requires rewriting earlier lines.
{"eventid":"f1a2b3c4-...","timestamp":"2026-03-04T09:12:44Z","actor":"c9f1a3b5-...","role":"LEAVEADMIN","ip":"10.44.2.19","sessionid":"sess8f3c...","action":"CASESTATUSCHANGE","before":{"status":"UNDERREVIEW"},"after":{"status":"APPROVED"},"caseid":"a3f8c2e1-..."}
Captured fields, non-negotiable:
-
event_id— UUID v4, unique per event -
timestamp— ISO 8601 UTC -
actor— user UUID (never a display name) -
role— the role in effect at action time, not the user's current role -
ipandsession_id— for forensic correlation -
action— enum from a fixed vocabulary -
before/after— JSON diffs of the changed fields only -
case_id— the object acted on
Store role as it was at the moment of the action. People get promoted, change teams, or lose access. If your log joins to the current role table at query time, your audit trail quietly rewrites history every time someone's permissions change — which defeats the entire purpose.
Retention and export rules:
-
Write events to append-only storage (object storage with versioning and object-lock, or an immutable log table with no
UPDATE/DELETEgrants) -
Retain for the longest applicable statute — typically 4–7 years for leave records depending on jurisdiction
-
Export as NDJSON so each line is independently parseable; a truncated file still yields valid complete events up to the break
Checksum the exported log file and record it in the manifest like any other artifact
## 5. Payroll linkage fields for automated tie-outs
This is where reconciliation actually lives or dies. A leave case says "240 hours approved." Payroll says "we paid X." Nobody can tie those together automatically unless the linkage fields exist and match a known vocabulary.
{ "caseid": "a3f8c2e1-...", "payrollrunid": "pr202604run07", "payrolljournalid": "jrnl202604000912", "paycode": "FMLAPAID", "payperiodstart": "2026-04-01", "payperiodend": "2026-04-15", "hours": 72.0, "amount": 2184.00, "currency": "USD" }
Rules that keep tie-outs honest:
-
amountis numeric with exactly two decimal places for USD-style currencies — never a formatted string with a$ -
currencyis an ISO 4217 code, always present even in single-currency environments (it saves you later when you aren't) -
pay_codemaps to a controlled list shared with payroll; a free-text pay code guarantees manual reconciliation forever -
One leave case can link to many payroll rows — model it as a one-to-many, keyed on
case_id
Maintain a small, version-controlled crosswalk between paycode values and leavetype values. When an automated tie-out finds a paycode that doesn't map to the case's leavetype, that's a flag, not an error to swallow. A typical example: a SICK case paying under a PTO pay code because someone applied the wrong code during entry. Caught at month-end, that's a five-minute fix. Caught during an audit eight months later, it's a finding.
## 6. Month-end reconciliation CSV schema with worked example
The reconciliation file is what a controller signs. It needs to be flat, human-readable, and still machine-parseable. CSV wins here precisely because a finance reviewer can open it in Excel and a script can validate it in CI.
Columns, in order:
employeeid, openingbalance, accruals, adjustments, leavetakenhours, payrollhours, payrollamounts, transfers, closingbalance, reconcilerid, signofftimestamp, varianceflag, variance_reason
employeeid,openingbalance,accruals,adjustments,leavetakenhours,payrollhours,payrollamounts,transfers,closingbalance,reconcilerid,signofftimestamp,varianceflag,variancereason b7e9d4a2-...,96.00,13.33,0.00,72.00,72.00,2184.00,0.00,37.33,c9f1a3b5-...,2026-05-02T16:40:11Z,false, d2c4e6f8-...,120.00,13.33,-8.00,40.00,48.00,1456.00,0.00,85.33,c9f1a3b5-...,2026-05-02T16:41:55Z,true,PAYROLLHOURSEXCEEDLEAVE_TAKEN
The second row is the one you want your script to catch: leavetakenhours of 40 but payrollhours of 48. That eight-hour gap means payroll paid leave that the case doesn't account for — exactly the kind of leakage that adds up across a few hundred employees. varianceflag is boolean, and variance_reason must be non-empty whenever the flag is true. Enforce that pairing in validation; an unexplained variance flag is the first thing auditors circle.
Make
variance_reasonvalidation part of the CSV schema check so flagged rows can't be signed without an explanation.
The balance identity your validator should assert on every row:
openingbalance + accruals + adjustments + transfers − leavetakenhours = closingbalance
Rows failing that identity shouldn't be signable until someone resolves them.
## 7. Redaction implementation details
Redaction on leave documents fails in two predictable ways: the visible text gets covered but the underlying data survives, or the OCR layer misses PHI buried in a scanned doctor's note. Both produce an export that looks redacted and isn't.
A sound redaction pipeline does three things:
-
OCR-based PHI detection across scanned attachments — flagging diagnoses, provider names, and dates that fall outside the leave window
-
Embedded metadata removal — stripping EXIF, document author fields, revision history, and comments that routinely leak names and internal notes
-
True content removal, not visual overlay — the redacted bytes are gone, replaced, and the file is re-serialized
Every redaction produces a manifest entry:
{ "documentid": "doc7f3a...", "whoredacted": "c9f1a3b5-...", "reasoncode": "PHIDIAGNOSIS", "timestamp": "2026-04-28T11:03:00Z", "redactionmethod": "OCRDETECTPLUSCONTENTREMOVAL", "retainedoriginallocation": "vault://originals/2026/doc_7f3a" }
The unredacted original never leaves the vault. retainedoriginallocation points to access-controlled storage with its own audit log, so if legal needs the original under a hold, there's a clean chain to it. The export carries only the redacted version plus this manifest entry.
This ties directly into the folder conventions in the audit-ready leave documentation checklist — the redaction manifest is what makes those folders defensible rather than just organized.
## 8. Sampling framework and cadence
Nobody reconciles every record line-by-line indefinitely. You reconcile fully at month-end, then sample between cycles to catch drift. The sampling design is what separates "we check our data" from something an auditor will actually accept.
Define these before you sample anything:
-
Population — all leave cases with payroll activity in the period
-
Sampling unit — one leave case (not one payroll row, or you'll double-count multi-row cases)
-
Selection method — random for unbiased baseline, systematic (every nth) for even coverage, stratified when leave types carry different risk
Recommended cadence:
| Activity | Frequency | Scope |
|---|---|---|
| Full reconciliation | Monthly | Entire population with payroll activity |
| Exception sampling | Weekly | All variance-flagged cases + random sample of clean ones |
| Audit sample | Quarterly | Stratified by leave type, deeper document-level review |
For exception sampling, review 100% of flagged cases plus roughly 5–10% of clean cases to confirm the clean ones are genuinely clean. For the quarterly audit sample, size to the population — a few dozen cases out of a few hundred usually gives enough confidence to spot systemic issues.
Failure thresholds and escalation:
-
If more than roughly 2% of sampled clean cases turn out to have variances, treat the month's reconciliation as suspect and widen the sample
-
Any single variance above a materiality threshold (set it in dollars — anything over a few hundred is a reasonable starting point) escalates to the controller immediately, not at the next monthly cycle
-
Three consecutive weeks of rising exception rates triggers a root-cause review of the upstream pay-code mapping
Failure thresholds should be set in policy and wired into your escalation playbooks so operational teams act consistently when samples show drift.
## 9. RBAC, segregation of duties, and legal-hold workflow
The person who approves a leave case should not be the same person who signs off its reconciliation. That's the core of segregation of duties, and it's surprisingly common for a small HR team to collapse both roles into one login.
Minimum role split:
-
Leave admin — creates and approves cases; cannot sign reconciliations
-
Reconciler — signs month-end files; cannot alter case data
-
Auditor — read-only across everything, including the audit log
-
Legal hold officer — can set and clear freeze flags; cannot edit records
When a hold is placed, set a legalhold freeze flag on the affected cases. The flag does three things: it blocks deletion and modification regardless of role, it suspends normal retention expiry, and it attaches chain-of-custody metadata to any forensic export (exportedby, exportedat, exporthash, custodytransferto). That metadata is what lets you testify, if it comes to that, that the export handed to counsel is bit-for-bit what the system held.
A forensic export under hold reuses the same manifest and checksum machinery from sections 2 and 3 — plus the signature. The difference is the custody fields and the fact that nothing in the freeze set can change while the hold is active.
## A small CI test dataset
Ship a tiny fixture so your validation pipeline has something to run against every build:
-
3 leave cases
one
APPROVEDwith clean payroll linkage, oneDENIEDwith no payroll rows, oneAPPROVEDwith a deliberate hours mismatch -
1 manifest with correct checksums, plus a second copy with one corrupted checksum to confirm your verifier actually fails
-
1 reconciliation CSV with the balance identity holding on two rows and broken on a third
-
1 audit log NDJSON with a before/after status change and a redaction event
Your CI should assert: schema validation passes on the good records and fails on the broken ones, the checksum verifier rejects the corrupted manifest, and the reconciliation validator flags the mismatched row with a non-empty reason. If all four pass, your export machinery is doing what it claims.
Where this actually pays off
The value of all this isn't passing one audit. It's that month-end stops being a manual hunt. When the schemas are locked, the checksums verify, and the pay-code crosswalk is maintained honestly, a reconciler spends their time on the handful of genuine variances instead of re-keying data just to figure out whether the numbers line up at all.
Eight-hour mismatches surface in the weekly exception sample instead of eight months later in a deposition. Build the schema contract first, wire the checksums and audit log in early, and treat the reconciliation CSV as the signed artifact it is. Everything else — sampling cadence, legal holds, redaction — hangs off those foundations. Get them right and audit-readiness stops being a quarterly fire drill and becomes a byproduct of how the data already flows.
Ready to optimize your workforce absence management?
Join 2,000+ HR teams using Absencely to reduce administrative burden, improve compliance, and boost employee satisfaction.