# Gate 05 — The audit log can reconstruct the run

**Check:** `checks/audit_check.py` · **Stops the build:** yes

## What this gate asserts

Every automated step wrote a log entry, and every entry carries enough to let
somebody who was not present work out what happened.

Required fields on every entry:

| Field | Why it is required |
|---|---|
| `run_id` | Ties every entry from one run together. |
| `timestamp` | UTC, ISO 8601. Local times in a log are a source of argument later. |
| `step` | Which step ran. |
| `inputs` | Each input by identifier **and content hash**. A filename alone does not establish which version was read. |
| `outputs` | What was written. |
| `outcome` | `ok` or `failed`, with a reason on failure. |

Required additionally where a model was involved:

| Field | Why |
|---|---|
| `provider`, `model` | Model behaviour changes between versions; "we used AI" does not reconstruct anything. |
| `tokens_in`, `tokens_out` | The basis of the cost line, and the thing that reconciles against the provider's bill. |

The check also asserts the log is **append-only in practice**: entries are in
non-decreasing timestamp order, and no `run_id` appears with two different
timestamps for the same step.

## Why it exists

A log recording that a step succeeded, and nothing else, satisfies a developer
and fails an auditor. The question at a client is never "did it run" — it is
"which version of the data did it read, when, and can you show me that the
figure in this report came out of that run rather than the one before it".

Content hashes are what make that answerable. Filenames get overwritten; a hash
does not. This is also what allows the byte-identical-rerun property in the
report pipeline to mean something: same inputs by hash, same output.

## What a client can challenge

- *"Show me the run that produced this report."* By `run_id`, with its inputs
  and their hashes.
- *"Prove the data has not changed since."* Re-hash the source and compare.
- *"What did the model do here?"* If a model was in the loop at all, the entry
  names the provider, the model and the token counts. If a step has no model
  fields, no model was involved in it — which for the render step is the
  point of the design.
- *"Has this log been edited?"* The ordering checks make casual editing
  visible. They do not make it impossible; if a client needs tamper evidence
  rather than tamper visibility, that is a different control and should be
  scoped as one rather than claimed here.

## Known limits

This gate checks the shape of the log, not the truth of it. A step that logged
inputs it did not read would pass. The defence against that is that the same
code writes the log and does the work, in one place, which is a design choice
rather than something this gate verifies.

## How it fails

```
gate-05  FAIL  logs/run.jsonl:7   missing fields: inputs.hash, tokens_out
gate-05  FAIL  logs/run.jsonl:12  timestamp goes backwards (2026-08-27T09:14:02Z after 2026-08-27T09:15:40Z)
gate-05  FAIL  logs/run.jsonl:19  run_id 4f2c step "render" logged twice with different timestamps
```

## Remediation

1. **Missing fields.** Fix the step that writes the entry, not the entry. An
   entry corrected by hand is worse than a missing one, because it looks
   complete.
2. **Ordering.** Usually a step writing local time instead of UTC. Fix the
   step; leave the historical entries and note the correction.
3. **Duplicate step in one run.** Either a retry that was not labelled as one —
   add the attempt number — or two runs sharing an id, which is a bug in id
   generation.
4. Re-run `python3 qa-gates/checks/run_gates.py`.

## Configuration

| Setting | Default | Where to change it |
|---|---|---|
| Log path | `logs/run.jsonl` | `--log` |
| Model fields required | when `provider` is present | `--require-model-fields` |
