CloseOps
Five governed agents with a provable no-duplicate guarantee, applied to month-end close
- multi-agent
- governance
- human-in-the-loop
- A2A protocol
0
duplicate postings
under a forced storm of 36 concurrent retries. Lower is better; the guarantee is that it cannot be anything else.
169
tests, green
5
human gates per close, none decidable by the orchestrator
£0
to run, no API key, no cloud
in short
CloseOps runs a synthetic month-end close with five governed agents. An orchestrator dispatches three specialists over the A2A protocol, humans approve at every risk boundary, and the orchestrator is structurally incapable of approving anything itself. A forced storm of thirty-six concurrent retries produced zero duplicate postings.
the problem
Why this is hard
Month-end close is accounting's biggest bottleneck and the obvious thing to point agents at. It is also the worst possible place to let one act unsupervised. The failure mode is not a bad answer. It is a duplicate journal entry that somebody finds three weeks later, when the books have already been signed.
So CloseOps is not a demo of agents doing a close. It is a demo of the machinery you need around them before you would let them: a checklist with a DAG and risk tiers, human gates at every risk boundary, idempotent dispatch, a variance narrative where every number is cross-checked against a query result, and an append-only close trace that reads as a handoff timeline.
constraints
- Every ledger, counterparty and transaction is synthetic, generated by ledgerfab from a profile and an integer seed. Nothing is ever posted to a real system.
- A retry must be a no-op, not a second set of postings. Retries are most likely exactly where something has already succeeded.
- The orchestrator must not be able to approve its own work, and that must survive a refactor rather than resting on one assertion.
- Runs on a laptop with no API key, no Docker daemon and no cloud subscription.
where else this pattern applies
The shape is idempotent multi-agent writes behind a separation-of-duties boundary, and the ledger is incidental. Claims adjudication has the same structure. An assessor agent that must not also authorise the payment, and a retry that must not pay twice.
Procurement approval chains need the identical guarantee: a duplicate purchase order raised under a network timeout is the same defect as a duplicate journal entry, and the fix is the same deterministic key rather than a prompt asking the model to be careful.
Anywhere an agent writes to a system of record more than once. Inventory adjustments, payroll runs, provisioning. The orchestrator-cannot-self-approve rule and the forced-retry test transfer unchanged.
architecture
How it runs
A LangGraph orchestrator holds no in-memory state. Postgres tables are the truth, and a close advances by invoking the graph repeatedly. It dispatches three specialists (reconciliation, accruals, variance) over the A2A protocol, each carrying an idempotency key. Gates open to a human and are decided through a signed callback that the Close Board and a Teams card both post to, identically.
01Checklist loadedplan
A DAG of seven tasks with risk tiers. The checklist declares requires_gate per task, and it is not believed.
02Gate derived from policypolicy
config/risk.yaml defines the tiers and the threshold. The loader refuses to start on a checklist that disagrees with it.
03Dispatch over A2Atool
Each specialist is called with an idempotency key. The A2A layer covers in-flight work; the database carries the guarantee.
04Planted anomaly haltshalt
Account 1000: 17 of 17 transactions unmatched against a 55% tolerance. Not exceptions for review. The account's data is materially wrong. The task stops.
05Escalated to a humangate
The gate opens. The orchestrator can open it and can never decide it, enforced at runtime and by a static import test.
06Narrative groundedroute
Every number in the variance draft is matched against a value a query actually returned. An unmatched figure raises, and nothing is emitted.
07Posting writtencommit
UNIQUE(postings.idempotency_key) makes the write idempotent. Thirty-six concurrent retries produce one row.
08Close trace appendedaudit
Append-only, exportable as JSON or CSV, and legible as a timeline of who handed what to whom.
- passed
- awaiting a person
- stopped
decisions
What was considered, and why it was ruled out
Every one of these had a reasonable alternative. The alternative is named.
ADR-01
Derive the gate requirement from a policy file, never from the checklist
- Alternative
- Trust the checklist YAML's own requires_gate flag, which is what it is there for.
- Why not
- The gate-correctness eval would then read requires_gate out of the checklist and assert that requires_gate is true. Grading the file that defines the answer. It would pass forever, including on the day somebody deletes a gate. Separating the policy from the checklist is what makes the eval capable of failing.
ADR-02
Make the orchestrator structurally incapable of approving, in three layers
- Alternative
- A single runtime assertion, which is what the spec asks for.
- Why not
- A hard rule enforced by one assert is one refactor away from being decoration. So: one module writes gates.decision; a runtime check refuses machine principals with normalisation; and a static test walks the orchestrator's imports and fails if that module is reachable at all. Its Dockerfile does not even copy the backend.
ADR-03
No LangGraph checkpointer, tables are the truth
- Alternative
- The Postgres checkpointer the spec explicitly asks for. It was written, wired, and removed.
- Why not
- CloseOps advances a close by invoking the graph repeatedly, and a checkpointer breaks that in four observed ways. The sharpest: re-invoking one thread_id resumes the finished graph, so the close silently stopped advancing and the retry drill fell from three postings to one with no error anywhere. The property the spec actually wants is proven harder instead. SIGKILL the orchestrator mid-dispatch and a fresh process finishes the close, which works precisely because there is no in-memory state to lose.
ADR-04
Carry idempotency on a database unique constraint, not the A2A task id
- Alternative
- Derive the A2A task id from the idempotency key and let the protocol dedupe.
- Why not
- The protocol corrected this twice. A task id cannot be derived from the key. Message.task_id means continue this task, so a first dispatch carrying a derived id gets TaskNotFoundError. And it cannot be reused after completion, because continuing a finished task is refused. So the A2A layer covers in-flight work only, and a design resting on it alone would have had a hole exactly where retries are most likely: after something already succeeded.
ADR-05
An ungrounded number fails the draft rather than annotating it
- Alternative
- Emit the narrative with a caveat on unverified figures and let the reviewer judge.
- Why not
- Numbers are extracted, normalised across thousands separators, currency, k/m suffixes, percentages and accounting parentheses, then matched against what the queries returned. Anything unmatched raises and nothing is emitted. The tolerance is deliberately absolute, not relative. 0.5% of a £53,000 opex line is £265, so a relative tolerance would wave through exactly the quietly-wrong figure the check exists to catch.
evaluation
How it was measured
Three evals plus two adversarial drills. The gate-correctness eval ships with a negative control that must fail, because an eval that cannot fail is decoration. The drills print which database they ran against and refuse to describe a SQLite pass as proof.
| metric | result | note |
|---|---|---|
| Gate-correctness eval | green | and its negative control fails, as it must |
| Idempotency drill | 0 duplicates | 36 concurrent forced retries, on SQLite |
| Restart drill | close completes | real SIGKILL mid-dispatch, fresh process finishes |
| Grounding | 0 unqueried numbers | structural, tested in both directions |
| Test suite | 169 passing | |
| A2A contract suite | green | real protocol, real ASGI apps, no mocked transport |
screens
What it looks like running
Captures of the real thing. Projects without a capture show none, nothing here is a mockup.



stack
Built with
- Python
- LangGraph
- a2a-sdk 1.1.2
- FastAPI
- Postgres
- DuckDB
- React
Aneeq Khatri