Skip to content

CloseOps

Five governed agents with a provable no-duplicate guarantee, applied to month-end close

  • multi-agent
  • governance
  • human-in-the-loop
  • A2A protocol
Walkthrough on request

0

duplicate postings

under a forced storm of 36 concurrent retries. Lower is better; the guarantee is that it cannot be anything else.

169

tests, green

5

human gates per close, none decidable by the orchestrator

£0

to run, no API key, no cloud

in short

CloseOps runs a synthetic month-end close with five governed agents. An orchestrator dispatches three specialists over the A2A protocol, humans approve at every risk boundary, and the orchestrator is structurally incapable of approving anything itself. A forced storm of thirty-six concurrent retries produced zero duplicate postings.

the problem

Why this is hard

Month-end close is accounting's biggest bottleneck and the obvious thing to point agents at. It is also the worst possible place to let one act unsupervised. The failure mode is not a bad answer. It is a duplicate journal entry that somebody finds three weeks later, when the books have already been signed.

So CloseOps is not a demo of agents doing a close. It is a demo of the machinery you need around them before you would let them: a checklist with a DAG and risk tiers, human gates at every risk boundary, idempotent dispatch, a variance narrative where every number is cross-checked against a query result, and an append-only close trace that reads as a handoff timeline.

constraints

  • Every ledger, counterparty and transaction is synthetic, generated by ledgerfab from a profile and an integer seed. Nothing is ever posted to a real system.
  • A retry must be a no-op, not a second set of postings. Retries are most likely exactly where something has already succeeded.
  • The orchestrator must not be able to approve its own work, and that must survive a refactor rather than resting on one assertion.
  • Runs on a laptop with no API key, no Docker daemon and no cloud subscription.

where else this pattern applies

The shape is idempotent multi-agent writes behind a separation-of-duties boundary, and the ledger is incidental. Claims adjudication has the same structure. An assessor agent that must not also authorise the payment, and a retry that must not pay twice.

Procurement approval chains need the identical guarantee: a duplicate purchase order raised under a network timeout is the same defect as a duplicate journal entry, and the fix is the same deterministic key rather than a prompt asking the model to be careful.

Anywhere an agent writes to a system of record more than once. Inventory adjustments, payroll runs, provisioning. The orchestrator-cannot-self-approve rule and the forced-retry test transfer unchanged.

architecture

How it runs

A LangGraph orchestrator holds no in-memory state. Postgres tables are the truth, and a close advances by invoking the graph repeatedly. It dispatches three specialists (reconciliation, accruals, variance) over the A2A protocol, each carrying an idempotency key. Gates open to a human and are decided through a signed callback that the Close Board and a Teams card both post to, identically.

  1. 01Checklist loadedplan

    A DAG of seven tasks with risk tiers. The checklist declares requires_gate per task, and it is not believed.

  2. 02Gate derived from policypolicy

    config/risk.yaml defines the tiers and the threshold. The loader refuses to start on a checklist that disagrees with it.

  3. 03Dispatch over A2Atool

    Each specialist is called with an idempotency key. The A2A layer covers in-flight work; the database carries the guarantee.

  4. 04Planted anomaly haltshalt

    Account 1000: 17 of 17 transactions unmatched against a 55% tolerance. Not exceptions for review. The account's data is materially wrong. The task stops.

  5. 05Escalated to a humangate

    The gate opens. The orchestrator can open it and can never decide it, enforced at runtime and by a static import test.

  6. 06Narrative groundedroute

    Every number in the variance draft is matched against a value a query actually returned. An unmatched figure raises, and nothing is emitted.

  7. 07Posting writtencommit

    UNIQUE(postings.idempotency_key) makes the write idempotent. Thirty-six concurrent retries produce one row.

  8. 08Close trace appendedaudit

    Append-only, exportable as JSON or CSV, and legible as a timeline of who handed what to whom.

  • passed
  • awaiting a person
  • stopped

decisions

What was considered, and why it was ruled out

Every one of these had a reasonable alternative. The alternative is named.

ADR-01

Derive the gate requirement from a policy file, never from the checklist

Alternative
Trust the checklist YAML's own requires_gate flag, which is what it is there for.
Why not
The gate-correctness eval would then read requires_gate out of the checklist and assert that requires_gate is true. Grading the file that defines the answer. It would pass forever, including on the day somebody deletes a gate. Separating the policy from the checklist is what makes the eval capable of failing.

ADR-02

Make the orchestrator structurally incapable of approving, in three layers

Alternative
A single runtime assertion, which is what the spec asks for.
Why not
A hard rule enforced by one assert is one refactor away from being decoration. So: one module writes gates.decision; a runtime check refuses machine principals with normalisation; and a static test walks the orchestrator's imports and fails if that module is reachable at all. Its Dockerfile does not even copy the backend.

ADR-03

No LangGraph checkpointer, tables are the truth

Alternative
The Postgres checkpointer the spec explicitly asks for. It was written, wired, and removed.
Why not
CloseOps advances a close by invoking the graph repeatedly, and a checkpointer breaks that in four observed ways. The sharpest: re-invoking one thread_id resumes the finished graph, so the close silently stopped advancing and the retry drill fell from three postings to one with no error anywhere. The property the spec actually wants is proven harder instead. SIGKILL the orchestrator mid-dispatch and a fresh process finishes the close, which works precisely because there is no in-memory state to lose.

ADR-04

Carry idempotency on a database unique constraint, not the A2A task id

Alternative
Derive the A2A task id from the idempotency key and let the protocol dedupe.
Why not
The protocol corrected this twice. A task id cannot be derived from the key. Message.task_id means continue this task, so a first dispatch carrying a derived id gets TaskNotFoundError. And it cannot be reused after completion, because continuing a finished task is refused. So the A2A layer covers in-flight work only, and a design resting on it alone would have had a hole exactly where retries are most likely: after something already succeeded.

ADR-05

An ungrounded number fails the draft rather than annotating it

Alternative
Emit the narrative with a caveat on unverified figures and let the reviewer judge.
Why not
Numbers are extracted, normalised across thousands separators, currency, k/m suffixes, percentages and accounting parentheses, then matched against what the queries returned. Anything unmatched raises and nothing is emitted. The tolerance is deliberately absolute, not relative. 0.5% of a £53,000 opex line is £265, so a relative tolerance would wave through exactly the quietly-wrong figure the check exists to catch.

evaluation

How it was measured

Three evals plus two adversarial drills. The gate-correctness eval ships with a negative control that must fail, because an eval that cannot fail is decoration. The drills print which database they ran against and refuse to describe a SQLite pass as proof.

metricresultnote
Gate-correctness evalgreenand its negative control fails, as it must
Idempotency drill0 duplicates36 concurrent forced retries, on SQLite
Restart drillclose completesreal SIGKILL mid-dispatch, fresh process finishes
Grounding0 unqueried numbersstructural, tested in both directions
Test suite169 passing
A2A contract suitegreenreal protocol, real ASGI apps, no mocked transport

screens

What it looks like running

Captures of the real thing. Projects without a capture show none, nothing here is a mockup.

The CloseOps board: seven tasks in four columns by dependency depth, each with risk tier and gate chips. One reconciliation task is red.
A completed close. Seven tasks, three specialists, five gates, and the planted anomaly in red: account 1000, 17 of 17 transactions unmatched against a 55% tolerance, escalated and resolved by a human. View full size
The gate queue, listing tasks awaiting a human decision with their risk tier and the reason each was gated.
The gate queue. It mirrors the Teams card and posts the same signed callback, so the console has no privileged path of its own. View full size
The close summary: time to close, tasks auto-completed versus gated, anomalies caught, and duplicate postings.
The summary the whole project exists to produce: auto versus gated, anomalies caught, escalations, and duplicate postings at zero. View full size

stack

Built with

  • Python
  • LangGraph
  • a2a-sdk 1.1.2
  • FastAPI
  • Postgres
  • DuckDB
  • React
Walkthrough on request