Skip to content

LedgerLab

An open MCP server that hands any agent a realistically messy world, applied to bank and ledger reconciliation

  • MCP
  • open source
  • synthetic data
  • developer tooling
A nightmare-profile world of 105 transactions, a live GPT-5.6 Luna agent reconciling it over MCP, then the answer key revealed: 52.1% coverage and 100% precision. It matched about half of what it could, and every match it made was correct. That is a live model run; the table below is the deterministic reference heuristic. The presenter is AI-generated; the project and the numbers are real.

399

tests, green

covering determinism, both transports, and 20 concurrent sessions.

7

MCP tools, fenced by a test that fails CI if an eighth appears

3

messiness profiles, byte-identical from profile and seed

£0

to run, it calls no LLMs at all

in short

An open MCP server that hands any client a realistically messy synthetic bank and ledger: mixed date formats, missing references, aliased payee names, duplicate payments. Seven tools, no ground truth reachable through any of them, and a live viewer that streams every tool call as your agent works.

the problem

Why this is hard

Finance AI cannot be demoed on real data, and the synthetic datasets that exist are toy-clean: every reference present, every name spelled the same way, every payment matching an invoice to the penny. Agents that look brilliant on that data fall over on the first real bank feed.

LedgerLab generates books that are wrong in the ways real books are actually wrong, hands them to any MCP client in about a minute, and lets you watch the agent work. The messiness is not invented. It is the set of things that went wrong in books an ex-accountant used to close.

constraints

  • Same profile plus same seed must produce a byte-identical world on any machine, every time. The dataset hashes are committed and asserted in CI.
  • No tool may reach the answer key. This is checked three ways: by AST, by published schema, and by inspecting live responses.
  • Seven tools, not eight. Tool bloat measurably degrades tool selection, so the count is enforced by a test.
  • Results are paginated and capped at 50 rows. Rejected, not silently truncated. Because a payload is what a tool call costs an agent's context window.

where else this pattern applies

The idea is a deterministic, seeded, deliberately messy fixture world served over MCP, with the answer key held where no tool can reach it. Nothing about that is financial. The same server shape would give a support agent a messy ticket history, or a clinical agent a set of notes with inconsistent names and missing referrals.

Withholding the joins that would make the task a string comparison is the reusable design move. Any agent benchmark that hands over a clean foreign key is measuring lookup rather than reasoning, whichever domain it is dressed in.

Errors written as recoverable instructions. What went wrong, what is true, what to call next. Are a general MCP authoring practice, and the live tool-call feed showing payload size per call is how any team makes context cost visible.

architecture

How it runs

The MCP server is mounted inside the viewer's FastAPI process, which is what makes the event bus genuinely in-process: a tool call reaches the browser without a broker, a queue or a second service. Sessions are one SQLite world each, seeded on creation by a stdlib-only generator that also emits an answer key no tool can read.

  1. 01Seed a worldplan

    A profile and an integer seed produce a bank, a ledger, counterparties, invoices and notes, plus an answer key held out of reach.

  2. 02Any MCP client connectstool

    stdio for Claude Desktop, Streamable HTTP for LangGraph and everything else. Both tested end to end against real servers.

  3. 03The agent has to workroute

    A bank transaction carries a payee name, not a foreign key. Resolving it is a tool call, which is what stops the alias messiness being cosmetic.

  4. 04Errors are part of the productpolicy

    Each one says what went wrong, what is actually true, and what to call next. So a wrong guess becomes a recoverable step rather than a dead end.

  5. 05Some payments have no explanationhalt

    Flagging one unknown is the correct answer. This is what stops an agent scoring well by confidently matching everything to something.

  6. 06Every call streams to the viewercommit

    Tool badge, argument preview, latency, payload size. Pausing buffers rather than disconnecting, so nothing is lost.

  7. 07Grade what is in the sessionaudit

    The acceptance gate grades what actually ended up in the database, not what the client claims it did.

  • passed
  • awaiting a person
  • stopped

decisions

What was considered, and why it was ruled out

Every one of these had a reasonable alternative. The alternative is named.

ADR-01

Withhold the joins that would make reconciliation a string comparison

Alternative
Put counterparty_id on bank transactions and a transaction id on GL entries, as a clean schema would.
Why not
A bank feed gives you a name, not a foreign key, and the GL holds what was booked while the bank side is unposted. Adding either join would make propose_match solvable by string equality and the acceptance test meaningless. The omissions are the product; they are documented as design decisions rather than left to look like bugs.

ADR-02

Cap results at 50 rows and reject rather than truncate

Alternative
Return everything, or truncate silently with a note in the response.
Why not
Silent truncation gives an agent a wrong answer that looks complete, which is the worst available failure. Rejecting forces pagination to be handled. Payload size is shown on every row of the live feed for the same reason. It is what a tool call costs an agent's context window, and it is the reason the cap exists at all.

ADR-03

Make the acceptance gate's headroom problem public rather than tightening the floor

Alternative
Raise the thresholds to something that looks discriminating, or quietly leave the 70% floor unexplained.
Why not
A roughly 200-line deterministic bookkeeping heuristic scores 100% coverage and 100% precision on all three profiles. So the gate proves an agent can complete this job reproducibly with zero API misuse. A genuine regression test. But it does not distinguish a good agent from a great one. Presenting a 70% floor as if it were tight would be dishonest, and making the world genuinely hard is the first item on the roadmap instead.

ADR-04

Mount the MCP server inside the viewer process

Alternative
Run the MCP server and the viewer API as two services with a message broker between them.
Why not
The event bus is the feature. Watching tool calls arrive live is most of the value. In-process means a call reaches the browser with no broker, no queue and no second service to run, which keeps the whole thing to one uvicorn process and a make target. Fan-out to two viewers, per-session isolation and backfill for a late joiner are all tested.

evaluation

How it was measured

The acceptance gate runs a reference LangGraph agent against a real server over Streamable HTTP and grades what ended up in the session rather than what the client reported. It passes on all three profiles, and the honest reading of that is in the limits below.

metricresultnote
clean profile100% coverage · 100% precision0 tool errors
realistic profile100% / 100%cause accuracy 23/23
nightmare profile100% / 100%cause accuracy 66/66
Determinismhashes identicalacross interpreters, against a patched clock, against committed hashes
Concurrency20 sessionsreads and writes, no cross-session bleed
Tests398 passed, 1 skippedthe skip is the live-LLM test

try it

Connect your own client

Not hosted yet. These are the exact steps against a local instance, and the same ones against the hosted URL when it lands.

LedgerLab is an MCP server, so the useful demo is not a screenshot. It is pointing your own client at it and watching your agent work a messy set of books. Seven tools, no key, no account, read-only against a seeded synthetic bank.

  1. 01Claude Desktop, or any client over Streamable HTTP

    {
      "mcpServers": {
        "ledgerlab": { "type": "http", "url": "http://localhost:8000/mcp" }
      }
    }

    Restart the client, then ask it to reconcile the 2025-Q1 bank statement.

  2. 02LangGraph, or anything using the MCP adapters

    client = MultiServerMCPClient({
        "ledgerlab": {"transport": "streamable_http", "url": "http://localhost:8000/mcp"},
    }, handle_tool_errors=False)
    tools = await client.get_tools()   # all 7

    handle_tool_errors=False matters: the errors are part of the product, and your agent should be able to tell one from an answer.

  3. 03Just list the tools

    npx -y @modelcontextprotocol/inspector --cli http://localhost:8000/mcp \
      --transport http --method tools/list

screens

What it looks like running

Captures of the real thing. Projects without a capture show none, nothing here is a mockup.

The session detail screen: 64 transactions, 61 matches proposed, 24 exceptions flagged, and a table of the agent's proposed matches with its stated reason for each.
One reference-agent run against a seeded world. Note the line under the progress bar: it is progress, not accuracy. The tools never tell an agent whether a match is correct. Grading happens against a ground-truth table no tool can reach, and the answer key is viewer-only. View full size
The live tool-call feed: each row shows the tool name, its arguments, latency in milliseconds and payload size in bytes.
Every tool call as it happens. Payload size is on each row because that is what the call costs the agent's context window. Which is the reason results are paginated and capped at 50 rows in the first place. View full size
The world browser, showing counts per entity: bank feed, general ledger, invoices, counterparties, notes and chart of accounts.
The generated world, entity by entity, with the messiness annotated. Aliased payees, missing references, chaotic date formats, so you can see what the agent is actually up against. View full size
The sessions screen: a card per world showing profile, seed, transaction count and the dataset hash.
A card per world. The dataset hash is on the card because same profile plus same seed must produce a byte-identical world, here it matches the value committed in CI. View full size

stack

Built with

  • Python
  • FastMCP 3
  • FastAPI
  • SQLite
  • WebSockets
  • React
  • Docker