LedgerLab
An open MCP server that hands any agent a realistically messy world, applied to bank and ledger reconciliation
- MCP
- open source
- synthetic data
- developer tooling
399
tests, green
covering determinism, both transports, and 20 concurrent sessions.
7
MCP tools, fenced by a test that fails CI if an eighth appears
3
messiness profiles, byte-identical from profile and seed
£0
to run, it calls no LLMs at all
in short
An open MCP server that hands any client a realistically messy synthetic bank and ledger: mixed date formats, missing references, aliased payee names, duplicate payments. Seven tools, no ground truth reachable through any of them, and a live viewer that streams every tool call as your agent works.
the problem
Why this is hard
Finance AI cannot be demoed on real data, and the synthetic datasets that exist are toy-clean: every reference present, every name spelled the same way, every payment matching an invoice to the penny. Agents that look brilliant on that data fall over on the first real bank feed.
LedgerLab generates books that are wrong in the ways real books are actually wrong, hands them to any MCP client in about a minute, and lets you watch the agent work. The messiness is not invented. It is the set of things that went wrong in books an ex-accountant used to close.
constraints
- Same profile plus same seed must produce a byte-identical world on any machine, every time. The dataset hashes are committed and asserted in CI.
- No tool may reach the answer key. This is checked three ways: by AST, by published schema, and by inspecting live responses.
- Seven tools, not eight. Tool bloat measurably degrades tool selection, so the count is enforced by a test.
- Results are paginated and capped at 50 rows. Rejected, not silently truncated. Because a payload is what a tool call costs an agent's context window.
where else this pattern applies
The idea is a deterministic, seeded, deliberately messy fixture world served over MCP, with the answer key held where no tool can reach it. Nothing about that is financial. The same server shape would give a support agent a messy ticket history, or a clinical agent a set of notes with inconsistent names and missing referrals.
Withholding the joins that would make the task a string comparison is the reusable design move. Any agent benchmark that hands over a clean foreign key is measuring lookup rather than reasoning, whichever domain it is dressed in.
Errors written as recoverable instructions. What went wrong, what is true, what to call next. Are a general MCP authoring practice, and the live tool-call feed showing payload size per call is how any team makes context cost visible.
architecture
How it runs
The MCP server is mounted inside the viewer's FastAPI process, which is what makes the event bus genuinely in-process: a tool call reaches the browser without a broker, a queue or a second service. Sessions are one SQLite world each, seeded on creation by a stdlib-only generator that also emits an answer key no tool can read.
01Seed a worldplan
A profile and an integer seed produce a bank, a ledger, counterparties, invoices and notes, plus an answer key held out of reach.
02Any MCP client connectstool
stdio for Claude Desktop, Streamable HTTP for LangGraph and everything else. Both tested end to end against real servers.
03The agent has to workroute
A bank transaction carries a payee name, not a foreign key. Resolving it is a tool call, which is what stops the alias messiness being cosmetic.
04Errors are part of the productpolicy
Each one says what went wrong, what is actually true, and what to call next. So a wrong guess becomes a recoverable step rather than a dead end.
05Some payments have no explanationhalt
Flagging one unknown is the correct answer. This is what stops an agent scoring well by confidently matching everything to something.
06Every call streams to the viewercommit
Tool badge, argument preview, latency, payload size. Pausing buffers rather than disconnecting, so nothing is lost.
07Grade what is in the sessionaudit
The acceptance gate grades what actually ended up in the database, not what the client claims it did.
- passed
- awaiting a person
- stopped
decisions
What was considered, and why it was ruled out
Every one of these had a reasonable alternative. The alternative is named.
ADR-01
Withhold the joins that would make reconciliation a string comparison
- Alternative
- Put counterparty_id on bank transactions and a transaction id on GL entries, as a clean schema would.
- Why not
- A bank feed gives you a name, not a foreign key, and the GL holds what was booked while the bank side is unposted. Adding either join would make propose_match solvable by string equality and the acceptance test meaningless. The omissions are the product; they are documented as design decisions rather than left to look like bugs.
ADR-02
Cap results at 50 rows and reject rather than truncate
- Alternative
- Return everything, or truncate silently with a note in the response.
- Why not
- Silent truncation gives an agent a wrong answer that looks complete, which is the worst available failure. Rejecting forces pagination to be handled. Payload size is shown on every row of the live feed for the same reason. It is what a tool call costs an agent's context window, and it is the reason the cap exists at all.
ADR-03
Make the acceptance gate's headroom problem public rather than tightening the floor
- Alternative
- Raise the thresholds to something that looks discriminating, or quietly leave the 70% floor unexplained.
- Why not
- A roughly 200-line deterministic bookkeeping heuristic scores 100% coverage and 100% precision on all three profiles. So the gate proves an agent can complete this job reproducibly with zero API misuse. A genuine regression test. But it does not distinguish a good agent from a great one. Presenting a 70% floor as if it were tight would be dishonest, and making the world genuinely hard is the first item on the roadmap instead.
ADR-04
Mount the MCP server inside the viewer process
- Alternative
- Run the MCP server and the viewer API as two services with a message broker between them.
- Why not
- The event bus is the feature. Watching tool calls arrive live is most of the value. In-process means a call reaches the browser with no broker, no queue and no second service to run, which keeps the whole thing to one uvicorn process and a make target. Fan-out to two viewers, per-session isolation and backfill for a late joiner are all tested.
evaluation
How it was measured
The acceptance gate runs a reference LangGraph agent against a real server over Streamable HTTP and grades what ended up in the session rather than what the client reported. It passes on all three profiles, and the honest reading of that is in the limits below.
| metric | result | note |
|---|---|---|
| clean profile | 100% coverage · 100% precision | 0 tool errors |
| realistic profile | 100% / 100% | cause accuracy 23/23 |
| nightmare profile | 100% / 100% | cause accuracy 66/66 |
| Determinism | hashes identical | across interpreters, against a patched clock, against committed hashes |
| Concurrency | 20 sessions | reads and writes, no cross-session bleed |
| Tests | 398 passed, 1 skipped | the skip is the live-LLM test |
try it
Connect your own client
Not hosted yet. These are the exact steps against a local instance, and the same ones against the hosted URL when it lands.
LedgerLab is an MCP server, so the useful demo is not a screenshot. It is pointing your own client at it and watching your agent work a messy set of books. Seven tools, no key, no account, read-only against a seeded synthetic bank.
01Claude Desktop, or any client over Streamable HTTP
{ "mcpServers": { "ledgerlab": { "type": "http", "url": "http://localhost:8000/mcp" } } }Restart the client, then ask it to reconcile the 2025-Q1 bank statement.
02LangGraph, or anything using the MCP adapters
client = MultiServerMCPClient({ "ledgerlab": {"transport": "streamable_http", "url": "http://localhost:8000/mcp"}, }, handle_tool_errors=False) tools = await client.get_tools() # all 7handle_tool_errors=False matters: the errors are part of the product, and your agent should be able to tell one from an answer.
03Just list the tools
npx -y @modelcontextprotocol/inspector --cli http://localhost:8000/mcp \ --transport http --method tools/list
screens
What it looks like running
Captures of the real thing. Projects without a capture show none, nothing here is a mockup.




stack
Built with
- Python
- FastMCP 3
- FastAPI
- SQLite
- WebSockets
- React
- Docker
Aneeq Khatri