domain
AI for Finance & Accounting
Agents for the one domain where a wrong number is money lost and a broken audit.
I build AI agents for accounting and finance. Month-end close, bank reconciliation, invoice approval. With the governance the work actually requires: postings that can be proven after the fact, humans gating every commit, and refusal instead of guessing when the evidence is thin. I reconciled real books before I automated any of it.
- passed
- awaiting a person
- stopped
the bar
What finance demands of an agent
Seven requirements the domain imposes before an agent is allowed anywhere near the books. Each one links to the system on this site that demonstrates it.
Every posting must be provable after the fact, not merely logged.
The model may pick the tool and explain the result. It must never be the thing that computes the number.
The thing that orchestrates the work cannot also approve it.
When the evidence is weak, the correct output is a refusal, not a plausible answer.
A human gates the commit. The agent prepares it.
Right once is not right. A run has to survive being repeated.
Uncertainty must be triageable, not hidden behind a confidence number.
before any of this
I started as an accounts assistant, which in practice meant I was the person the bank statement landed on.
I reconciled bank feeds against the ledger, chased purchase orders that authorised one amount while the invoice billed another, matched payments to counterparties whose names were spelled three different ways across two systems, and wrote off the ones nothing in the books could explain.
A close that goes wrong does not announce itself. It is quiet, the books balance, the queue is shorter than yesterday, and three weeks later someone finds the duplicate posting after the numbers have already been signed.
I moved into software because most of that work was repetitive in a way that obviously should not have needed a person, and then into AI because agents are the first tools that can actually do the judgement-shaped parts of it.
Which is why everything I build is organised around proving the agent is right rather than making it capable: I already know exactly what it costs when nobody checks.
the work
Four things that go wrong, and what was built for them
Every dataset behind these is synthetic. The failure modes are real: each one happens in a real finance function, and each links to the system built against it.
month-end close
what goes wrong
A duplicate journal entry gets posted during a close and nobody finds it for three weeks, by which point the books are signed. The failure is silent: the ledger balances and the queue looks shorter than yesterday.
what was built
A checklist with a DAG and risk tiers, idempotent dispatch, and an orchestrator that is structurally incapable of approving its own work. Five agents, a human gate at every risk boundary.
bank reconciliation
what goes wrong
An exception queue that a person has to work through one item at a time, where most of the effort is not deciding what to do but reconstructing what happened: which payment, which counterparty, spelled which way, in which of two systems.
what was built
Root-cause hypotheses that must cite the records they came from. An uncited hypothesis cannot reach the database at all, so the reviewer is triaging evidence rather than trusting a score.
auditability
what goes wrong
Someone asks how a number was arrived at, months later. A log tells you what the system did; it does not tell you what evidence the decision rested on, or why a reviewer overrode it.
what was built
An audit trail that is an architecture rather than a log, plus a reviewer console where a human decision is captured as data, including the disagreements.
trusting the agent at all
what goes wrong
An agent demo works. It works again. Then it is put in front of a month of real volume and the behaviour that was fine in a single run turns out not to be stable across repeats. That stability is what matters once it is unattended.
what was built
An evaluation harness that scores the same case repeatedly and reports the gap between one attempt and four, with a frozen, hashed case set and evals that check the evals.
open source
Infrastructure I've given the field
Most people have built things for finance AI. These two are for anyone building it. A public benchmark and a public fixture, both free to run.
public benchmark
FinAgent-Evals
An open benchmark for agentic finance behaviour rather than model quality. It holds the model constant and asks whether the configuration around it picks the right tool, escalates when evidence runs out, and gives the same answer four times running. Every case runs k=4 and publishes pass@1 beside pass^4.
public fixture
LedgerLab
An open MCP server that hands any client a realistically messy synthetic bank and ledger: mixed date formats, missing references, aliased payee names, duplicate payments. Seven tools, no ground truth reachable through any of them, and a live viewer that streams every tool call as your agent works.
point your own client at it
{
"mcpServers": {
"ledgerlab": { "type": "http", "url": "<LEDGERLAB_URL>/mcp" }
}
}work together
Building AI for a finance team? Let’s talk.
Whether you are putting agents near the close for the first time, or you already have some and need to know whether they can be trusted, that is the conversation I want. I have been on both sides of it. The one building the system, and the one who finds the duplicate three weeks later.
Aneeq Khatri