Skip to content

domain

AI for Finance & Accounting

Agents for the one domain where a wrong number is money lost and a broken audit.

I build AI agents for accounting and finance. Month-end close, bank reconciliation, invoice approval. With the governance the work actually requires: postings that can be proven after the fact, humans gating every commit, and refusal instead of guessing when the evidence is thin. I reconciled real books before I automated any of it.

  • passed
  • awaiting a person
  • stopped
Illustrative replay. The approval step is simulated; in the real systems a person signs it.

the bar

What finance demands of an agent

Seven requirements the domain imposes before an agent is allowed anywhere near the books. Each one links to the system on this site that demonstrates it.

Every posting must be provable after the fact, not merely logged.

LedgerGuard

The model may pick the tool and explain the result. It must never be the thing that computes the number.

RevLedger

The thing that orchestrates the work cannot also approve it.

CloseOps

When the evidence is weak, the correct output is a refusal, not a plausible answer.

PolicyGround

A human gates the commit. The agent prepares it.

InvoiceAudit

Right once is not right. A run has to survive being repeated.

FinAgent-Evals

Uncertainty must be triageable, not hidden behind a confidence number.

LedgerLens

before any of this

I started as an accounts assistant, which in practice meant I was the person the bank statement landed on.

I reconciled bank feeds against the ledger, chased purchase orders that authorised one amount while the invoice billed another, matched payments to counterparties whose names were spelled three different ways across two systems, and wrote off the ones nothing in the books could explain.

A close that goes wrong does not announce itself. It is quiet, the books balance, the queue is shorter than yesterday, and three weeks later someone finds the duplicate posting after the numbers have already been signed.

I moved into software because most of that work was repetitive in a way that obviously should not have needed a person, and then into AI because agents are the first tools that can actually do the judgement-shaped parts of it.

Which is why everything I build is organised around proving the agent is right rather than making it capable: I already know exactly what it costs when nobody checks.

the work

Four things that go wrong, and what was built for them

Every dataset behind these is synthetic. The failure modes are real: each one happens in a real finance function, and each links to the system built against it.

month-end close

what goes wrong

A duplicate journal entry gets posted during a close and nobody finds it for three weeks, by which point the books are signed. The failure is silent: the ledger balances and the queue looks shorter than yesterday.

what was built

A checklist with a DAG and risk tiers, idempotent dispatch, and an orchestrator that is structurally incapable of approving its own work. Five agents, a human gate at every risk boundary.

evidence

0

duplicate postings

CloseOps

bank reconciliation

what goes wrong

An exception queue that a person has to work through one item at a time, where most of the effort is not deciding what to do but reconstructing what happened: which payment, which counterparty, spelled which way, in which of two systems.

what was built

Root-cause hypotheses that must cite the records they came from. An uncited hypothesis cannot reach the database at all, so the reviewer is triaging evidence rather than trusting a score.

evidence

96%

top-1 root-cause accuracy

LedgerLens

auditability

what goes wrong

Someone asks how a number was arrived at, months later. A log tells you what the system did; it does not tell you what evidence the decision rested on, or why a reviewer overrode it.

what was built

An audit trail that is an architecture rather than a log, plus a reviewer console where a human decision is captured as data, including the disagreements.

evidence

38

decisions in one reconciliation

LedgerGuard

trusting the agent at all

what goes wrong

An agent demo works. It works again. Then it is put in front of a month of real volume and the behaviour that was fine in a single run turns out not to be stable across repeats. That stability is what matters once it is unattended.

what was built

An evaluation harness that scores the same case repeatedly and reports the gap between one attempt and four, with a frozen, hashed case set and evals that check the evals.

evidence

−37.8%

gap between pass@1 and pass^4

FinAgent-Evals

open source

Infrastructure I've given the field

Most people have built things for finance AI. These two are for anyone building it. A public benchmark and a public fixture, both free to run.

public benchmark

FinAgent-Evals

An open benchmark for agentic finance behaviour rather than model quality. It holds the model constant and asks whether the configuration around it picks the right tool, escalates when evidence runs out, and gives the same answer four times running. Every case runs k=4 and publishes pass@1 beside pass^4.

public fixture

LedgerLab

An open MCP server that hands any client a realistically messy synthetic bank and ledger: mixed date formats, missing references, aliased payee names, duplicate payments. Seven tools, no ground truth reachable through any of them, and a live viewer that streams every tool call as your agent works.

point your own client at it

{
  "mcpServers": {
    "ledgerlab": { "type": "http", "url": "<LEDGERLAB_URL>/mcp" }
  }
}

work together

Building AI for a finance team? Let’s talk.

Whether you are putting agents near the close for the first time, or you already have some and need to know whether they can be trusted, that is the conversation I want. I have been on both sides of it. The one building the system, and the one who finds the duplicate three weeks later.