Skip to content

FinAgent-Evals

A vendor-neutral benchmark for reliability under repetition, applied to finance reconciliation tasks

  • evals
  • benchmark
  • reliability
  • open source
Walkthrough on request

−37.8%

gap between pass@1 and pass^4

on the mock-flaky configuration, which scores 45.8% on one attempt and 8.0% across four. A smaller gap means a more reliable agent; a single-shot eval reports none of this.

100

cases across four difficulty tiers, frozen and hashed

12

properties the evals-of-the-evals check in CI

298

tests, 273 Python, 25 SPA

in short

An open benchmark for agentic finance behaviour rather than model quality. It holds the model constant and asks whether the configuration around it picks the right tool, escalates when evidence runs out, and gives the same answer four times running. Every case runs k=4 and publishes pass@1 beside pass^4.

the problem

Why this is hard

Most finance-AI benchmarks ask how good the model is. That is the right question, and somebody else is asking it. This one holds the model constant and asks a different one: does the agent configuration around it behave responsibly? Does it pick the right tool or guess from the prompt? When the evidence runs out, does it escalate or confabulate? Does it give the same answer four times running?

The headline result is one configuration measured two ways. Asked once, it passes 45.8% of cases. Asked four times, the share it gets right every single time drops to 8.0%. That is not 46% good. It is unreliable, and a single-shot evaluation would never have told you.

constraints

  • Vendor-neutral by construction: a row on the leaderboard is a fully-declared configuration committed as a manifest, and a test asserts no adapter reads a behavioural parameter its manifest does not declare.
  • The reference adapters share a verbatim identical prompt, model, temperature and budget. Fair floors, not tuned champions. No framework is being crowned.
  • An adapter never receives the case, only a metered tool broker, so the tool path is recorded by the harness rather than self-reported.
  • The static leaderboard has no backend. It reads committed run artefacts and deploys to GitHub Pages.

where else this pattern applies

The central finding is about agents, not about ledgers: single-shot evaluation systematically overstates reliability, and the pass@1 versus pass^k gap is the amount of a score that does not survive being asked again. That is true of a coding agent, a support agent or a research agent, and any of them can adopt the k-repetition harness unchanged.

The tier-4 trap design. Cases where escalating is the only correct action, scored so that confident guessing loses. Is reusable wherever an agent must know the limits of its evidence. Clinical triage and incident response both need that measured rather than assumed.

Benchmarking configurations rather than frameworks is the methodological transfer. Any team comparing agent stacks needs the manifest discipline here, or it publishes a framework ranking it cannot defend.

architecture

How it runs

A generator turns seeded synthetic worlds into a frozen, hashed suite. Adapters receive only a metered tool broker, never the case itself. Which is the load-bearing decision, because everything else follows from it: budgets become enforceable, the tool path is observed rather than claimed, and "the agent only knows what it looked up" becomes a fact instead of a promise.

  1. 01Generate and freezeplan

    100 cases in four tiers of 25, hashed. A case is a whole reconciliation workspace, not one transaction, MatchF1 needs a real denominator.

  2. 02Metered tool brokertool

    Seven tools with enforced budgets. The adapter never sees the case; it can only ask, and every ask is recorded by the harness.

  3. 03Tier 4, escalation is the only right answerhalt

    Three trap families: a payment nothing explains, two invoices that fit equally, and a dispute whose deciding note has been removed.

  4. 04Run it four timesroute

    Every case, k=4. This is where a configuration that looked competent stops looking competent.

  5. 05Deterministic graders firstpolicy

    MatchF1, EscalationScore and PathValidity are computed offline. The expected outcome never reaches the judge.

  6. 06Pinned judge, cachedcommit

    A rubric-only LLM judge with a committed cache, so scores reproduce byte-identically.

  7. 07Publish both numbersaudit

    pass@1 and pass^4 side by side. The gap between them is the entire point of the project.

  • passed
  • awaiting a person
  • stopped

decisions

What was considered, and why it was ruled out

Every one of these had a reasonable alternative. The alternative is named.

ADR-01

Publish pass^4 next to pass@1, always

Alternative
Report pass@1 like nearly every other benchmark does.
Why not
pass@1 45.8% and pass^4 8.0% describe the same configuration. Reporting only the first would describe it as roughly half-competent, when what it actually is is unreliable. For a finance agent, a system that is right once in two attempts and consistent one time in twelve is not half a solution. It is a different category of thing, and only the second number says so.

ADR-02

The adapter never receives the case, only a metered broker

Alternative
Hand the adapter the case and trust it to report which tools it used.
Why not
Self-reported tool paths cannot be audited, and budgets cannot be enforced against an agent that already has the data. Routing everything through a broker makes the budget a hard ceiling, makes the path an observation rather than a claim, and turns "the agent only knows what it looked up" into a structural fact. The adversarial grader cases include a lying adapter, and it is caught.

ADR-03

Write evals for the evals

Alternative
Rely on the unit test suite, which is already at 273 tests.
Why not
Twelve properties are declared in prose. Each one something that, if it broke, would make every published number wrong while the unit tests stayed green. CI checks them with no network. This caught a systematic answer leak where status == paid correlated with the right answer across the whole suite, which in the ambiguity traps made supposedly-undecidable cases decidable. The benchmark would have been punishing sound reasoning.

ADR-04

Benchmark configurations, not frameworks

Alternative
Tune each framework's adapter and publish a league table, which would get far more attention.
Why not
The reference adapters are deliberately simple canonical loops sharing one verbatim identical prompt, model, temperature and budget. A framework scoring higher here means that configuration scored higher. Anyone reading a result as a framework ranking is reading it wrong, and the manifest requirement is what makes that checkable rather than a disclaimer.

evaluation

How it was measured

The checks found four bugs during the build that would each have invalidated every published number while the unit tests stayed green. That is the argument for grading the grader.

metricresultnote
mock-flaky pass@145.8%
mock-flaky pass^48.0%the gap is the finding
Answer leakfound and fixedstatus == paid correlated with the answer suite-wide
Budget realismfound and fixed60,000 tokens against a measured ~93,000 for a competent run
Holdout commitmentfound and fixedgitignore pattern silently excluded the sealed hash
Score reproducibilitybyte-identicaljudge cache-hit counts removed from scores.json

screens

What it looks like running

Captures of the real thing. Projects without a capture show none, nothing here is a mockup.

The leaderboard: two configurations with pass@1 → pass^k columns, a consistency-gap badge on each row, and three disclosure banners above the table.
The board leads with the gap, not the rank. mock holds 94.0% across all four repetitions; mock-flaky drops 45.8% → 8.0%, and the −37.8% badge is the number a single-shot eval would never have shown you. The three banners above the table are not a footnote. Synthetic data, a release-candidate suite, and no live runs are stated before any score is. View full size
The compare view, showing headline, pass@1, pass^4, consistency gap, MatchF1, EscalationScore and PathValidity side by side for two configurations.
The same two configurations metric by metric. pass@1 falls 48.2% while pass^4 falls 86.0%. The spread between those two numbers is the whole argument for running every case k times. View full size
The case explorer filtered to tier-4 traps, with one case open showing its seed, trap family, and a table of bank transactions in four different date formats with no references.
A tier-4 trap, open. Every case here contains a transaction where escalating is the only correct action, and the answer stays behind a Reveal button so you see exactly what the agent sees. Note the dates in the transaction table. 26/03/2025, 04-16-25, 03 Mar 2025, 28.02.25, and the references reading none. View full size
The methodology page, explaining what the benchmark measures and justifying each weight in the headline formula.
Every weight in the formula is argued for on the page rather than asserted. pass^k takes the largest share because an agent that is right once in four is not 25% useful in a month-end close, it is unusable. View full size

stack

Built with

  • Python
  • LangGraph
  • OpenAI Agents SDK
  • uv
  • Vite
  • GitHub Pages
Walkthrough on request