FinAgent-Evals
A vendor-neutral benchmark for reliability under repetition, applied to finance reconciliation tasks
- evals
- benchmark
- reliability
- open source
−37.8%
gap between pass@1 and pass^4
on the mock-flaky configuration, which scores 45.8% on one attempt and 8.0% across four. A smaller gap means a more reliable agent; a single-shot eval reports none of this.
100
cases across four difficulty tiers, frozen and hashed
12
properties the evals-of-the-evals check in CI
298
tests, 273 Python, 25 SPA
in short
An open benchmark for agentic finance behaviour rather than model quality. It holds the model constant and asks whether the configuration around it picks the right tool, escalates when evidence runs out, and gives the same answer four times running. Every case runs k=4 and publishes pass@1 beside pass^4.
the problem
Why this is hard
Most finance-AI benchmarks ask how good the model is. That is the right question, and somebody else is asking it. This one holds the model constant and asks a different one: does the agent configuration around it behave responsibly? Does it pick the right tool or guess from the prompt? When the evidence runs out, does it escalate or confabulate? Does it give the same answer four times running?
The headline result is one configuration measured two ways. Asked once, it passes 45.8% of cases. Asked four times, the share it gets right every single time drops to 8.0%. That is not 46% good. It is unreliable, and a single-shot evaluation would never have told you.
constraints
- Vendor-neutral by construction: a row on the leaderboard is a fully-declared configuration committed as a manifest, and a test asserts no adapter reads a behavioural parameter its manifest does not declare.
- The reference adapters share a verbatim identical prompt, model, temperature and budget. Fair floors, not tuned champions. No framework is being crowned.
- An adapter never receives the case, only a metered tool broker, so the tool path is recorded by the harness rather than self-reported.
- The static leaderboard has no backend. It reads committed run artefacts and deploys to GitHub Pages.
where else this pattern applies
The central finding is about agents, not about ledgers: single-shot evaluation systematically overstates reliability, and the pass@1 versus pass^k gap is the amount of a score that does not survive being asked again. That is true of a coding agent, a support agent or a research agent, and any of them can adopt the k-repetition harness unchanged.
The tier-4 trap design. Cases where escalating is the only correct action, scored so that confident guessing loses. Is reusable wherever an agent must know the limits of its evidence. Clinical triage and incident response both need that measured rather than assumed.
Benchmarking configurations rather than frameworks is the methodological transfer. Any team comparing agent stacks needs the manifest discipline here, or it publishes a framework ranking it cannot defend.
architecture
How it runs
A generator turns seeded synthetic worlds into a frozen, hashed suite. Adapters receive only a metered tool broker, never the case itself. Which is the load-bearing decision, because everything else follows from it: budgets become enforceable, the tool path is observed rather than claimed, and "the agent only knows what it looked up" becomes a fact instead of a promise.
01Generate and freezeplan
100 cases in four tiers of 25, hashed. A case is a whole reconciliation workspace, not one transaction, MatchF1 needs a real denominator.
02Metered tool brokertool
Seven tools with enforced budgets. The adapter never sees the case; it can only ask, and every ask is recorded by the harness.
03Tier 4, escalation is the only right answerhalt
Three trap families: a payment nothing explains, two invoices that fit equally, and a dispute whose deciding note has been removed.
04Run it four timesroute
Every case, k=4. This is where a configuration that looked competent stops looking competent.
05Deterministic graders firstpolicy
MatchF1, EscalationScore and PathValidity are computed offline. The expected outcome never reaches the judge.
06Pinned judge, cachedcommit
A rubric-only LLM judge with a committed cache, so scores reproduce byte-identically.
07Publish both numbersaudit
pass@1 and pass^4 side by side. The gap between them is the entire point of the project.
- passed
- awaiting a person
- stopped
decisions
What was considered, and why it was ruled out
Every one of these had a reasonable alternative. The alternative is named.
ADR-01
Publish pass^4 next to pass@1, always
- Alternative
- Report pass@1 like nearly every other benchmark does.
- Why not
- pass@1 45.8% and pass^4 8.0% describe the same configuration. Reporting only the first would describe it as roughly half-competent, when what it actually is is unreliable. For a finance agent, a system that is right once in two attempts and consistent one time in twelve is not half a solution. It is a different category of thing, and only the second number says so.
ADR-02
The adapter never receives the case, only a metered broker
- Alternative
- Hand the adapter the case and trust it to report which tools it used.
- Why not
- Self-reported tool paths cannot be audited, and budgets cannot be enforced against an agent that already has the data. Routing everything through a broker makes the budget a hard ceiling, makes the path an observation rather than a claim, and turns "the agent only knows what it looked up" into a structural fact. The adversarial grader cases include a lying adapter, and it is caught.
ADR-03
Write evals for the evals
- Alternative
- Rely on the unit test suite, which is already at 273 tests.
- Why not
- Twelve properties are declared in prose. Each one something that, if it broke, would make every published number wrong while the unit tests stayed green. CI checks them with no network. This caught a systematic answer leak where status == paid correlated with the right answer across the whole suite, which in the ambiguity traps made supposedly-undecidable cases decidable. The benchmark would have been punishing sound reasoning.
ADR-04
Benchmark configurations, not frameworks
- Alternative
- Tune each framework's adapter and publish a league table, which would get far more attention.
- Why not
- The reference adapters are deliberately simple canonical loops sharing one verbatim identical prompt, model, temperature and budget. A framework scoring higher here means that configuration scored higher. Anyone reading a result as a framework ranking is reading it wrong, and the manifest requirement is what makes that checkable rather than a disclaimer.
evaluation
How it was measured
The checks found four bugs during the build that would each have invalidated every published number while the unit tests stayed green. That is the argument for grading the grader.
| metric | result | note |
|---|---|---|
| mock-flaky pass@1 | 45.8% | |
| mock-flaky pass^4 | 8.0% | the gap is the finding |
| Answer leak | found and fixed | status == paid correlated with the answer suite-wide |
| Budget realism | found and fixed | 60,000 tokens against a measured ~93,000 for a competent run |
| Holdout commitment | found and fixed | gitignore pattern silently excluded the sealed hash |
| Score reproducibility | byte-identical | judge cache-hit counts removed from scores.json |
screens
What it looks like running
Captures of the real thing. Projects without a capture show none, nothing here is a mockup.




stack
Built with
- Python
- LangGraph
- OpenAI Agents SDK
- uv
- Vite
- GitHub Pages
Aneeq Khatri