Benchmarks
Which agent should run your books?
Every cell below is a real agent, driven by a real CLI, doing real bookkeeping against a live Postgres ledger — then graded on the books it left behind. Score on one axis, what it cost to run on the other.
11 skills · 7 harness/model pairings · 77 of 77 cells measured · last run 2026-08-18T07:04:36Z
A founder who wants an agent to run their books picks two things before getting any value: which agent harness, and which model inside it. That choice is usually made on brand or on whatever was already open. We measure it for our own reasons — the eval harness below is what tells us whether a change to one of our skills helped — so we publish what it says.
Each case hands one agent one skill, one prompt, and a ledger seeded to a fixture state. The agent works through Economico's MCP tools like any customer's would. Then the case is graded on the ledger it left — did the contract end up active, did the metered line land in the right account, did the fee post as its own leg, did the read-only skill stay read-only. Nobody scores the prose.
The numbers are as much a scoreboard for our own guidance as for anyone's model. A low cell may mean the model struggled; it may equally mean our skill did not tell it enough. Both are worth publishing.
Read this before you quote a number
- Costs are not comparable across harnesses. Some cells carry the vendor's own figure and others are derived from published list prices, and a run's token count partly measures how verbose our own instructions are rather than how efficient the model is. Read the dollar column as an order of magnitude, not a price list.
- Every cell is one to three rollouts. A case is pass or fail, so a three-case skill moves in 33-point steps. Treat a gap of one case as noise.
- The cases move too, and so does the harness. A cell measured under an older case manifest, or under a tool surface we have since pinned, is marked stale — a score that dropped may mean we made the case stricter or the agent had a different set of tools, not that the model got worse.
- Empty cells are honest. A pairing we never ran on a skill renders as a gap, and a harness that reports no token usage gets no dollar figure rather than a zero.
By pairing
What each agent scores, and what it burned
| Model | Harness | Skills | Mean | Range | Rollouts | Cost | Cost source | Last run |
|---|---|---|---|---|---|---|---|---|
claude-haiku-4-5 | claude | 11/11 | 86.1% | 33–100% | 1 | $6.61 | harness 11/11 skills priced | 2026-08-17T19:56:57Z |
claude-opus-5 medium | claude | 11/11 | 100.0% | 100–100% | 1–3 | $35.20 | harness 11/11 skills priced | 2026-08-18T07:04:36Z |
claude-sonnet-5 medium | claude | 11/11 | 100.0% | 100–100% | 1 | $14.94 | harness 11/11 skills priced | 2026-08-17T12:56:50Z |
gpt-5.6-luna | codex | 11/11 | 89.9% | 22–100% | 1–3 | $0.89 | estimated 11/11 skills priced | 2026-08-18T06:37:11Z |
gpt-5.6-sol | codex | 11/11 | 100.0% | 100–100% | 1 | $14.64 | estimated 11/11 skills priced | 2026-08-17T12:40:04Z |
gpt-5.6-terra | codex | 11/11 | 98.0% | 78–100% | 1–3 | $8.45 | estimated 11/11 skills priced | 2026-08-18T05:15:40Z |
grok-4.6 | grok | 11/11 | 100.0% | 100–100% | 1 | $4.33 | harness 11/11 skills priced | 2026-08-17T13:26:29Z |
Cost totals what that pairing's recorded skills consumed end to end.
harness means the CLI reported the dollar figure itself;
estimated means we derived it from the list prices checked in
at skills/evals/skillopt/report/prices.json, as of
2026-08-16. It is an estimate for this set of eval cases, not
a bill anyone received.
The matrix
Skill by pairing
| Skill | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
company-setup | 67% 2/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
cpa | 100% 2/2 | 100% 2/2 | 100% 2/2 | 100% 2/2 | 100% 2/2 | 100% 2/2 | 100% 2/2 |
creating-contracts | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
expense-tracking | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
financial-analyst | 33% 1/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
forecasting | 100% 3/3 | 100% 3/3 | 100% 3/3 | 22% 0/3 | 100% 3/3 | 78% 3/3 | 100% 3/3 |
investor-reporting | 100% 3/3 | 100% 3/3 | 100% 3/3 | 67% 2/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
invoicing | 80% 4/5 | 100% 5/5 | 100% 5/5 | 100% 5/5 | 100% 5/5 | 100% 5/5 | 100% 5/5 |
payment-rails | 67% 2/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
pricing | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
setup-economico | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 | 100% 3/3 |
A percentage is the mean of that skill's held-out cases; the small number under it is cases passed out of cases scored.
Case by case
What each skill is actually being asked to do
company-setup
First-run setup of the business itself — the legal-entity profile, named bank and wallet accounts as one pool of value each, and a cap table with founder shares and an early SAFE.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
company-accounts-one-pool 6/7 passed | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
company-captable-safe 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
company-profile-wrong-entity 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.30 2026-08-17 | $1.20 2026-08-17 | $0.60 2026-08-17 | $0.04 2026-08-17 | $1.18 2026-08-17 | $0.40 2026-08-17 | $0.24 2026-08-17 |
cpa
An auditor's read: period-end accruals to recommend, and COGS-versus-OpEx classification calls — findings only, never posting.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
cpa-cogs-opex-classification 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
cpa-period-end-accruals-recommend 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.29 2026-08-17 | $4.32 2026-08-18 | $0.80 2026-08-17 | $0.04 2026-08-17 | $0.99 2026-08-17 | $0.36 2026-08-17 | $0.22 2026-08-17 |
creating-contracts
Build the contract spine a bill or invoice hangs off — an annual SaaS order form, a usage-metered vendor contract, and an upsell that must amend the existing contract instead of replacing it.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
contracts-annual-saas 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
contracts-upsell-amendment 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
contracts-vendor-llm-usage 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $1.07 2026-08-17 | $1.77 2026-08-17 | $0.74 2026-08-17 | $0.07 2026-08-17 | $1.37 2026-08-17 | $0.66 2026-08-17 | $0.29 2026-08-17 |
expense-tracking
Accounts payable from a vendor receipt: split infrastructure spend between production COGS and R&D, build the vendor spine for metered LLM usage, and notice when a bill is already booked.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
expense-aws-already-booked 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
expense-openai-usage-spine 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
expense-render-dev-rnd 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.70 2026-08-17 | $2.39 2026-08-17 | $1.49 2026-08-17 | $0.09 2026-08-17 | $1.90 2026-08-17 | $0.86 2026-08-17 | $0.51 2026-08-17 |
financial-analyst
Read-only analysis that has to stay grounded in the ledger: committed run-rate, a dead annual contract that needs writing off, and a quarter-over-quarter trend the books cannot actually support — where the right answer is to refuse.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
analyst-committed-runrate 6/7 passed | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
analyst-dead-annual-writeoff 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
analyst-quarter-trend-refusal 6/7 passed | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.39 2026-08-17 | $1.01 2026-08-17 | $0.99 2026-08-17 | $0.04 2026-08-17 | $0.96 2026-08-17 | $0.42 2026-08-17 | $0.21 2026-08-17 |
forecasting
What-ifs that must stay off the real ledger: a price-hike overlay, a hiring-plan overlay, and a signed vendor bill that belongs on the real books rather than in a scenario.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
forecasting-hiring-plan-overlay 6/7 passed | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ |
forecasting-price-hike-overlay 6/7 passed | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ |
forecasting-signed-vendor-bill-real-books 6/7 passed | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ |
| Run | $0.80 2026-08-17 | $17.00 2026-08-18 | $5.96 2026-08-17 | $0.25 2026-08-18 | $2.22 2026-08-17 | $2.58 2026-08-18 | $1.38 2026-08-17 |
investor-reporting
Investor-facing metrics from the same books — a consulting retainer base, a SaaS cost base, and a board-ready report share that must be scoped and revocable.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
investor-reporting-board-share-link 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
investor-reporting-consulting-retainer-base 6/7 passed | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ |
investor-reporting-saas-cost-base 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.20 2026-08-17 | $0.72 2026-08-17 | $0.37 2026-08-17 | $0.02 2026-08-17 | $0.63 2026-08-17 | $0.28 2026-08-17 | $0.12 2026-08-17 |
invoicing
Bill a customer correctly: grant tranches, an ad-hoc charge that still needs an obligation, a milestone that was already billed and must not be billed again, and a payment recorded against the right funding account.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
invoicing-ad-hoc-obligation 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
invoicing-already-billed-milestone 6/7 passed | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
invoicing-grant-funded 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
invoicing-kinoko-paid 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
invoicing-kinoko-paid-no-account 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.87 2026-08-17 | $1.74 2026-08-17 | $1.27 2026-08-17 | $0.13 2026-08-17 | $1.71 2026-08-17 | $0.63 2026-08-17 | $0.44 2026-08-17 |
payment-rails
Model how money moves — one address reachable on two chains, quoting routes before paying, and posting a wire fee as its own journal leg rather than netting it into the payment.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
rails-one-address-two-chains 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
rails-quote-before-paying 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
rails-wire-fee-account 6/7 passed | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.72 2026-08-17 | $1.42 2026-08-17 | $0.76 2026-08-17 | $0.05 2026-08-17 | $1.12 2026-08-17 | $0.51 2026-08-17 | $0.27 2026-08-17 |
pricing
Turn a business model into priced catalog items and obligations: a bespoke enterprise deal, per-call agent pricing, and prepaid packs that have to draw down rather than bill twice.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
pricing-agent-per-call 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
pricing-bespoke-enterprise 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
pricing-prepaid-packs 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $1.05 2026-08-17 | $1.77 2026-08-17 | $1.00 2026-08-17 | $0.05 2026-08-17 | $1.26 2026-08-17 | $1.22 2026-08-17 | $0.33 2026-08-17 |
setup-economico
Connect the ledger and prove the connection: rehearse in a sandbox scenario without touching the real books, take a first real customer live, and register a headless agent without handing it write access it does not need.
| Case | claude-haiku-4-5 | claude-opus-5 medium | claude-sonnet-5 medium | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | grok-4.6 |
|---|---|---|---|---|---|---|---|
setup-go-live-real-customer 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
setup-headless-agent 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
setup-sandbox-dry-run 7/7 passed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Run | $0.23 2026-08-17 | $1.85 2026-08-17 | $0.96 2026-08-17 | $0.10 2026-08-17 | $1.29 2026-08-17 | $0.52 2026-08-17 | $0.32 2026-08-17 |
Methodology
How these numbers are made
What a case is
One prompt, one skill, one seeded ledger. The agent runs in its own CLI (Claude Code, Codex, Grok) against a throwaway Postgres-backed instance of Economico with the real MCP tool surface, and is given the same skill file we publish. Nothing is mocked.
What a pass means
Each case asserts against the ledger the agent leaves behind — the parties, contracts, obligations, invoices, bills, journals, and scenarios that exist afterwards, and the account codes they landed on. A few cases assert on a grounded answer instead, where the right behaviour is to report a number or to refuse. A case is pass or fail; a skill's score is the mean of its cases.
Which runs get published here
Only held-out cases, only runs of the checked-in skill, and only runs where every case actually ran — a sweep where a case skipped or errored proves nothing and is discarded rather than averaged. Optimizer scratch runs of candidate skill files never appear.
How many rollouts back a number
One to three per case, shown per pairing above. That is enough to separate a model that cannot do the work from one that can, and not enough to rank two close cells. We do not publish error bars we cannot honestly compute from this many samples.
How a cell can be reproduced
Every cell records the repository commit it ran at, the hash of the case manifest, the tool surface the agent was given, and the wall-clock window of the run. Compare two cells only when their case-manifest hashes and tool surfaces match; the machine-readable benchmarks feed carries all of them per cell.
What we deliberately do not publish
The case prompts, the fixtures, and the grader's reasons stay private, and a held-out set stays unpublished entirely. A published eval is a gameable eval, and these numbers are only worth anything while nobody can train against them.
Where the case manifests stand
383a5380ba3b— 77 cells, last run 2026-08-18 (the manifest in the tree today)
Get started
The finance team you don't have to hire.
Whichever agent you land on, it talks to the same ledger. Connect the one you already use and run your own books through it.