Benchmarks

Which agent should run your books?

Every cell below is a real agent, driven by a real CLI, doing real bookkeeping against a live Postgres ledger — then graded on the books it left behind. Score on one axis, what it cost to run on the other.

11 skills · 7 harness/model pairings · 77 of 77 cells measured · last run 2026-08-18T07:04:36Z

A founder who wants an agent to run their books picks two things before getting any value: which agent harness, and which model inside it. That choice is usually made on brand or on whatever was already open. We measure it for our own reasons — the eval harness below is what tells us whether a change to one of our skills helped — so we publish what it says.

Each case hands one agent one skill, one prompt, and a ledger seeded to a fixture state. The agent works through Economico's MCP tools like any customer's would. Then the case is graded on the ledger it left — did the contract end up active, did the metered line land in the right account, did the fee post as its own leg, did the read-only skill stay read-only. Nobody scores the prose.

The numbers are as much a scoreboard for our own guidance as for anyone's model. A low cell may mean the model struggled; it may equally mean our skill did not tell it enough. Both are worth publishing.

Read this before you quote a number

  • Costs are not comparable across harnesses. Some cells carry the vendor's own figure and others are derived from published list prices, and a run's token count partly measures how verbose our own instructions are rather than how efficient the model is. Read the dollar column as an order of magnitude, not a price list.
  • Every cell is one to three rollouts. A case is pass or fail, so a three-case skill moves in 33-point steps. Treat a gap of one case as noise.
  • The cases move too, and so does the harness. A cell measured under an older case manifest, or under a tool surface we have since pinned, is marked stale — a score that dropped may mean we made the case stricter or the agent had a different set of tools, not that the model got worse.
  • Empty cells are honest. A pairing we never ran on a skill renders as a gap, and a harness that reports no token usage gets no dollar figure rather than a zero.

By pairing

What each agent scores, and what it burned

ModelHarnessSkillsMeanRangeRolloutsCostCost sourceLast run
claude-haiku-4-5claude11/1186.1%33–100%1$6.61harness 11/11 skills priced2026-08-17T19:56:57Z
claude-opus-5 mediumclaude11/11100.0%100–100%1–3$35.20harness 11/11 skills priced2026-08-18T07:04:36Z
claude-sonnet-5 mediumclaude11/11100.0%100–100%1$14.94harness 11/11 skills priced2026-08-17T12:56:50Z
gpt-5.6-lunacodex11/1189.9%22–100%1–3$0.89estimated 11/11 skills priced2026-08-18T06:37:11Z
gpt-5.6-solcodex11/11100.0%100–100%1$14.64estimated 11/11 skills priced2026-08-17T12:40:04Z
gpt-5.6-terracodex11/1198.0%78–100%1–3$8.45estimated 11/11 skills priced2026-08-18T05:15:40Z
grok-4.6grok11/11100.0%100–100%1$4.33harness 11/11 skills priced2026-08-17T13:26:29Z

Cost totals what that pairing's recorded skills consumed end to end. harness means the CLI reported the dollar figure itself; estimated means we derived it from the list prices checked in at skills/evals/skillopt/report/prices.json, as of 2026-08-16. It is an estimate for this set of eval cases, not a bill anyone received.

The matrix

Skill by pairing

Skillclaude-haiku-4-5claude-opus-5
medium
claude-sonnet-5
medium
gpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
company-setup67% 2/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3
cpa100% 2/2100% 2/2100% 2/2100% 2/2100% 2/2100% 2/2100% 2/2
creating-contracts100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3
expense-tracking100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3
financial-analyst33% 1/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3
forecasting100% 3/3100% 3/3100% 3/322% 0/3100% 3/378% 3/3100% 3/3
investor-reporting100% 3/3100% 3/3100% 3/367% 2/3100% 3/3100% 3/3100% 3/3
invoicing80% 4/5100% 5/5100% 5/5100% 5/5100% 5/5100% 5/5100% 5/5
payment-rails67% 2/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3
pricing100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3
setup-economico100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3100% 3/3

A percentage is the mean of that skill's held-out cases; the small number under it is cases passed out of cases scored.

Case by case

What each skill is actually being asked to do

company-setup

First-run setup of the business itself — the legal-entity profile, named bank and wallet accounts as one pool of value each, and a cap table with founder shares and an early SAFE.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
company-accounts-one-pool 6/7 passed
company-captable-safe 7/7 passed
company-profile-wrong-entity 7/7 passed
Run$0.30 2026-08-17$1.20 2026-08-17$0.60 2026-08-17$0.04 2026-08-17$1.18 2026-08-17$0.40 2026-08-17$0.24 2026-08-17

cpa

An auditor's read: period-end accruals to recommend, and COGS-versus-OpEx classification calls — findings only, never posting.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
cpa-cogs-opex-classification 7/7 passed
cpa-period-end-accruals-recommend 7/7 passed
Run$0.29 2026-08-17$4.32 2026-08-18$0.80 2026-08-17$0.04 2026-08-17$0.99 2026-08-17$0.36 2026-08-17$0.22 2026-08-17

creating-contracts

Build the contract spine a bill or invoice hangs off — an annual SaaS order form, a usage-metered vendor contract, and an upsell that must amend the existing contract instead of replacing it.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
contracts-annual-saas 7/7 passed
contracts-upsell-amendment 7/7 passed
contracts-vendor-llm-usage 7/7 passed
Run$1.07 2026-08-17$1.77 2026-08-17$0.74 2026-08-17$0.07 2026-08-17$1.37 2026-08-17$0.66 2026-08-17$0.29 2026-08-17

expense-tracking

Accounts payable from a vendor receipt: split infrastructure spend between production COGS and R&D, build the vendor spine for metered LLM usage, and notice when a bill is already booked.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
expense-aws-already-booked 7/7 passed
expense-openai-usage-spine 7/7 passed
expense-render-dev-rnd 7/7 passed
Run$0.70 2026-08-17$2.39 2026-08-17$1.49 2026-08-17$0.09 2026-08-17$1.90 2026-08-17$0.86 2026-08-17$0.51 2026-08-17

financial-analyst

Read-only analysis that has to stay grounded in the ledger: committed run-rate, a dead annual contract that needs writing off, and a quarter-over-quarter trend the books cannot actually support — where the right answer is to refuse.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
analyst-committed-runrate 6/7 passed
analyst-dead-annual-writeoff 7/7 passed
analyst-quarter-trend-refusal 6/7 passed
Run$0.39 2026-08-17$1.01 2026-08-17$0.99 2026-08-17$0.04 2026-08-17$0.96 2026-08-17$0.42 2026-08-17$0.21 2026-08-17

forecasting

What-ifs that must stay off the real ledger: a price-hike overlay, a hiring-plan overlay, and a signed vendor bill that belongs on the real books rather than in a scenario.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
forecasting-hiring-plan-overlay 6/7 passed
forecasting-price-hike-overlay 6/7 passed
forecasting-signed-vendor-bill-real-books 6/7 passed
Run$0.80 2026-08-17$17.00 2026-08-18$5.96 2026-08-17$0.25 2026-08-18$2.22 2026-08-17$2.58 2026-08-18$1.38 2026-08-17

investor-reporting

Investor-facing metrics from the same books — a consulting retainer base, a SaaS cost base, and a board-ready report share that must be scoped and revocable.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
investor-reporting-board-share-link 7/7 passed
investor-reporting-consulting-retainer-base 6/7 passed
investor-reporting-saas-cost-base 7/7 passed
Run$0.20 2026-08-17$0.72 2026-08-17$0.37 2026-08-17$0.02 2026-08-17$0.63 2026-08-17$0.28 2026-08-17$0.12 2026-08-17

invoicing

Bill a customer correctly: grant tranches, an ad-hoc charge that still needs an obligation, a milestone that was already billed and must not be billed again, and a payment recorded against the right funding account.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
invoicing-ad-hoc-obligation 7/7 passed
invoicing-already-billed-milestone 6/7 passed
invoicing-grant-funded 7/7 passed
invoicing-kinoko-paid 7/7 passed
invoicing-kinoko-paid-no-account 7/7 passed
Run$0.87 2026-08-17$1.74 2026-08-17$1.27 2026-08-17$0.13 2026-08-17$1.71 2026-08-17$0.63 2026-08-17$0.44 2026-08-17

payment-rails

Model how money moves — one address reachable on two chains, quoting routes before paying, and posting a wire fee as its own journal leg rather than netting it into the payment.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
rails-one-address-two-chains 7/7 passed
rails-quote-before-paying 7/7 passed
rails-wire-fee-account 6/7 passed
Run$0.72 2026-08-17$1.42 2026-08-17$0.76 2026-08-17$0.05 2026-08-17$1.12 2026-08-17$0.51 2026-08-17$0.27 2026-08-17

pricing

Turn a business model into priced catalog items and obligations: a bespoke enterprise deal, per-call agent pricing, and prepaid packs that have to draw down rather than bill twice.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
pricing-agent-per-call 7/7 passed
pricing-bespoke-enterprise 7/7 passed
pricing-prepaid-packs 7/7 passed
Run$1.05 2026-08-17$1.77 2026-08-17$1.00 2026-08-17$0.05 2026-08-17$1.26 2026-08-17$1.22 2026-08-17$0.33 2026-08-17

setup-economico

Connect the ledger and prove the connection: rehearse in a sandbox scenario without touching the real books, take a first real customer live, and register a headless agent without handing it write access it does not need.

Caseclaude-haiku-4-5claude-opus-5 mediumclaude-sonnet-5 mediumgpt-5.6-lunagpt-5.6-solgpt-5.6-terragrok-4.6
setup-go-live-real-customer 7/7 passed
setup-headless-agent 7/7 passed
setup-sandbox-dry-run 7/7 passed
Run$0.23 2026-08-17$1.85 2026-08-17$0.96 2026-08-17$0.10 2026-08-17$1.29 2026-08-17$0.52 2026-08-17$0.32 2026-08-17

Methodology

How these numbers are made

What a case is

One prompt, one skill, one seeded ledger. The agent runs in its own CLI (Claude Code, Codex, Grok) against a throwaway Postgres-backed instance of Economico with the real MCP tool surface, and is given the same skill file we publish. Nothing is mocked.

What a pass means

Each case asserts against the ledger the agent leaves behind — the parties, contracts, obligations, invoices, bills, journals, and scenarios that exist afterwards, and the account codes they landed on. A few cases assert on a grounded answer instead, where the right behaviour is to report a number or to refuse. A case is pass or fail; a skill's score is the mean of its cases.

Which runs get published here

Only held-out cases, only runs of the checked-in skill, and only runs where every case actually ran — a sweep where a case skipped or errored proves nothing and is discarded rather than averaged. Optimizer scratch runs of candidate skill files never appear.

How many rollouts back a number

One to three per case, shown per pairing above. That is enough to separate a model that cannot do the work from one that can, and not enough to rank two close cells. We do not publish error bars we cannot honestly compute from this many samples.

How a cell can be reproduced

Every cell records the repository commit it ran at, the hash of the case manifest, the tool surface the agent was given, and the wall-clock window of the run. Compare two cells only when their case-manifest hashes and tool surfaces match; the machine-readable benchmarks feed carries all of them per cell.

What we deliberately do not publish

The case prompts, the fixtures, and the grader's reasons stay private, and a held-out set stays unpublished entirely. A published eval is a gameable eval, and these numbers are only worth anything while nobody can train against them.

Where the case manifests stand

  • 383a5380ba3b — 77 cells, last run 2026-08-18 (the manifest in the tree today)

Get started

The finance team you don't have to hire.

Whichever agent you land on, it talks to the same ledger. Connect the one you already use and run your own books through it.