KVAi Research · Evals · First edition, October 2026

How we measure AI models on real work

A new model comes out every month. The question is always the same: should we switch?

To answer it we built a test bench out of tasks from our own work. It measures the quality, cost and time of every combination of model, agent and reasoning effort.

7
configurations compared
model, agent and effort
96
runs
78 valid, each one isolated
520
pairwise judgments
over 118 comparisons
3
types of work
one task each in this round

What we compare

A configuration is made of three choices.

The model
Claude Opus 5.5

It is the brain. It writes, reasons and decides what to do.

The agent
Claude Code

It is the tool that puts the model to work: it reads files, runs commands and writes code. Practitioners call it the harness.

The effort
high

It sets how much the model reasons before it answers. More effort costs more and takes longer. Technically: reasoning effort.

The seven configurations of the first round

The problem

Public leaderboards are a poor guide for anyone choosing a model for their own work.

  1. They measure purpose-built tasks.

    The tasks are short and have one right answer. Real work is long and open-ended, and it is done with tools.

  2. They stop separating anything.

    When every frontier model scores above 90%, the ranking no longer tells them apart.

  3. They ignore cost, time and the agent.

    The same model costs and performs differently depending on the agent it runs in and the effort it is asked for.

  4. They are exposed to contamination.

    A public task can end up in the data that models are trained on. From then on it no longer measures anything.

Why we don't publish a leaderboard

A single leaderboard averages over a set of tasks. Whoever builds the leaderboard picks that set, and changing the set changes the winner.

Model A stronger on short work67
Model B stronger on long work59

Top of the leaderboard: Model A

An illustrative example. A and B are hypothetical models. average over the set

Choosing a model takes three more things: the type of work, the cost and time you accept, and the agent and effort you use. That is why we publish results one type of work at a time.

First-round results

The usage card, version 0

This is what our method produces: what to use for each type of work, at what quality, cost and time.

The first round has one task per type of work and three runs per configuration. The directions are clear, the distances are not: the intervals are about ±400 points. The second round brings at least three tasks per area.

How far to trust each result

  • solidThe comparison is within the same family or with the Chinese model. On these comparisons the two judges agree.
  • provisionalThe comparison is between Claude and GPT. A tie-break from a third lab decided it, and it is waiting for a person to arbitrate.
  • referenceIt is Claude Opus 5.5 in Claude Code at high effort, and it scores 0 points. A configuration at +100 points beats it in 64% of comparisons.

Is it better? What does it cost, and how long does it take?

Quality and cost by type of work

An analysis based on a meeting with a client

Top left: more quality for less money.

The bars show the 95% interval. They are dashed when the result is provisional.

Costs are at API list price, and we recompute them from the tokens. For agents used on a subscription, it is the pay-per-use equivalent.

Claude Code Codex OpenCode

Analysis · Revision · Long work: all the numbers
An analysis based on a meeting with a client: all the numbers
ConfigurationPoints [95% CI]P(better)Won · tied · lostDecided byCost per runMinutesSolidity
Claude Opus 5.5 · Claude Code0———$1.10 [$0.94–$1.30]5.6 min [4.6–6.7]reference
Claude Sonnet 5.5 · Claude Code−106 [−441, +211]25%0 · 0 · 2agreement 2$0.56 [$0.51–$0.61]4.2 min [4.1–4.9]solid
Claude Fable 5.1 · Claude Code−319 [−919, +143]9%0 · 0 · 3agreement 3$3.19 [$2.93–$3.49]10.9 min [8.8–11.1]solid
GPT-6.1 Sol · Codex+265 [−30, +600]96%2 · 0 · 0tie-break 2$0.30 [$0.28–$0.30]4.9 min [4.8–5.2]provisional
GPT-6.1 Sol · OpenCode+305 [−42, +708]96%2 · 0 · 0tie-break 2$0.33 [$0.32–$0.41]7.9 min [6.7–9.2]provisional
GPT-6 Astra · Codex+170 [−138, +509]86%2 · 1 · 0tie-break 3$1.35 [$1.32–$1.57]5.1 min [4.3–5.6]provisional
GLM-5.3 · OpenCode−358 [−818, +21]3%0 · 0 · 2agreement 2$0.22 [$0.20–$0.68]3.7 min [3.5–22.8]solid
A content revision on a repository: all the numbers
ConfigurationPoints [95% CI]P(better)Won · tied · lostDecided byCost per runMinutesSolidity
Claude Opus 5.5 · Claude Code0———$1.34 [$1.22–$1.59]6.2 min [5.5–6.3]reference
Claude Sonnet 5.5 · Claude Code−161 [−543, +191]19%0 · 1 · 2agreement 2, external judge 1$0.64 [$0.57–$0.73]3.9 min [3.7–4.8]solid
Claude Fable 5.1 · Claude Code−89 [−454, +263]31%1 · 0 · 2agreement 2, external judge 1$2.24 [$1.98–$2.41]7.0 min [5.0–8.3]solid
GPT-6.1 Sol · Codex+286 [−77, +753]94%3 · 0 · 0tie-break 3$0.56 [$0.51–$0.58]6.8 min [5.8–7.5]provisional
GPT-6.1 Sol · OpenCode+322 [−139, +920]91%3 · 0 · 0agreement 1, tie-break 2$1.01 [$0.86–$1.04]9.9 min [9.2–12.3]provisional
GPT-6 Astra · Codex+370 [−15, +869]97%3 · 0 · 0agreement 1, tie-break 2$3.38 [$3.34–$4.22]6.4 min [6.3–6.6]provisional
GLM-5.3 · OpenCode−256 [−701, +127]10%0 · 0 · 3agreement 3$0.48 [$0.38–$1.31]5.3 min [4.2–18.5]solid
A long development job on a large repository: all the numbers
ConfigurationPoints [95% CI]P(better)Won · tied · lostDecided byCost per runMinutesSolidity
Claude Opus 5.5 · Claude Code0———$23.79 [$23.07–$26.92]36.7 min [35.1–49.6]reference
Claude Sonnet 5.5 · Claude Code+6 [−336, +347]51%1 · 1 · 1agreement 3$17.38 [$17.18–$34.70]42.5 min [30.2–55.2]solid
Claude Fable 5.1 · Claude Code−319 [−919, +143]9%0 · 0 · 3agreement 3$25.10 [$18.91–$29.59]36.3 min [32.6–43.1]solid
GPT-6.1 Sol · Codex−143 [−469, +157]18%0 · 1 · 2tie-break 3$3.94 [$3.46–$4.11]37.4 min [35.2–39.5]provisional
GPT-6.1 Sol · OpenCode−272 [−741, +121]9%0 · 0 · 3agreement 2, tie-break 1$4.16 [$3.10–$4.82]36.8 min [30.3–47.4]provisional
GPT-6 Astra · Codex−224 [−578, +96]9%0 · 1 · 2tie-break 3$19.15 [$16.64–$22.69]20.6 min [17.1–23.9]provisional
GLM-5.3 · OpenCode−272 [−741, +121]9%0 · 0 · 3agreement 3$6.76 [$4.52–$10.11]28.9 min [24.2–50.2]solid

What to use

If you work with Claude

solid

Use Claude Opus 5.5 at high effort. On short work Sonnet 5.5 does worse: P(better) is between 19% and 25%. On long work it matches Opus at 73% of the cost. Fable 5.1 did no better on any type of work: it won 1 comparison and lost 8. It costs 1.1 to 2.9 times as much.

If you work with GPT

solid

Use GPT-6.1 Sol in Codex at high effort. GPT-6 Astra did not clearly do better: it won 3 comparisons, tied 2 and lost 4. It costs 4.5 to 6 times as much. Sol in OpenCode has the same quality as in Codex. It costs more on every type of work and takes longer on short work.

The Chinese model

solid

GLM-5.3 lost all 14 comparisons against the frontier configurations. The judges always agreed. It is the cheapest on short work, but on long work it costs more than Sol. On 7 October it was the top Chinese model on Terminal-Bench 4.0.

Claude or GPT?

provisional

On short work the GPT configurations are ahead: the probability that they do better than Opus lies between 86% and 97%. Sol in Codex costs less than half as much as Opus. On long work Opus is ahead: the probability that it does better lies between 82% and 91%, but Sol costs a sixth as much.

This result is provisional. A tie-break from a third lab decided these pairs, and the two main judges agreed in only 58% of cases. In the checks on the second round's tasks, with other judges and other tasks, the same direction is not confirmed. A person's arbitration and the second round will settle it.

Is more effort worth it?

On the type of work where we varied it, Opus at maximum effort cost 5.1 times as much and took 6.3 times as long. The two judges did not see the gain in quality the same way. Lowering it to medium saves 31% but loses all three comparisons. For Sol, neither medium nor xhigh improves on high.

What it means for you

Use high effort as the default setting. Keep the maximum for cases where a slightly better result is worth five times the cost.
Go deeper

Opus at max against high: +89 points [−263, +454], P(better) 69%. HAL, across 21,730 agent runs, also finds that more reasoning effort gives equal or lower accuracy in 21 of 36 combinations of model, agent and benchmark. The test covers few models and has no significance tests (Kapoor et al., ICLR 2026).

How much does the agent matter?

We ran the same model, GPT-6.1 Sol, in two agents. Quality is indistinguishable; cost and time are not. Codex costs the same or less and is faster on short work.

Go deeper

Sol in Codex against Sol in OpenCode: −41, −39 and +131 points across the three types of work. The intervals range from ±350 to ±650.

Analysis

−41 [−406, +317]

Revision

−39 [−687, +600]

Long work

+131 [−365, +663]

The bars show the points of Sol in Codex against Sol in OpenCode, with the 95% interval. Every interval crosses zero: the difference is not visible.

The comparison between Claude and GPT will go through a person's arbitration. With the second round, the usage card will move to one card per area.

The method

How we measure

Every round follows the same six steps. We write them down before anything runs.

  1. The rules are written first.

    Before anything runs we fix what is measured, which pairs are judged and who decides when the judges disagree. A rule written after seeing the results ends up picking the winner instead of measuring it.

  2. The tasks come from real work.

    They are real requests from people at KVA. We rebuild them with the repository and the documents of that moment. A task enters the bench only after it passes a check. The tasks stay private. Publishing them would make them useless.

  3. The chain is tested on a dummy task.

    Before the round, every configuration runs on a synthetic task. We check that the effort applied is the effort requested and that the cost we compute is the real one. That is how we found a privacy setting that never reached its destination.

  4. Every run is isolated and repeated three times.

    Each run happens in its own environment, and the network reaches only the model's servers. Two runs of the same configuration give different results, so we do three. Only infrastructure errors are rerun, never the agent's.

  5. Cost and time are read at the source.

    Cost is recomputed from the tokens consumed and the day's list prices. It differs from the provider's invoice by 0.5%.

  6. Quality is measured in two ways.

    Automatic checks say whether the work functions: the build passes, the tests pass, the application does what it should. Pairwise comparisons say which piece of work is better. Every result comes with its uncertainty interval, and a tie is a result.

Go deeper

Points come from a Bayesian Bradley-Terry model. It is the idea behind chess Elo ratings: +100 points means winning 64% of comparisons. The reference is Claude Opus 5.5 in Claude Code at high effort, and it scores 0. For each configuration we report the 95% interval and P(better). P(better) is the probability of being better than the reference on that type of work.

The judge sees the task, the starting material, a quality brief written in advance and the two deliveries. The deliveries are labelled 1 and 2. It never sees the expected solution or the author's name. It reads up to 600,000 characters per side. Before judging, we measure what actually reaches it.

The judge answers in a fixed format: its preference, its preference on each dimension of the brief, and the errors in each delivery with the passage that proves them.

The strongest signal against judge bias is the concessions. These are the dimensions on which the judge from the losing family sides with the other family.

The noise between two runs of the same configuration is large. Almost all of it comes from the model doing the work, not from the judge.

Each agent loads its own instruction file: Claude Code reads CLAUDE.md, Codex reads AGENTS.md. On the bench the repository's instructions sit under both names.

The first round also included a small model of the previous generation. It served as the bottom of the scale. In the second round Claude Haiku 5.5 takes its place. It was released on 7 October.

What we did not expect

AI judges prefer their own family.

To compare two deliveries we use two AI judges from different labs: Claude Opus 5.5 from Anthropic and GPT-6.1 Sol from OpenAI. They read the two deliveries without knowing who wrote them, and they read them in both orders.

  1. Step 1 of 5

    58%

    agreement between Claude and GPT

    On pairs from the same family the two judges agree in 88% of cases. With the Chinese model they always agree. When one delivery comes from Claude and the other from GPT, agreement drops to 58%.

  2. Step 2 of 5

    14 of 14

    disagreements in favour of their own family

    In the 14 clear disagreements between Claude and GPT, each judge picked its own family's delivery. It never happened the other way round.

  3. Step 3 of 5

    1 of 24

    pairs they agree on when the instructions are identical

    We gave the two agents exactly the same instructions and ran the test again on 24 new pairs. The two judges agreed on a single pair. The instructions are not the cause.

  4. Step 4 of 5

    20 of 20

    times each one recognises the author

    The judge never gets the author's name. When we asked them to guess it, each of the two recognised the right family 20 times out of 20.

  5. Step 5 of 5

    11–12 of 12

    even with deliveries rewritten into a fixed template

    We rewrote the deliveries into a fixed template to strip out the style. The judges recognised them anyway. Recognition goes through substance: what one chooses to say, how much one dwells on risks, how firm the recommendation is.

Why it happens

  • Each family of models finds the text of its relatives more natural. Text that reads naturally also looks better.
  • Each family tackles a task in its own way. A rewrite changes the words but keeps the choices of substance. In our test those choices were enough to recognise the author.
  • The two judges do not disagree about facts. They disagree about what makes a delivery better: how much to say, how much to dwell on risks, how firm to be.
  • That is why every judge's instructions now say what the person who asked for the work prefers. With equal substance and equal honesty about the checks, they prefer the simplest delivery that is ready to use. Risks and caveats count only when they change what gets decided or done.

What we did about it

  1. Where an automatic check measures what matters most, the check decides. No judge can override it.
  2. Otherwise two judges from two other labs decide, DeepSeek V4 Pro and Muse Spark 1.3. Their vote counts only if they agree. When both vote clearly, they agree 18 times out of 18.DeepSeek V4 ProMuse Spark 1.3
  3. A person arbitrates a sample of pairs blind. The two judges' vote enters the results only if it matches the person's at least 7 times out of 10.

These rules apply from the second round. In the first, the pairs between Claude and GPT were decided by the tie-break of a single judge from a third lab. That is why those results remain provisional.

Agreeing is not the same as being right. That is why the automatic checks and the person count too.

What it means for you

If you use one AI to judge the work of another:

  • never use a judge from the same family as one of the contestants;
  • have every pair read in both orders, and count the vote only if it does not change;
  • check what the judge actually sees;
  • keep a person on a sample of pairs.
Go deeper

Panel agreement on pairs with two clear votes, with 95% Wilson intervals: 0.80 [0.71–0.87] over 95 pairs; 0.88 [0.74–0.95] within the same family; 1.00 [0.84–1.00] with the Chinese model; 0.58 [0.41–0.73] between Claude and GPT. On the Claude–GPT disagreements the sign test gives p = 0.0001.

In the recognition test Opus guessed the family 20 times out of 20, and so did Sol. We then had two different models rewrite the deliveries into the fixed template. Opus recognised them 12 times out of 12 in both versions, Sol 12 of 12 in one and 11 of 12 in the other. The threshold written before the test was 10 of 12.

Recent research points the same way. For each study we also state its limit.

  • Across four open-weight model families and 9,312 judgments, judges give their own family 3.4 to 8.4 percentage points more support. Once you account for how familiar the text sounds to the judge, the effect shrinks by 61%. The study is a preprint and uses open-weight models of up to 72 billion parameters. Awuni et al., arXiv:2609.17857, preprint
  • A judge favours models trained on data from a model related to it: the same model, one derived from it or one from the same family. It happens in most of the pairs studied. The study uses three judge models and two benchmarks. Li et al., ICLR 2026
  • The counterpoint: once the judge's own quality is taken into account, only 51% of the published self-preference cases remain significant. The study mostly uses open-weight judges. That is why we add automatic checks, judges from other families and a person. Roytburg et al., ICML 2026
  • A classifier tells which of five chat assistants wrote an answer in 97.1% of cases. Retrained on answers paraphrased by another model, it still reaches 91.4%. Sun et al., ICML 2025

What we learned

The other lessons

They do not apply only to us: they apply to anyone who evaluates AI models. They come from the first round and from the checks on the second round's tasks.

16% · 5%

pairs where Opus and Sol change their verdict

A judge can change its mind when the order is swapped.

As a judge, Opus changed its verdict on 16% of pairs just because the deliveries were shown in reverse order. Sol did so on 5%. It happens almost only on the closest pairs.

Why. When two deliveries are close, a detail such as their position is enough to tip the verdict. Sol judged with more reasoning effort than Opus, and part of the gap may come from that.

What it means for you

Read every pair in both orders, and count the vote only if it stays the same.
Go deeper

Opus changed its verdict on 19 pairs out of 118, Sol on 6. In a panel of open-weight models, swapping the order changes the winner in 55.4% of pairs (Awuni et al., 2026 preprint). Across 15 judges and over 150,000 judgments, the order effect is strongest when the two answers are close in quality. The length of the prompt matters little (Shi et al., IJCNLP-AACL 2025).

1 in 5

verdicts that changed once we fixed what the judge read

A judge only sees what it is shown.

Long deliveries exceed what a judge can read. On one task the judge was grading the agent's report instead of its code. On another, the agents left up to 200 working files behind, and the requested document fell outside the limit in 3 deliveries out of 4.

Why. The tool decides what fits within the reading limit, not the judge. The verdicts looked reasonable: we found out by measuring what reached the judge.

What it means for you

Before you trust a judge, measure what it reads. State which files make up the delivery.
Go deeper

Fixing the view of the long task changed 10 votes out of 48. Our paper on judges that cannot see the source shows the same problem from another angle.

Read the paper: Summary Is Not Enough

5 of 12

tasks with an error in the automatic check

Automatic checks get things wrong too, and not at random.

Checking the twelve tasks of the second round, we found an error in the automatic check of five. In one, the check looked for a value in the wrong table and counted 24 correct citations out of 30 as errors.

Why. Whoever writes a check imagines one way of doing the work. Each error penalised a way of working, not wrong work.

What it means for you

Before you trust a check, read by hand the errors it flags on a few deliveries from different models.
Go deeper

Public benchmarks have the same problem. A study of ten agentic benchmarks found seven with flaws in the validity of their outcomes: on SWE-bench Verified, for example, a wrong patch can pass the tests (Zhu et al., NeurIPS 2025). When we correct a check, it becomes a new version of the task. The version records the date and the reason.

7 of 8

repositories with instructions for one agent only

Agents don't always do what you think.

Seven repositories out of eight had only Claude Code's instruction file, so the project's rules reached only one agent. One agent sent the task text to a third model just to write the session title.

Why. Each agent loads its own files and settings, and some of its choices stay invisible to the person using it.

What it means for you

Check what an agent actually sends, not what it says it sends. Give every agent the same instructions.
Go deeper

We found out by intercepting the requests. With the same instructions under both CLAUDE.md and AGENTS.md we reran the comparisons, and the direction of the verdict from judges of other labs did not change.

$0.38 → $3.94

the cost of one run, before and after counting the sub-agents

Numbers have to be read at the source.

An agent that launches sub-agents appeared to cost a tenth of the real figure. Only one of them was being counted. A model reached through an intermediary cost 3.6 times its list price. An output limit set too low left 3 judgments out of 7 empty. We paid for them all the same.

Why. Every tool records cost and time in its own way, and often records only part of them.

What it means for you

Read time and cost from the source closest to the model, and check them against the invoice.
Go deeper

The cost we recompute differs from the invoice by 0.5%: $298.31 against $299.85. A run that finished in five minutes was recorded as closing after an hour. A wait had hung.

+79% · +63%

extra cost and time, same model in another agent

The same model costs differently depending on the agent.

With the same model, GPT-6.1 Sol, one agent cost 79% more than another on one task and took 63% more time on another. Quality is indistinguishable.

Why. The agent decides how much context to resend to the model, how many tools to call and whether to use the cache. Those choices change the tokens consumed.

What it means for you

Measure cost with the agent you will actually use. The model's price per token is not enough.
Go deeper

Recent studies confirm it. With the same model in three open-source agents, tokens per solved task vary by up to 40 times. The share of tasks solved moves by only 0–8 points (Vats and Golev, ICML 2026 workshop, two models and 50 tasks). HAL also finds that the agent changes both accuracy and cost (Kapoor et al., ICLR 2026).

Three comparisons are not proof.

Winning 3 comparisons out of 3 looks decisive. With three comparisons you can see who is ahead, not by how much: the uncertainty runs from “about even” to “clearly better”. Telling apart a modest difference takes about 46 comparisons per type of work.

How many comparisons does it take?

modest difference (64%) · clear difference (76%)

With 3 comparisons the uncertainty is ±393 points.

Between two configurations that are actually even, the share of wins could be anywhere between 9% and 91%.

With this many comparisons, a modest difference is not visible yet.

Why. Each comparison is like tossing a slightly loaded coin. To find out how loaded it is, you need many tosses.

What it means for you

Don't trust a winner without its interval. Repeat the runs and count the comparisons.
Go deeper

We simulated hundreds of rounds with a known truth. Up to a difference of about 85% of wins, the 95% interval contains the true value 95–99% of the time. Beyond that, the direction stays right but the distance is underestimated. Recent research shows it too. On HotpotQA, over 30 repeated runs, the best reasoning strategy beats the runner-up in only 77% of them (Potamitis et al., ReasonBENCH). A benchmark with more noise relative to its signal leads to less reliable decisions (Heineman et al., NeurIPS 2025).

You need automatic checks, comparisons and a person.

The checks say whether the work is usable. The comparisons say which is better. A person looks at the cases where the judges disagree.

Why. Sometimes the check no longer separates anyone: on one task every run fixed every defect measured. Sometimes it is the comparison that misses something: on another, the judges did not see a clear advantage the check had measured.

What it means for you

Where a check measures what matters most, the judges' vote does not replace it.
Go deeper

A measure that is the same for everyone is saturated: we flag it and leave it out of any average.

Use them today

Test models on your own work: eight rules

We turned our bench's lessons into rules. They hold even if you don't use our method.

  1. Use tasks from your own work, and keep them private.

    A public task ends up in the training data and stops measuring anything.

    From the lesson: Automatic checks get things wrong too, and not at random.
  2. Measure cost and time with your agent.

    The price per token does not tell you what the work will cost.

    From the lesson: The same model costs differently depending on the agent.
  3. Run each configuration at least three times.

    A single run cannot tell a better model from a lucky run.

    From the lesson: Three comparisons are not proof.
  4. Compare in pairs and read each pair in both orders.

    A vote counts only if it stays the same when the order is swapped.

    From the lesson: A judge can change its mind when the order is swapped.
  5. Use judges from another family, and a person on a sample.

    A judge prefers its own family, and recognises it.

    From the lesson: AI judges prefer their own family.
  6. Check what the judge sees.

    Long deliveries exceed its reading limit.

    From the lesson: A judge only sees what it is shown.
  7. Don't use maximum effort as a fixed setting.

    It costs a lot, slows everything down and has not shown that it pays off.

    From the lesson: Three comparisons are not proof.
  8. Give every number with its interval, and measure again at every release.

    A number without an interval promises a precision it does not have.

    From the lesson: Three comparisons are not proof.

The second round

Built to decide

  • Five work areas: analysis and discovery, changes to an existing product, interfaces, long work from scratch, research with sources. An area is published only with at least three tasks.
  • More than half of the tasks require real tools: builds and tests, an application running with its database, the browser, web search.
  • The latest models join. The first is Claude Haiku 5.5. Anthropic released it on 7 October.
  • Between labs, the automatic check decides first, then two judges from other labs who must agree, then a person on a sample.

How we decide whether to switch models

For each area we compare the new model with the one in use. The rules are written in advance, and they give one of six verdicts.

Adopt
It is almost certainly better (P(better) at least 95%), costs at most one and a half times as much, and does not collapse on any task in the area.
Better but costlier
It is better, but costs more than one and a half times as much. A person decides.
Better, but one task collapses
It is better, but much worse on one task in the area. A person looks at that task and decides.
Do not adopt
It is almost certainly worse (P(better) at most 5%).
Adopt for cost
It is not worse beyond a margin fixed in advance, and it costs or takes at least 30% less. The margin is 150 points. On long work a mistake costs more, and the margin drops to 100.
Inconclusive
In every other case we stay with the model in use.

Try the rule

Verdict

Adopt for cost

It is not worse beyond a margin fixed in advance, and it costs or takes at least 30% less. The margin is 150 points. On long work a mistake costs more, and the margin drops to 100.

Go deeper

We measured the rule's errors by simulating rounds with three tasks per area. False alarms run between 4% and 6%. A true advantage of 200 points is adopted in 73% of rounds, and a disadvantage of 200 points is rejected in 67%. A fourth task per area is the first improvement planned.

What we will publish

For each area we will publish a recommendation with its level of solidity, the quality and cost chart and the full table. The chart shows how a result will read.

First edition after the second round

wins 70% of comparisons, between 52% and 83%

For your company

We apply the same method to your company's tasks

The bench works with the real tasks of any company. We take them from its documents, its repositories and its processes. It answers practical questions:

  • which model and which agent to use for that type of work;
  • what each run costs and how long it takes;
  • when a cheaper model is good enough and when it is not;
  • when it is worth asking for more reasoning.

At every major release, the new model is measured on the same tasks and with the same rules. First a few short tasks show whether it is in the top tier, then the decision tasks give a verdict area by area.

Frequently asked questions

Why don't you publish the tasks?

A published task ends up in the data models are trained on and stops measuring anything. We publish the method, the lessons and the results by type of work.

Which model is the best?

It depends on the type of work, the cost you accept and the agent you use it with. That is why we publish a usage card by type of work instead of a leaderboard.

Who judges the quality?

When an automatic check measures what matters most, the check decides. Otherwise AI judges from different labs decide. They read every pair in both orders. A person arbitrates a sample of pairs blind.

Can you apply the method to my company?

Yes. We build the bench on the company's real tasks, and at every major release we measure the new model against the one in use.

When are the next results out?

They come with the second round. It will have at least three tasks per area. Each edition says what changed from the previous one.

Notes

The first round ran on 5 and 6 October 2026. It has one task per type of work and three runs per configuration: 96 runs, 78 of them valid, 118 pairwise comparisons and 520 judgments. The tasks and their detailed results stay private.

The Terminal-Bench 4.0 figure comes from the vals.ai leaderboard updated on 7 October 2026. Public leaderboards change every month.

The names and marks of models, agents and labs belong to their owners and are used here only to identify them. KVA is not affiliated with any of them, and none of them commissioned or approved these measurements.

This first edition is dated 9 October 2026. The next one will bring the results of the second round.

Research cited

  1. Awuni et al., Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels, arXiv:2609.17857, preprint.
  2. Roytburg et al., Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations, ICML 2026.
  3. Li et al., Preference Leakage: A Contamination Problem in LLM-as-a-judge, ICLR 2026.
  4. Sun et al., Idiosyncrasies in Large Language Models, ICML 2025.
  5. Shi et al., Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge, IJCNLP-AACL 2025.
  6. Zhu et al., Establishing Best Practices in Building Rigorous Agentic Benchmarks, NeurIPS 2025 Datasets & Benchmarks.
  7. Vats & Golev, The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation, ICML 2026 workshop (Deep Learning for Code), arXiv:2607.22585.
  8. Kapoor et al., Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation, ICLR 2026.
  9. Potamitis et al., ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning, arXiv:2512.07795, v2 May 2026.
  10. Heineman et al., Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation, NeurIPS 2025.
  11. Fooladi & Bottino (KVAi Research), Summary Is Not Enough: Source-Blind LLM Judges Mistake Faithful Citation for Hallucination, under review.