Points come from a Bayesian Bradley-Terry model. It is the idea behind chess Elo ratings: +100 points means winning 64% of comparisons. The reference is Claude Opus 5.5 in Claude Code at high effort, and it scores 0. For each configuration we report the 95% interval and P(better). P(better) is the probability of being better than the reference on that type of work.
The judge sees the task, the starting material, a quality brief written in advance and the two deliveries. The deliveries are labelled 1 and 2. It never sees the expected solution or the author's name. It reads up to 600,000 characters per side. Before judging, we measure what actually reaches it.
The judge answers in a fixed format: its preference, its preference on each dimension of the brief, and the errors in each delivery with the passage that proves them.
The strongest signal against judge bias is the concessions. These are the dimensions on which the judge from the losing family sides with the other family.
The noise between two runs of the same configuration is large. Almost all of it comes from the model doing the work, not from the judge.
Each agent loads its own instruction file: Claude Code reads CLAUDE.md, Codex reads AGENTS.md. On the bench the repository's instructions sit under both names.
The first round also included a small model of the previous generation. It served as the bottom of the scale. In the second round Claude Haiku 5.5 takes its place. It was released on 7 October.