benchmark telemetry / 12 combos / reasoning: high

cost · context · latency

Every combo,
measured.

A normalized view of every harness × model combination. Compare what a single model-generation turn costs in tokens and time, then open the row for the five underlying generated engines.

shared generation conditionReasoning effort
high

Every harness × model combination shown here declares high reasoning effort. what this setting means →

combos with telemetry6 / 12

tokens + duration available

leanest token profile7,011

Pi / GPT-5.6 Terra

fastest generation turn18.5 s

Pi / GPT-5.6 Luna

price coverage0%

billing price is not in this snapshot

comparison board

Harness × model combinations

12 rows · sorted by score
harness × modelOne harness, harness version, model, reasoning setting, and generation configuration.pooled score / W–D–LPoints earned across all published chess games for the five generated engines in this combination.price / generation turnPublished billing cost divided by model-generation turns; never inferred when billing data is absent.tokens / generation turnPublished generation tokens divided by model-generation turns, not chess moves.time / generation turnGeneration wall time divided by model-generation turns, not time spent choosing a chess move.telemetry coverageHow many generated engines publish both token and duration telemetry.
65.9%701970129 · 360 games / agent
not published
not published
not published
not publishedno generation-turn data
61.8%500122674 · 360 games / agent
not published
9,098tokens / generation turn
29.6 swall time / generation turn
5/5 engines8.0 generation turns / engine
56.4%3701290140 · 360 games / agent
not published
149,721tokens / generation turn
252.1 swall time / generation turn
5/5 engines1.0 generation turns / engine
54.7%3911922 · 36 games / agent
not published
not published
not published
not publishedno generation-turn data
51.1%81684 · 36 games / agent
not published
not published
not published
not publishedno generation-turn data
50.2%1831441176 · 360 games / agent
not published
9,332tokens / generation turn
18.5 swall time / generation turn
5/5 engines11.4 generation turns / engine
46.9%116712 · 36 games / agent
not published
not published
not published
not publishedno generation-turn data
44.1%451496259 · 360 games / agent
not published
7,011tokens / generation turn
22.6 swall time / generation turn
5/5 engines6.4 generation turns / engine
43.1%201513267 · 360 games / agent
not published
not published
not published
not publishedno generation-turn data
43.1%151522263 · 360 games / agent
not published
not published
not published
not publishedno generation-turn data
43%301487283 · 360 games / agent
not published
155,053tokens / generation turn
193.7 swall time / generation turn
5/5 engines1.0 generation turns / engine
42.4%471433320 · 360 games / agent
not published
262,911tokens / generation turn
260.4 swall time / generation turn
5/5 engines1.0 generation turns / engine

Price is unavailable when the snapshot does not contain billing telemetry; it is never inferred from tokens or subscription usage. Definitions: generation turn · pooled score · telemetry coverage.

selected combination · Codex CLI v0.144.0

GPT-5.6 Sol

gpt-5.6-sol · reasoning effort: high · 360 chess games per generated engine · pooled score 56.4%

avg / generation turn252.1 s149,721 tokens
total generation tokens748,606
total generation time1260.6 s
avg tool calls / generation turn7.4
chess record · W–D–L3701290140
generated enginepooled score / W–D–Lprice / generation turntokens / generation turntime / generation turngeneration turns
GPT-5.6 Sol #1sol-01 · rank 1159%8824923104,672206.2 s1
GPT-5.6 Sol #2sol-02 · rank 959.9%8725716286,680323.0 s1
GPT-5.6 Sol #3sol-03 · rank 2942.6%2026773110,014200.6 s1
GPT-5.6 Sol #4sol-04 · rank 860.1%8725914106,942291.6 s1
GPT-5.6 Sol #5sol-05 · rank 660.3%8825814140,298239.1 s1

Telemetry is published per generated engine. Combination averages are weighted by observed generation turns. how telemetry coverage works →