cost · context · latency
Every combo,
measured.
A normalized view of every harness × model combination. Compare what a single model-generation turn costs in tokens and time, then open the row for the five underlying generated engines.
Every harness × model combination shown here declares high reasoning effort. what this setting means →
tokens + duration available
Pi / GPT-5.6 Terra
Pi / GPT-5.6 Luna
billing price is not in this snapshot
comparison board
Harness × model combinations
| harness × modelOne harness, harness version, model, reasoning setting, and generation configuration. | pooled score / W–D–LPoints earned across all published chess games for the five generated engines in this combination. | price / generation turnPublished billing cost divided by model-generation turns; never inferred when billing data is absent. | tokens / generation turnPublished generation tokens divided by model-generation turns, not chess moves. | time / generation turnGeneration wall time divided by model-generation turns, not time spent choosing a chess move. | telemetry coverageHow many generated engines publish both token and duration telemetry. |
|---|---|---|---|---|---|
65.9%701–970–129 · 360 games / agent | —not published | —not published | —not published | not publishedno generation-turn data | |
61.8%500–1226–74 · 360 games / agent | —not published | 9,098tokens / generation turn | 29.6 swall time / generation turn | 5/5 engines8.0 generation turns / engine | |
56.4%370–1290–140 · 360 games / agent | —not published | 149,721tokens / generation turn | 252.1 swall time / generation turn | 5/5 engines1.0 generation turns / engine | |
54.7%39–119–22 · 36 games / agent | —not published | —not published | —not published | not publishedno generation-turn data | |
51.1%8–168–4 · 36 games / agent | —not published | —not published | —not published | not publishedno generation-turn data | |
50.2%183–1441–176 · 360 games / agent | —not published | 9,332tokens / generation turn | 18.5 swall time / generation turn | 5/5 engines11.4 generation turns / engine | |
46.9%1–167–12 · 36 games / agent | —not published | —not published | —not published | not publishedno generation-turn data | |
44.1%45–1496–259 · 360 games / agent | —not published | 7,011tokens / generation turn | 22.6 swall time / generation turn | 5/5 engines6.4 generation turns / engine | |
43.1%20–1513–267 · 360 games / agent | —not published | —not published | —not published | not publishedno generation-turn data | |
43.1%15–1522–263 · 360 games / agent | —not published | —not published | —not published | not publishedno generation-turn data | |
43%30–1487–283 · 360 games / agent | —not published | 155,053tokens / generation turn | 193.7 swall time / generation turn | 5/5 engines1.0 generation turns / engine | |
42.4%47–1433–320 · 360 games / agent | —not published | 262,911tokens / generation turn | 260.4 swall time / generation turn | 5/5 engines1.0 generation turns / engine |
Price is unavailable when the snapshot does not contain billing telemetry; it is never inferred from tokens or subscription usage. Definitions: generation turn · pooled score · telemetry coverage.
selected combination · Codex CLI v0.144.0
GPT-5.6 Sol
gpt-5.6-sol · reasoning effort: high · 360 chess games per generated engine · pooled score 56.4%
Telemetry is published per generated engine. Combination averages are weighted by observed generation turns. how telemetry coverage works →