Harness × model interaction
Claude Code with Sol led at 73.44. The same model family did not hold the same rank across harnesses, so neither a harness-only nor model-only summary explains the result.
Mini Ledger v4 · long-horizon terminal benchmark
Four coding harnesses ran the same three models through a continuous 15-turn engineering task. Five independent generations per combination. Nobody solved it perfectly—and the harness mattered.
Score combines 70 visible stage points with 30 points from 11 private holdout checks. Bars show the five-run range; markers are individual generations.
Claude Code with Sol led at 73.44. The same model family did not hold the same rank across harnesses, so neither a harness-only nor model-only summary explains the result.
Codex CLI × Terra ranged from 3.00 to 66.36—a 63.36-point swing under the same declared condition.
The fastest average condition was Claude Code × Luna at 23m. Some slower conditions scored lower; runtime is reported as evidence, not interpreted as causal.
The best individual run scored 81.82. Zero runs reached 100, leaving headroom across concurrency, recovery, compaction, validation, and integrated stress behavior.
All 60 jobs were declared before execution. Each run used the same prompt sequence, verifier hashes, high reasoning request, isolated workspace, disabled network, and one continuous harness session. A run only scores when all 15 turns complete and both verifier layers execute.
Claude Code and DotAgents reached the pinned model endpoints through CLIProxyAPI; Codex CLI and Pi used their native routes. That transport difference is part of the published harness condition. Token telemetry is retained but not ranked because harnesses report cache and context usage differently.
49 MB of downloadable semantic traces preserve all distinct messages, tool calls, results, usage events, and stderr from 25.0 GB of cumulative raw streams.
Pi and early DotAgents streams repeatedly emitted the entire growing conversation, producing 26.8 GB from 900 turns. The release removes only those cumulative streaming snapshots after retaining their final messages and tool calls. Every source trace is hashed, the transformation is open source, 225 secret-shaped fields or values were redacted, and host paths were normalized.
inspect the exporter ↗the battle lane · three models · four harnesses · 11,340 games
The original chess lane remains published alongside the terminal study: 12 harness and model combinations, immutable generated engines, and deterministic game replays.
All 60 engines used the same prompt and requested high reasoning. The leaderboard compares same-model games across 4 harnesses — Codex CLI v0.144.0 · Pi v0.80.7 · Claude Code v2.1.211 · DotAgents v1.1.6 — and shows each combo’s schedule size; independent Harbor reproduction is not claimed.
Each row is one harness and model combination, pooling its 5 independently generated engines. Scores use same-model cross-harness games so model identity stays fixed; the dots keep generation variance visible. How pooled score works →
Open the complete move log, step through every board state, and inspect both competing artifacts.
watch this battle →Isolated Codex CLI v0.144.0 · Pi v0.80.7 · Claude Code v2.1.211 · DotAgents v1.1.6 harnesses each write executable chess agents.
Known positions catch malformed output before competition.
Deterministic positions, colors, seeds, and move limits are recorded.
Source hashes, telemetry, standings, and replay traces travel together.