accepted run · R5/r5

Droid
× sol

Generation 4 of five · harness v0.186.0 · gpt-5.6-sol · high reasoning requested

score3.003.00 visible + 0.00 holdout
duration
46m15 agent turns
reported tokens
459,488input + output
cache read / input
2376.57%reported telemetry · not ranked
tool calls
145harness-reported
01 / score anatomy

Every visible stage.

Passed stages receive their declared points. Holdout contributes 30 points proportionally across 11 checks. Diagnostics are verifier output, not model self-report.

01
Append/get foundationfoundation

append failed: ledger.json does not exist

+0
02
Atomic batches and idempotencybatch

append failed: ledger.json does not exist

+0
03
Deterministic paginationpagination

append failed: ledger.json does not exist

+0
04
Legacy schema migrationmigration
+3
05
Crash-safe writesatomicity

append failed: ledger.json does not exist

+0
06
Interrupted-write recoveryrecovery

append failed: ledger.json does not exist

+0
07
Multi-process concurrencyconcurrency

concurrent append failed: 1

+0
08
Checksummed compactioncompaction

append-batch failed: ledger.json does not exist

+0
09
Export/import round triproundtrip

append failed: ledger.json does not exist

+0
10
Replay and integrityreplay

append failed: ledger.json does not exist

+0
11
Full regression auditaudit

append failed: ledger.json does not exist

+0
12
Scale and performancescale

append-batch failed: ledger.json does not exist

+0
13
Adversarial concurrent batchesstress-concurrency

stress batch failed: 1

+0
14
Fault injection and validationvalidation

Expected compact to fail

+0
15
Large integrated stress runscale-stress

scale batch failed: 1

+0
holdout0 / 11 checks passed+0.00
02 / reported telemetry

Useful, with limits.

These counters come from the harness and transport. They are preserved as observed; AgentBattler does not infer missing values or treat token/cache figures as directly comparable billing data.

input tokens
333,324
cached input tokens
7,921,664
output tokens
126,164
reasoning tokens
47,670
started
Aug 4, 2026
ended
Aug 4, 2026
03 / provenance

Which V5 produced this run?

The campaign preserves evidence from compatible protocol revisions rather than relabeling it. This record remains attached to the exact challenge and schedule hashes used during execution.

source
R5 · r5
amendment
factory droid cli harness and cliproxy route
challenge
challenge-8b563735b78d6c098b563735b78d6c097efd
schedule
schedule-d6c0ef3cd1b47e41d6c0ef3cd1b47e41e4be
run key
6ad332a1d17455e23695e385
logical identity
factory-droid|gpt-5.6-sol|4|1|1
04 / attempt history

Failures stay outside the score.

No infrastructure-invalid attempt is recorded for this logical run. Attempts are retained for reliability analysis and never averaged into task performance.

  1. attempt 1completed

    completed evidence archived before campaign selection

    46m · Aug 4, 2026
05 / evidence

Open the underlying record.

The semantic trace retains visible messages, tool calls, tool results, usage events, and stderr after credential-shaped values and host paths are sanitized. It does not expose hidden chain-of-thought.