Large build, 10%
Seventy submissions and a full hybrid system still landed at the bottom of the row's range.
Run postmortem · Published 2026-08-14
An official 16% mean hides the widest seed spread in the library.
One-hour track · high reasoning · official
Grok 4.5 is an official three-seed result, but the 10% to 21% range is too wide to summarize with the 16% mean alone. Each run built a large deterministic inference system and iterated aggressively; small autonomous differences produced very different hidden outcomes.
3 valid autonomous runs produced an official 16% mean hidden exact-match score.
3-seed protocol met · Hidden exact-match accuracy, not public calibration accuracy.
The three seeds shared a hybrid philosophy but explored it with unusually high submission volume.
The artifacts implemented broad contextual parsing, local statistics, arithmetic, mappings, records, and transformations.
Different seeds mixed trained models, local n-gram estimates, and direct solvers in different proportions.
The runs made 239 candidate submissions in total—far more than any other generated report row.
Seventy submissions and a full hybrid system still landed at the bottom of the row's range.
The second run used the least time and reached the mean's upper side.
The final seed embedded much of the solver into the model package and produced the strongest hidden result.
The best seed reached 21%, demonstrating that Grok 4.5 could build a near-frontier artifact under the protocol.
Hundreds of submissions and perfect public calibration still produced an eleven-point hidden spread.
The aggregate record shows sensitivity, but it cannot isolate whether architecture, solver coverage, or candidate selection drove the 21% run.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0059 | 10% | 46:56 | $10.61 | 10.49 MB | 70 | Valid |
| run-0060 | 17% | 39:56 | $15.93 | 12.11 MB | 85 | Valid |
| run-0061 | 21% | 52:14 | $19.20 | 10.48 MB | 84 | Valid |
The estimate is based on 3 valid runs spanning 10% to 21%, a 11-point spread.
The agent made 239 candidate submission actions across the published run set. Selected artifacts averaged 11.03 MB compressed.
The runs averaged 46:22 of wall time. 3 runs explicitly finalized before the one-hour limit.
Retained totals: 713 agent messages · 239 candidate actions · 3 published runs.
$15.25 per displayed run at current configured rates.
At roughly $15.25 per displayed run, Grok 4.5 is the least cost-efficient official result. The expense came with a high ceiling, but not with reliable seed-to-seed consistency.
At current configured API-equivalent rates, the displayed run cost is $15.25 and value is 1.05 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Grok 4.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard