Run postmortem · Published 2026-08-14

Grok 4.5high reasoning

An official 16% mean hides the widest seed spread in the library.

16/ 100
hidden exact

One-hour track · high reasoning · official

ARI score scaleexact-match accuracy
16%
0255075100
StatusOfficial
Valid seeds3 / 3
Score rank#8
Current cost$15.25
Points / $1.05
Average time46:22
01 · Verdict

Capable, expensive, and highly seed-sensitive

Grok 4.5 is an official three-seed result, but the 10% to 21% range is too wide to summarize with the 16% mean alone. Each run built a large deterministic inference system and iterated aggressively; small autonomous differences produced very different hidden outcomes.

3 valid autonomous runs produced an official 16% mean hidden exact-match score.

#8
Score position16% on the common 0–100 scale. Tied valid scores share a rank.

3-seed protocol met · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Grok 4.5 built

The three seeds shared a hybrid philosophy but explored it with unusually high submission volume.

01

Large symbolic solver surface

The artifacts implemented broad contextual parsing, local statistics, arithmetic, mappings, records, and transformations.

02

Neural and n-gram fallbacks

Different seeds mixed trained models, local n-gram estimates, and direct solvers in different proportions.

03

Relentless candidate iteration

The runs made 239 candidate submissions in total—far more than any other generated report row.

03 · Trajectory

How the run unfolded

Seed 1

Large build, 10%

Seventy submissions and a full hybrid system still landed at the bottom of the row's range.

Seed 2

Forty minutes, 17%

The second run used the least time and reached the mean's upper side.

Seed 3

Best build, 21%

The final seed embedded much of the solver into the model package and produced the strongest hidden result.

04 · Assessment

Good, bad, and unresolved

What worked

A high ceiling

The best seed reached 21%, demonstrating that Grok 4.5 could build a near-frontier artifact under the protocol.

What hurt

Iteration volume did not create stability

Hundreds of submissions and perfect public calibration still produced an eleven-point hidden spread.

What remains unknown

Which seed behavior was causal

The aggregate record shows sensitivity, but it cannot isolate whether architecture, solver coverage, or candidate selection drove the 21% run.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-005910%46:56$10.6110.49 MB70Valid
run-006017%39:56$15.9312.11 MB85Valid
run-006121%52:14$19.2010.48 MB84Valid
726.5KInput tokens
307.8KOutput tokens
84.9MCache-read tokens
01

Seeded evidence

The estimate is based on 3 valid runs spanning 10% to 21%, a 11-point spread.

02

Candidate behavior

The agent made 239 candidate submission actions across the published run set. Selected artifacts averaged 11.03 MB compressed.

03

Time use

The runs averaged 46:22 of wall time. 3 runs explicitly finalized before the one-hour limit.

Retained totals: 713 agent messages · 239 candidate actions · 3 published runs.

06 · Economics

What the result cost

$15.25 per displayed run at current configured rates.

At roughly $15.25 per displayed run, Grok 4.5 is the least cost-efficient official result. The expense came with a high ceiling, but not with reliable seed-to-seed consistency.

At current configured API-equivalent rates, the displayed run cost is $15.25 and value is 1.05 score points per dollar.

#17
Value positionValue rank #17 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Grok 4.5 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard