Run postmortem · Published 2026-08-26

GLM-5.3 Flashmax reasoning

It reached 64% public calibration, recovered the right package, and still scored only 4% hidden.

4/ 100
hidden exact

One-hour track · max reasoning · provisional

ARI score scaleexact-match accuracy
4%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#20
Current cost$0.05
Points / $75.33
Average time55:21
01 · Verdict

Package recovery could not overcome severe public overfitting

GLM-5.3 Flash produced a valid 4% hidden exact-match result in its first published max-effort run. The agent built several compact neural candidates, raised official public calibration from 0% to 64%, briefly promoted a weaker 60% package, then restored and finalized the 64% candidate. Independent grading reproduced the hidden score from the exact stored artifact, so the 60-point public-to-hidden gap reflects poor transfer rather than a stale-package failure. One seed remains provisional.

One valid autonomous run produced a provisional 4% hidden exact-match score.

#20
Score position4% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GLM-5.3 Flash built

The final artifact was a compact byte-level transformer trained through a broad synthetic-curriculum search, with unusually careful late-stage package selection.

01

Synthetic task curriculum

The run generated varied examples for copying, arithmetic, mappings, records, schedules, sorting, and text transformations.

02

Compact neural candidates

The final model used a roughly 16.8-million-parameter byte transformer whose compressed package stayed just below the 16 MB cap.

03

Exact package verification

The agent rebuilt and rescored the selected bundle, restored the best official candidate after a weaker promotion, and finalized matching code and weights.

03 · Trajectory

How the run unfolded

Opening 9m

Moved from smoke test to 44%

A zero-percent safety checkpoint gave way to the first valid main package at 44% public calibration.

36m

Reached 64%, then regressed selection

Candidate m7 reached the best official public score. A later 60% package was temporarily promoted despite scoring lower.

55:21

Restored m7 and finalized

The agent rechecked the 64% package, selected it again, finished before the deadline, and received a valid 4% hidden grade.

04 · Assessment

Good, bad, and unresolved

What worked

The final artifact is auditable

The selected package passed the size and inference checks, matched its recorded sources, and reproduced 4% in an independent Docker re-grade.

What hurt

Calibration did not predict generalization

Sixteen of twenty-five public questions were correct, but only four of one hundred hidden questions matched exactly.

What remains unknown

Replication and the Ox Alpha relationship

Two more valid seeds are needed, and public model metadata does not establish whether this release is literally identical to the former anonymous Ox Alpha endpoint.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-01044%55:21$0.0515.46 MB10Valid
50.8KInput tokens
23KOutput tokens
2.9MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 10 candidate submission actions across the published run set. Selected artifacts averaged 15.46 MB compressed.

03

Time use

The runs averaged 55:21 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 80 agent messages · 10 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.05 per displayed run at current configured rates.

The official run cost an estimated $0.0531 at its launch pricing. That creates high displayed points per dollar, but the value ratio is driven by very low token prices and should not obscure the weak 4% hidden result. Preflight smoke requests are excluded.

At current configured API-equivalent rates, the displayed run cost is $0.05 and value is 75.33 score points per dollar.

#1
Value positionValue rank #1 of 25 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-26.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GLM-5.3 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard