Moved from smoke test to 44%
A zero-percent safety checkpoint gave way to the first valid main package at 44% public calibration.
Run postmortem · Published 2026-08-26
It reached 64% public calibration, recovered the right package, and still scored only 4% hidden.
One-hour track · max reasoning · provisional
GLM-5.3 Flash produced a valid 4% hidden exact-match result in its first published max-effort run. The agent built several compact neural candidates, raised official public calibration from 0% to 64%, briefly promoted a weaker 60% package, then restored and finalized the 64% candidate. Independent grading reproduced the hidden score from the exact stored artifact, so the 60-point public-to-hidden gap reflects poor transfer rather than a stale-package failure. One seed remains provisional.
One valid autonomous run produced a provisional 4% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The final artifact was a compact byte-level transformer trained through a broad synthetic-curriculum search, with unusually careful late-stage package selection.
The run generated varied examples for copying, arithmetic, mappings, records, schedules, sorting, and text transformations.
The final model used a roughly 16.8-million-parameter byte transformer whose compressed package stayed just below the 16 MB cap.
The agent rebuilt and rescored the selected bundle, restored the best official candidate after a weaker promotion, and finalized matching code and weights.
A zero-percent safety checkpoint gave way to the first valid main package at 44% public calibration.
Candidate m7 reached the best official public score. A later 60% package was temporarily promoted despite scoring lower.
The agent rechecked the 64% package, selected it again, finished before the deadline, and received a valid 4% hidden grade.
The selected package passed the size and inference checks, matched its recorded sources, and reproduced 4% in an independent Docker re-grade.
Sixteen of twenty-five public questions were correct, but only four of one hundred hidden questions matched exactly.
Two more valid seeds are needed, and public model metadata does not establish whether this release is literally identical to the former anonymous Ox Alpha endpoint.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0104 | 4% | 55:21 | $0.05 | 15.46 MB | 10 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 10 candidate submission actions across the published run set. Selected artifacts averaged 15.46 MB compressed.
The runs averaged 55:21 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 80 agent messages · 10 candidate actions · 1 published run.
$0.05 per displayed run at current configured rates.
The official run cost an estimated $0.0531 at its launch pricing. That creates high displayed points per dollar, but the value ratio is driven by very low token prices and should not obscure the weak 4% hidden result. Preflight smoke requests are excluded.
At current configured API-equivalent rates, the displayed run cost is $0.05 and value is 75.33 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-08-26.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare GLM-5.3 Flash with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard