Saved by checkpoint discipline
A valid candidate survived the later agent failure and still scored 14%.
Run postmortem · Published 2026-08-14
Two radically different packages landed one point apart.
One-hour track · xhigh reasoning · provisional
GLM-5.2 produced valid 14% and 13% scores across two runs even though one agent session failed after leaving a gradeable checkpoint. The final artifacts differed by more than twenty times in size, yet their hidden outcomes were nearly identical.
2 valid autonomous runs averaged 13.5%; one more seed is required for an official estimate.
2 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
Both seeds emphasized direct contextual problem solving, but they packaged it very differently.
The systems covered arithmetic, sorting, sequences, transformations, time operations, mappings, and records before neural fallback.
The first run retained a valid sub-megabyte artifact even though the agent process later failed.
The second seed expanded the inference implementation and produced a roughly 10.5 MB final package.
A valid candidate survived the later agent failure and still scored 14%.
The second run kept iterating, reached perfect calibration, and finished at 13%.
The one-point spread suggests that package size alone was not the limiting factor.
Promotion worked as intended: an agent failure did not erase the first run's valid artifact.
The substantially larger second artifact scored one point lower than the compact checkpoint.
The two runs are close, but a third is still needed to determine whether 13.5% is a stable estimate.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0035 | 14% | 55:39 | $1.39 | 0.48 MB | 4 | Valid |
| run-0058 | 13% | 58:17 | — | 10.48 MB | 11 | Valid |
The valid scores span 13% to 14%. A third run is required for an official estimate.
The agent made 15 candidate submission actions across the published run set. Selected artifacts averaged 5.48 MB compressed.
The runs averaged 56:58 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 176 agent messages · 15 candidate actions · 2 published runs.
$3.29 per displayed run at current configured rates.
GLM-5.2 occupies the middle of the value table. The more important economic lesson is that the much larger second package did not buy a higher score.
At current configured API-equivalent rates, the displayed run cost is $3.29 and value is 4.1 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare GLM-5.2 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard