Run postmortem · Published 2026-08-14

GLM-5.2xhigh reasoning

Two radically different packages landed one point apart.

13.5/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
13.5%
0255075100
StatusProvisional
Valid seeds2 / 3
Score rank#12
Current cost$3.29
Points / $4.1
Average time56:58
01 · Verdict

Resilient execution, modest hidden ceiling

GLM-5.2 produced valid 14% and 13% scores across two runs even though one agent session failed after leaving a gradeable checkpoint. The final artifacts differed by more than twenty times in size, yet their hidden outcomes were nearly identical.

2 valid autonomous runs averaged 13.5%; one more seed is required for an official estimate.

#12
Score position13.5% on the common 0–100 scale. Tied valid scores share a rank.

2 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GLM-5.2 built

Both seeds emphasized direct contextual problem solving, but they packaged it very differently.

01

Rule-driven inference

The systems covered arithmetic, sorting, sequences, transformations, time operations, mappings, and records before neural fallback.

02

Tiny emergency checkpoint

The first run retained a valid sub-megabyte artifact even though the agent process later failed.

03

Larger second build

The second seed expanded the inference implementation and produced a roughly 10.5 MB final package.

03 · Trajectory

How the run unfolded

Seed 1

Saved by checkpoint discipline

A valid candidate survived the later agent failure and still scored 14%.

Seed 2

Used almost the full hour

The second run kept iterating, reached perfect calibration, and finished at 13%.

Outcome

Different size, same score band

The one-point spread suggests that package size alone was not the limiting factor.

04 · Assessment

Good, bad, and unresolved

What worked

The protocol preserved useful work

Promotion worked as intended: an agent failure did not erase the first run's valid artifact.

What hurt

More code did not improve generalization

The substantially larger second artifact scored one point lower than the compact checkpoint.

What remains unknown

Third-seed stability

The two runs are close, but a third is still needed to determine whether 13.5% is a stable estimate.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-003514%55:39$1.390.48 MB4Valid
run-005813%58:1710.48 MB11Valid
4.32MInput tokens
79.6KOutput tokens
11.1MCache-read tokens
01

Promising, still provisional

The valid scores span 13% to 14%. A third run is required for an official estimate.

02

Candidate behavior

The agent made 15 candidate submission actions across the published run set. Selected artifacts averaged 5.48 MB compressed.

03

Time use

The runs averaged 56:58 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 176 agent messages · 15 candidate actions · 2 published runs.

06 · Economics

What the result cost

$3.29 per displayed run at current configured rates.

GLM-5.2 occupies the middle of the value table. The more important economic lesson is that the much larger second package did not buy a higher score.

At current configured API-equivalent rates, the displayed run cost is $3.29 and value is 4.1 score points per dollar.

#12
Value positionValue rank #12 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GLM-5.2 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard