Run postmortem · Published 2026-08-18

GLM-5.3max reasoning

It diagnosed the packaging path, lifted calibration from 24% to 92%, and finished at 18% hidden accuracy.

18/ 100
hidden exact

One-hour track · max reasoning · provisional

ARI score scaleexact-match accuracy
18%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#6
Current cost$3.84
Points / $4.69
Average time52:28
01 · Verdict

Strong recovery turned a broken package into a competitive seed

GLM-5.3 produced a valid 18% result in its first published max-effort run, four points above the best prior GLM-5.2 seed. Its first two submitted hybrids remained at 24% public calibration because the packaged inference file and trained weights came from different source paths. The agent tested the actual packaged artifact, corrected that mismatch, and raised calibration to 92%. The outcome is promising, but one seed remains provisional.

One valid autonomous run produced a provisional 18% hidden exact-match score.

#6
Score position18% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What GLM-5.3 built

The final artifact combined a compact trained fallback with a broad deterministic solver and an unusually productive package-test-repair loop.

01

Compact byte-level fallback

The package retained a roughly five-million-parameter, six-layer byte transformer trained on broad synthetic examples.

02

Contextual inference layer

Deterministic handlers covered arithmetic, records, mappings, sequences, copying, and text transformations before neural fallback.

03

Packaged-path validation

The agent stopped trusting workspace-only checks, exercised the actual candidate package, and found that inference code and weights were being sourced differently.

03 · Trajectory

How the run unfolded

Opening 14m

Built the hybrid and hit 24%

Local evaluation looked much stronger, but the first official candidates stayed at 24% public calibration.

36:49

Found the real packaging path

Testing the assembled artifact exposed the split-source mismatch; the corrected candidate immediately reached 92% calibration.

52:28

Hardened and finalized

The run expanded synthetic coverage, strengthened deterministic rules, and finalized a valid 9.50 MB artifact that scored 18% hidden.

04 · Assessment

Good, bad, and unresolved

What worked

Excellent diagnosis under benchmark pressure

The run identified a real execution-path defect, recovered without a timeout or broker failure, and converted it into a strong hidden result.

What hurt

Public calibration still overstated transfer

The final package reached 92% on the visible set but only 18% hidden, leaving a large generalization gap and little signal for choosing among late revisions.

What remains unknown

Whether the result repeats

Two additional valid seeds are needed to determine whether 18% is a stable GLM-5.3 estimate or a favorable first run.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-000118%52:28$3.849.50 MB14Valid
155.6KInput tokens
49.4KOutput tokens
13.1MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 14 candidate submission actions across the published run set. Selected artifacts averaged 9.50 MB compressed.

03

Time use

The runs averaged 52:28 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 130 agent messages · 14 candidate actions · 1 published run.

06 · Economics

What the result cost

$3.84 per displayed run at current configured rates.

The run cost an estimated $3.84, or 4.69 hidden-score points per dollar. It outscored both prior GLM-5.2 seeds, but its value rank remains provisional until replication.

At current configured API-equivalent rates, the displayed run cost is $3.84 and value is 4.69 score points per dollar.

#8
Value positionValue rank #8 of 24 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-18.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare GLM-5.3 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard