Run postmortem · Published 2026-08-21

Ox Alpha · OpenCode 1.18.21max reasoning

A clean controller rerun fixed the transport uncertainty, then exposed a much harder model-selection failure.

1/ 100
hidden exact

One-hour track · max reasoning · provisional

ARI score scaleexact-match accuracy
1%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#23
Current cost$0.00
Points / $
Average time55:33
01 · Verdict

The harness worked; the selected strategy did not generalize

Ox Alpha completed the OpenCode 1.18.21 control run without the unknown-stop or network failure suspected in the earlier attempt. It finalized a valid 15.58 MB candidate that scored 52% on the 25-question public calibration set, but only 1% on the 100-question hidden exam. Matching hashes across the built, promoted, finalized, and graded artifacts rule out a packaging swap.

One valid autonomous run produced a provisional 1% hidden exact-match score.

#23
Score position1% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Ox Alpha · OpenCode 1.18.21 built

The run concentrated on fitting the largest byte transformer possible under the compressed artifact cap and training it on increasingly task-shaped synthetic corpora.

01

Cap-aware model search

An oversized first candidate was rejected, after which the agent reduced the network to an 8-layer, 288-wide model whose FP16 package fit below 16 MB.

02

Synthetic curriculum expansion

Successive corpora emphasized arithmetic, records, mappings, sorting, casing, schedules, and sequences based on visible calibration failures.

03

Checkpoint selection and restoration

Seven named candidates were retained; after later regressions, the agent restored and finalized c6, the best public checkpoint at 52%.

03 · Trajectory

How the run unfolded

Opening

Recovered from an 81 MB rejection

The broker correctly rejected the first package, and the agent iterated until a valid 15.57 MB checkpoint was promoted.

Late middle

Raised public calibration to 52%

Continuation training on a broader synthetic mix lifted the official calibration result from 36% to 52%.

55:33

Finalized c6 and scored 1% hidden

The run ended cleanly and early, but the hidden exam rewarded only one exact answer: a small numeric sort.

04 · Assessment

Good, bad, and unresolved

What worked

The OpenCode upgrade removed the original ambiguity

OpenCode 1.18.21 completed one continuous max-effort session with no broker, transport, or agent return-code error.

What hurt

Public calibration became a misleading selector

The selected neural system rose to 52% on visible calibration while collapsing to 1% on broader locally defined hidden tasks.

What remains unknown

How much the Pi harness changes execution

The next controlled run keeps the model, effort, exam, grader, budget, and artifact rules fixed while replacing OpenCode with Pi.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-01031%55:33$0.0015.58 MB8Valid
191.9KInput tokens
40.2KOutput tokens
2MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 8 candidate submission actions across the published run set. Selected artifacts averaged 15.58 MB compressed.

03

Time use

The runs averaged 55:33 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 75 agent messages · 8 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.00 per displayed run at current configured rates.

OpenRouter priced Ox Alpha at zero for this run, so the recorded API cost is $0.00. Hardware and electricity are excluded, and a zero-cost row does not receive a finite points-per-dollar ratio.

At current configured API-equivalent rates, the displayed run cost is $0.00 and value is — score points per dollar.

Value positionValue rank unavailable. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-08-21.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Ox Alpha · OpenCode 1.18.21 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard