Run postmortem · Published 2026-08-14

North Mini Codexhigh reasoning

Free inference did not make the run cheap in the resource that mattered: the hour.

0/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
0%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#22
Current cost$0.00
Points / $
Average time44:57
01 · Verdict

A busy run without a working selector

North Mini Code trained repeatedly and built a custom small GPT plus heuristic inference, but never established a non-zero public calibration score. The agent process later failed, leaving the promoted artifact to grade validly at 0%.

One valid autonomous run completed, but it produced no exact matches on the hidden exam.

#22
Score position0% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What North Mini Code built

The run tried both neural scale and direct rules, but neither produced a usable feedback loop.

01

Ten training attempts

The agent cycled through small transformers and several synthetic-data schedules rather than committing to one stable baseline.

02

Custom hybrid package

The final artifact combined a compact SimpleGPT implementation with lookup, translation, arithmetic, and sequence rules.

03

Promotion without calibration

Seven candidate submissions remained valid packages, but the best retained public accuracy was still 0%.

03 · Trajectory

How the run unfolded

Opening

Hit missing-tool friction

Expected inspection utilities were unavailable, slowing the initial attempt to understand the data path.

Middle

Trained broadly without a positive signal

The agent kept changing models and data schedules even though calibration never moved above zero.

44:58

Agent failed, checkpoint survived

The promoted artifact remained gradeable and returned a valid 0%, not an infrastructure invalidation.

04 · Assessment

Good, bad, and unresolved

What worked

A valid artifact survived failure

Checkpointing preserved a real benchmark outcome even after the agent process stopped successfully iterating.

What hurt

No evidence-guided progress

Repeated training without any public accuracy signal meant the run could not distinguish improvement from churn.

What remains unknown

Model capability versus execution failure

One failed agent run cannot separate the model's potential from its inability to establish a useful workflow here.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00380%44:57$0.001.39 MB7Valid
2.8MInput tokens
43.1KOutput tokens
0Cache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 7 candidate submission actions across the published run set. Selected artifacts averaged 1.39 MB compressed.

03

Time use

The runs averaged 44:57 of wall time. Runtime alone cannot show whether more time would have improved the result.

Retained totals: 121 agent messages · 7 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.00 per displayed run at current configured rates.

The API-equivalent price was zero, so points-per-dollar is undefined rather than infinite. The run still consumed GPU time and most of the one-hour benchmark window.

At current configured API-equivalent rates, the displayed run cost is $0.00 and value is — score points per dollar.

Value positionValue rank unavailable. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare North Mini Code with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard