Run postmortem · Published 2026-09-21

MiMo V2.6 Proreasoning enabled

MiMo V2.6 Pro completed the one-hour benchmark with 18/100 hidden answers correct and a $0.20 estimated API cost.

18/ 100
hidden exact

One-hour track · reasoning enabled · provisional

ARI score scaleexact-match accuracy
18%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#11
Current cost$0.20
Points / $92.12
Average time56:41
01 · Verdict

A completed rerun scored 18/100

MiMo V2.6 Pro finalized a valid artifact and answered 18 of 100 hidden questions correctly. It passed all 25 public validation questions and finished normally in 56 minutes 41 seconds. This is one valid seed in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the small system the agent builds, rather than asking the controller model to answer the hidden questions directly.

One valid autonomous run produced a provisional 18% hidden exact-match score.

#11
Score position18% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What MiMo V2.6 Pro built

The agent developed a compact inference package with task-specific solver logic and a neural fallback.

01

Iterative solver development

MiMo revised and tested its inference code throughout the run, recovering from packaging and code errors.

02

A compact submitted system

The finalized artifact was 3,299,124 bytes gzip, within the 16 MB limit.

03

Explicit final selection

The agent submitted nine candidates and explicitly finalized solver_v9. The hidden grader used that saved snapshot.

03 · Trajectory

How the run unfolded

Early candidates

Packaging and validation errors

The first candidate had an invalid output shape; the next two were valid but scored 0/25 on public validation.

Candidate four

Public validation reached 25/25

The fourth candidate passed all 25 public questions. Later candidates retained that public score.

Finalization

Candidate nine was selected

MiMo finalized solver_v9 and ended the agent session successfully before the one-hour deadline.

Hidden grading

18 of 100 exact matches

The final package passed validation and scored 18%. The controller exited with code 0 and no broker errors.

04 · Assessment

Good, bad, and unresolved

What worked

Completed without an approval interruption

OpenCode 1.18.32 and explicit permission for repeated allowed calls let this run finish normally. All 67 broker requests received responses.

What hurt

Public validation did not generalize

A perfect 25/25 public score became 18/100 on the separate hidden exam.

What remains unknown

One completed seed

The earlier run-0126 ended on an automatic permission rejection without a valid submitted result; it is retained as an interrupted attempt. This row uses completed run-0127 alone. The rerun changed both the OpenCode version and loop permission, so it does not isolate the effect of either change.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-012718%56:41$0.203.30 MB10Valid
210.5KInput tokens
84.2KOutput tokens
8.5MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 10 candidate submission actions across the published run set. Selected artifacts averaged 3.30 MB compressed.

03

Time use

The runs averaged 56:41 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 73 agent messages · 10 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.20 per displayed run at current configured rates.

The plotted cost is the run-record token-based estimate of US$0.1954, displayed as $0.20. Saved Xiaomi endpoint rates were $0.435/M input, $0.87/M output, and $0.0036/M cache reads. Recorded usage was 210,500 input, 84,200 output, and 8.5 million cache-read tokens. OpenCode separately reports US$0.26. Neither figure is a reconciled provider invoice. The run used OpenRouter xiaomi/mimo-v2.6-pro through Xiaomi fp8, with reasoning enabled and no named effort setting. OpenCode 1.18.32 is the default harness for new runs in the normal benchmark track; historical results retain their original versions.

At current configured API-equivalent rates, the displayed run cost is $0.20 and value is 92.12 score points per dollar.

#1
Value positionValue rank #1 of 33 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare MiMo V2.6 Pro with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard