Packaging and validation errors
The first candidate had an invalid output shape; the next two were valid but scored 0/25 on public validation.
Run postmortem · Published 2026-09-21
MiMo V2.6 Pro completed the one-hour benchmark with 18/100 hidden answers correct and a $0.20 estimated API cost.
One-hour track · reasoning enabled · provisional
MiMo V2.6 Pro finalized a valid artifact and answered 18 of 100 hidden questions correctly. It passed all 25 public validation questions and finished normally in 56 minutes 41 seconds. This is one valid seed in the normal one-hour benchmark, so the result remains provisional. ARI Bench scores the small system the agent builds, rather than asking the controller model to answer the hidden questions directly.
One valid autonomous run produced a provisional 18% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The agent developed a compact inference package with task-specific solver logic and a neural fallback.
MiMo revised and tested its inference code throughout the run, recovering from packaging and code errors.
The finalized artifact was 3,299,124 bytes gzip, within the 16 MB limit.
The agent submitted nine candidates and explicitly finalized solver_v9. The hidden grader used that saved snapshot.
The first candidate had an invalid output shape; the next two were valid but scored 0/25 on public validation.
The fourth candidate passed all 25 public questions. Later candidates retained that public score.
MiMo finalized solver_v9 and ended the agent session successfully before the one-hour deadline.
The final package passed validation and scored 18%. The controller exited with code 0 and no broker errors.
OpenCode 1.18.32 and explicit permission for repeated allowed calls let this run finish normally. All 67 broker requests received responses.
A perfect 25/25 public score became 18/100 on the separate hidden exam.
The earlier run-0126 ended on an automatic permission rejection without a valid submitted result; it is retained as an interrupted attempt. This row uses completed run-0127 alone. The rerun changed both the OpenCode version and loop permission, so it does not isolate the effect of either change.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0127 | 18% | 56:41 | $0.20 | 3.30 MB | 10 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 10 candidate submission actions across the published run set. Selected artifacts averaged 3.30 MB compressed.
The runs averaged 56:41 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 73 agent messages · 10 candidate actions · 1 published run.
$0.20 per displayed run at current configured rates.
The plotted cost is the run-record token-based estimate of US$0.1954, displayed as $0.20. Saved Xiaomi endpoint rates were $0.435/M input, $0.87/M output, and $0.0036/M cache reads. Recorded usage was 210,500 input, 84,200 output, and 8.5 million cache-read tokens. OpenCode separately reports US$0.26. Neither figure is a reconciled provider invoice. The run used OpenRouter xiaomi/mimo-v2.6-pro through Xiaomi fp8, with reasoning enabled and no named effort setting. OpenCode 1.18.32 is the default harness for new runs in the normal benchmark track; historical results retain their original versions.
At current configured API-equivalent rates, the displayed run cost is $0.20 and value is 92.12 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot 2026-09-22.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare MiMo V2.6 Pro with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard