Established a neural package
The agent concentrated on getting a trained artifact into the runner.
Run postmortem · Published 2026-08-14
Sixty percent public calibration collapsed to one hidden match.
One-hour track · xhigh reasoning · provisional
MiniMax M3 reached 60% public calibration and produced a valid 9.52 MB package, but the final inference path remained essentially neural. The hidden score of 1% shows that partial public success did not represent broad contextual generalization.
One valid autonomous run produced a provisional 1% hidden exact-match score.
1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.
The run made only light changes around the baseline model.
The final artifact used the conventional byte transformer supplied by the benchmark structure.
The inference file added configuration handling but no broad family of deterministic solvers.
Only four submissions were recorded across a 52-minute run.
The agent concentrated on getting a trained artifact into the runner.
That visible result suggested partial functionality but still left ten of twenty-five items wrong.
The final artifact produced a single exact match on the private exam.
At thirty-eight cents, MiniMax completed a valid near-full-hour benchmark run cheaply.
The final package lacked the direct contextual solving that characterized every upper-tier generated result.
The archive shows four candidates but cannot establish whether provider latency, agent decisions, or both limited the search.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Cost at run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0043 | 1% | 52:12 | $0.38 | 9.52 MB | 4 | Valid |
Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.
The agent made 4 candidate submission actions across the published run set. Selected artifacts averaged 9.52 MB compressed.
The runs averaged 52:12 of wall time. 1 run explicitly finalized before the one-hour limit.
Retained totals: 102 agent messages · 4 candidate actions · 1 published run.
$0.38 per displayed run at current configured rates.
MiniMax was cheap in dollars but expensive per useful outcome. A low price cannot rescue a system that scored only one point.
At current configured API-equivalent rates, the displayed run cost is $0.38 and value is 2.63 score points per dollar.
Current configured API-equivalent cost / run · pricing snapshot retained archive basis.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare MiniMax M3 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard