Run postmortem · Published 2026-08-14

MiniMax M3xhigh reasoning

Sixty percent public calibration collapsed to one hidden match.

1/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
1%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#19
Current cost$0.38
Points / $2.63
Average time52:12
01 · Verdict

The visible set overstated a brittle system

MiniMax M3 reached 60% public calibration and produced a valid 9.52 MB package, but the final inference path remained essentially neural. The hidden score of 1% shows that partial public success did not represent broad contextual generalization.

One valid autonomous run produced a provisional 1% hidden exact-match score.

#19
Score position1% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What MiniMax M3 built

The run made only light changes around the baseline model.

01

Standard GPT-style model

The final artifact used the conventional byte transformer supplied by the benchmark structure.

02

Minimal inference wrapper

The inference file added configuration handling but no broad family of deterministic solvers.

03

Sparse candidate search

Only four submissions were recorded across a 52-minute run.

03 · Trajectory

How the run unfolded

Opening

Established a neural package

The agent concentrated on getting a trained artifact into the runner.

Middle

Reached 60% public calibration

That visible result suggested partial functionality but still left ten of twenty-five items wrong.

52:11

Hidden transfer failed

The final artifact produced a single exact match on the private exam.

04 · Assessment

Good, bad, and unresolved

What worked

Low dollar cost

At thirty-eight cents, MiniMax completed a valid near-full-hour benchmark run cheaply.

What hurt

No task-aware inference layer

The final package lacked the direct contextual solving that characterized every upper-tier generated result.

What remains unknown

Why iteration stayed sparse

The archive shows four candidates but cannot establish whether provider latency, agent decisions, or both limited the search.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00431%52:12$0.389.52 MB4Valid
303.2KInput tokens
45.8KOutput tokens
3.9MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 4 candidate submission actions across the published run set. Selected artifacts averaged 9.52 MB compressed.

03

Time use

The runs averaged 52:12 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 102 agent messages · 4 candidate actions · 1 published run.

06 · Economics

What the result cost

$0.38 per displayed run at current configured rates.

MiniMax was cheap in dollars but expensive per useful outcome. A low price cannot rescue a system that scored only one point.

At current configured API-equivalent rates, the displayed run cost is $0.38 and value is 2.63 score points per dollar.

#15
Value positionValue rank #15 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare MiniMax M3 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard