Run postmortem · Published 2026-08-14

Kimi K2.7 Codexhigh reasoning

Twenty-nine submissions expanded a solver that never became broad enough.

3/ 100
hidden exact

One-hour track · xhigh reasoning · provisional

ARI score scaleexact-match accuracy
3%
0255075100
StatusProvisional
Valid seeds1 / 3
Score rank#18
Current cost$3.77
Points / $0.8
Average time52:13
01 · Verdict

A lot of patching, very little transfer

Kimi K2.7 Code built a compact rule-driven inference system and submitted candidates twenty-nine times. The solver could handle several visible task families, but it never moved beyond 52% public calibration and produced only three hidden exact matches.

One valid autonomous run produced a provisional 3% hidden exact-match score.

#18
Score position3% on the common 0–100 scale. Tied valid scores share a rank.

1 of 3 required seeds · Hidden exact-match accuracy, not public calibration accuracy.

02 · Approach

What Kimi K2.7 Code built

The final artifact prioritized explicit solvers over a large trained model.

01

Broad deterministic parser

It implemented records, mappings, arithmetic, sorting, sequences, time operations, text transforms, and odd-one-out logic.

02

Small final package

The selected artifact was only 1.84 MB, leaving size headroom for more coverage.

03

High submission cadence

Twenty-nine candidates show that the run repeatedly patched and tested the solver rather than waiting for a final build.

03 · Trajectory

How the run unfolded

Opening

Chose direct inference

The run treated the benchmark as a family of contextual transformations rather than a pure language-model training task.

Middle

Expanded task coverage

The agent kept adding solver branches and resubmitting, but public accuracy stalled near half.

52:13

Finalized below readiness

The hidden exam returned 3%, confirming that the visible parser remained brittle.

04 · Assessment

Good, bad, and unresolved

What worked

Correct strategic direction

Direct contextual solving was the same broad mechanism used by the benchmark's stronger systems.

What hurt

Patches did not create general rules

High submission volume without a strong public score suggests narrow fixes rather than robust abstractions.

What remains unknown

Whether the parser could be consolidated

The archive cannot show whether fewer, more general mechanisms would have transferred better than many special cases.

05 · Evidence

The retained run record

Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.

Swipe horizontally to see all run columns.

RunScoreTimeCost at runArtifactCandidatesGrade
run-00453%52:13$3.771.84 MB29Valid
338.2KInput tokens
72.2KOutput tokens
21.8MCache-read tokens
01

One run, one outcome

Autonomous runs are stochastic. Two more valid seeds are required before this can be treated as an official estimate.

02

Candidate behavior

The agent made 29 candidate submission actions across the published run set. Selected artifacts averaged 1.84 MB compressed.

03

Time use

The runs averaged 52:13 of wall time. 1 run explicitly finalized before the one-hour limit.

Retained totals: 192 agent messages · 29 candidate actions · 1 published run.

06 · Economics

What the result cost

$3.77 per displayed run at current configured rates.

Kimi K2.7 Code was neither a cheap experiment nor a competitive scorer. Most of the spend funded repeated solver iteration that did not generalize.

At current configured API-equivalent rates, the displayed run cost is $3.77 and value is 0.8 score points per dollar.

#19
Value positionValue rank #19 of 22 valid priced entries. Higher score points per dollar indicate better value.

Current configured API-equivalent cost / run · pricing snapshot retained archive basis.

07 · Context

Nearby results

Adjacent rows provide score context without treating small one-seed differences as settled model rankings.

Compare Kimi K2.7 Code with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.

View the leaderboard