Model run analyses

Open the run.
Keep the exam closed.

Evidence-based accounts of how frontier models used their time, what they built, what it cost, and what the final score can—and cannot—support.

Grok 4.6: the $9.80 run, explained

A fast hybrid build, the strongest observed single-run score, and 18 million cached-context tokens.

26%provisional

Claude Opus 5 on ARI Bench

An ambitious neural-symbolic system, perfect public calibration, and a valid hidden-exam result.

23%provisional

These pages disclose aggregate evidence and high-level agent behavior. Hidden questions, answers, private grading records, credentials, and raw private transcripts remain unpublished.