13/100 — original attempt
The agent finished in 51m 56s. An abandoned filesystem diagnostic blocked the compute queue for roughly 20 minutes before the checkpoint timeout released it.
Run postmortem · Published 2026-09-22
Grok 4.7 xhigh averaged 16.33/100 across all three graded runs: 13, 19, and 17.
One-hour track · xhigh reasoning · blended
The published Grok 4.7 xhigh score is the arithmetic mean of all three graded artifacts: (13 + 19 + 17) / 3 = 16.33%. Each run used the same 100-question hidden exam, giving 49 correct answers across 300 graded question attempts. All artifacts were valid. The runs span a harness fix and an OpenCode update, and one agent session ended on a provider error. This is a transparent blend of the three attempts, not a claim that three identical execution settings were repeated. It appears in the normal one-hour benchmark charts.
Three valid artifact grades, weighted equally: 13%, 19%, and 17%, averaging 16.33%.
All three graded attempts · Hidden exact-match accuracy, not public calibration accuracy.
The same Grok 4.7 model and xhigh reasoning setting were used through OpenRouter across all three runs.
All three runs used the same hidden exam, one-hour maximum budget, RTX 5090, and 16 MB compressed-artifact limit.
Each score comes from the saved artifact selected by the benchmark rules. The provider-interrupted run retained a valid submitted candidate.
Runs 0124 and 0125 used OpenCode 1.17.11. Run 0128 used the new default, OpenCode 1.18.32, with explicit repeated-call permission and the runner cancellation fix.
The agent finished in 51m 56s. An abandoned filesystem diagnostic blocked the compute queue for roughly 20 minutes before the checkpoint timeout released it.
After the runner cancellation fix, the saved artifact scored 19/100. The provider ended the agent session with a temporary-unavailability error after 42m 30s; the remaining budget was unused.
The agent finished normally in 53m 33s, explicitly finalized candidate cand68, and had no broker errors or cancelled compute jobs.
All three valid grades contribute equally. No attempt is omitted and the 19/100 result is not selected on its own.
The full sequence is represented, including both operationally affected attempts and the clean final run.
The first run lost compute time to a stalled diagnostic; the second ended early on a provider failure. Their valid artifact scores remain in the user-requested average.
This is not a controlled three-seed comparison under one frozen harness version. The blend does not isolate the effect of the OpenCode upgrade or establish a stable ranking.
Every archived run selected for this exact model-and-effort row is shown below. The narrative uses aggregate run evidence and final artifact structure; it does not expose hidden questions, answers, or raw private transcripts.
Swipe horizontally to see all run columns.
| Run | Score | Time | Estimated cost per run | Artifact | Candidates | Grade |
|---|---|---|---|---|---|---|
| run-0124 | 13% | 51:56 | $14.17 | 0.02 MB | 55 | Valid |
| run-0125 | 19% | 42:30 | $17.89 | 0.03 MB | 52 | Valid |
| run-0128 | 17% | 53:33 | $34.44 | 3.22 MB | 91 | Valid |
The valid scores span 13% to 19%. All three attempts are included; their execution versions and interruptions differ.
The agent made 198 candidate submission actions across the published run set. Selected artifacts averaged 1.09 MB compressed.
The runs averaged 49:20 of wall time. 2 runs explicitly finalized before the one-hour limit.
Retained totals: 946 agent messages · 198 candidate actions · 3 published runs.
$22.17 per displayed run at current configured rates.
The cost and value graphs use the mean estimated cost across all three attempts, $22.17 per run. This is the midpoint of an approximate $21.62–$22.72 average range, derived from retained per-request token snapshots plus remaining rounded token totals. Per-run ranges are $13.72–14.63, $17.84–17.94, and $33.30–35.58. Verified endpoint rates are $1.60/M input, $4.80/M output, and $0.40/M cache reads, doubled for prompts at or above 200,000 tokens. Some receipt detail is incomplete, including a truncated final export for run-0128. These are estimates, not reconciled invoices. Value is the blended score divided by the estimated mean cost, rather than a mean of individual score/cost ratios.
At current configured API-equivalent rates, the displayed run cost is $22.17 and value is 0.74 score points per dollar. The archived base-rate estimates average $18.97 and undercount long-context charges. The plotted $22.17 corrects that using an estimated range; it is not settled billing.
Estimated mean cost per run; long-context adjustment and incomplete receipts disclosed above.
Adjacent rows provide score context without treating small one-seed differences as settled model rankings.
Compare Grok 4.7 with the complete ARI Bench field, including score, current cost, value per dollar, effort, and seed status.
View the leaderboard