The same task as the blind run, but the prompt now hands every model a numbered plan and names the traps. What is left to measure is instruction-following — how far a structured plan lifts the models that stumbled blind. Scores compare only against the same model's blind run, never across the two tests.
Ranked by the weighted criteria in the second table. Cost is the metered vendor rate with what was really paid beneath it; local models are free. Peak context is the largest single request, against each model's own window.
| # | Model | Score | Cost | Wall | Tokens | Peak ctx | Window | Commits | Bugs |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GLM 5.3 Flashpi · fireworks · think high | 98 9 tool errors |
$0.25paid $0.25 | 24 min | 6.9 M | 72 k | 7 % | 18 | 0 |
| 2 | DeepSeek V4 Flash (0731)pi · fireworks · think high | 97 1 model nudge · 7 tool errors · 1 compaction · repetition loop 0.02 on text |
$2.10paid $2.47 | 249 min | 68.1 M | 996 k | 97 % | 16 | 0 |
| 3 | GPT-5.6 Lunapi · openai-codex · think high | 89 |
$0.60plan ≈$0.05 | 40 min | 29.3 M | 232 k | 85 % | 18 | 2 |
| 4 | Claude Sonnet 4.5pi · anthropic · think high | 88 16 tool errors |
$3.92paid $3.92 | 32 min | 10.3 M | 93 k | 9 % | 16 | 3 |
| 5 | Qwen3.6 35B-A3Bpi · llama · local · think high | 83 peak_context corrected 2026-09-04: the earlier 77849 was the value after the compaction. · 25 tool errors · 1 compaction |
$0plan ≈$0.00 | 92 min | 12.7 M | 94 k | 79 % | 16 | 5 |
| 6 | Claude Haiku 4.5pi · anthropic · think high | 76 17 tool errors |
—paid $1.73 | 27 min | 13.9 M | 108 k | 54 % | 19 | 3 |
| 7 | Qwen3.6 35B-A3Bpi · llama · local · think off | 63 24 tool errors · 1 compaction |
$0plan ≈$0.00 | 89 min | 13.0 M | 78 k | 95.1 % | 16 | 8 |
| 8 | Gemma 4 26B-A4Bpi · llama · local · think highpartial · 7 of 8 | 57 7/8 done · stopped without finishing |
$0plan ≈$0.00 | 115 min | 24.8 M | 209 k | 98.1 % | 13 | 16 |
| 9 | Qwen3.6 35B-A3Bpi · llama · local · think off | 47 2 model nudges · 37 tool errors · 12 compactions |
$0plan ≈$0.00 | 96 min | 9.5 M | 52 k | 104.9 % | 7 | 16 |
| 10 | Gemma-4-12B (llama.cpp, off)pi · llama · local · think offpartial · 3 of 8 · model-nudge budget spent | 38 raw 58 · 3/8 done · model-nudge budget spent · repetition loop 0.02 on text |
$0plan ≈$0.00 | 98 min | 6.5 M | 125 k | 48 % | 3 | 6 |
| 11 | Ternary Bonsai 27B (PrismML GGUF)pi · llama · local · think highpartial · 3 of 8 · hit the 300 min time budget | 32 3/8 done · hit the 300 min time budget |
$0plan ≈$0.00 | 300 min | 0.0 M | 63 k | 96.28 % | 3 | 16 |
| 12 | Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 1 of 8 · hit the 300 min time budget | 13 raw 59 · 1/8 done · hit the 300 min time budget |
$0plan ≈$0.00 | 300 min | 3.6 M | 46 k | 80 % | 1 | 1 |
| 13 | bonsai-prism (f16 KV)pi · llama · local · think highpartial · 1 of 8 | 13 raw 36 · 1/8 done · stopped without finishing |
$0plan ≈$0.00 | 195 min | 21.0 M | 127 k | 97 % | 2 | 10 |
| — | Gemma 4 26B-A4Bpi · llama · local · think offpartial · 2 of 8 · repetition loop, stopped by the runner | 25 invalid — live loop stop: repetition_loop on 5 identical edit tool calls, 18:12:06Z to 18:12:47Z; 2 of 8 libraries landed · repetition loop 0.22 on tool call |
$0plan ≈$0.00 | 20 min | 2.6 M | 73 k | 34.15 % | 3 | 11 |
| — | Qwen3.8 27B (MLX 4-bit)pi · mlx · local · think lowpartial · 0 of 8 · tooling-nudge budget spent | 0 invalid — three Metal OOM server crashes, tooling budget exhausted; zero commits |
$0plan ≈$0.00 | 261 min | 1.3 M | 30 k | — | 0 | 3 |
| — | Gemma 4 12Bpi · lmstudio · local · think highpartial · 0 of 8 · model-nudge budget spent | 0 invalid — repetition loop on the LM Studio MLX thinking-on entry google/gemma-4-12b, pre-fix chat template (research run 2); zero commits |
$0plan ≈$0.00 | 46 min | 0.3 M | 30 k | 19 % | 0 | 4 |
| — | Gemma 4 12B (low reasoning)pi · lmstudio · local · think lowpartial · 0 of 8 · model-nudge budget spent | 0 invalid — repetition loop on the LM Studio MLX thinking-on entry google/gemma-4-12b, pre-fix chat template (research run 2); zero commits · repetition loop 0.02 on tool call |
$0plan ≈$0.00 | 99 min | 2.0 M | 45 k | 28 % | 0 | 6 |
| — | Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think offpartial · 0 of 8 · tooling-nudge budget spent | 0 invalid — gh token invalid on the host (HTTP 401), model looped on interactive gh auth login, tooling budget exhausted; zero commits |
$0plan ≈$0.00 | 84 min | 0.0 M | 5.3 k | 9 % | 0 | 3 |
| — | Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think offpartial · 0 of 8 · stopped | 0 invalid — operator stop: non-terminating identical-command loop (one bash ls/cat of a missing .taprc, 85 times in a row); zero commits · repetition loop 0.02 on tool call |
$0plan ≈$0.00 | 187 min | 2.0 M | 27 k | 47 % | 0 | 3 |
| # | Model | Score | Cost | Wall | Tokens | Peak ctx | Window | Commits | Bugs |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5claude-code · anthropic | 99 1 tool error |
$5.73plan ≈$0.38 | 31 min | 36.1 M | 180 k † | 18 % | 18 | 0 |
| 2 | Qwen3.8 27B (MLX 4-bit)pi · mlx · local · think lowpartial · 6 of 8 · harness crash | 75 raw 84 · 6/8 done · harness crash |
$0plan ≈$0.00 | 154 min | 1.1 M | 23 k | 85 % | 6 | 4 |
| 3 | Claude Haiku 4.5claude-code · anthropic | 68 6 tool errors |
$2.56plan ≈$0.17 | 18 min | 20.3 M | 160 k † | 80 % | 12 | 4 |
| 4 | Qwen3.6 35B-A3Bpi · llama · local · think high | 66 18 tool errors |
$0plan ≈$0.00 | 76 min | 12.1 M | 94 k | 96 % | 8 | 8 |
| 5 | Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 3 of 8 · stuck | 38 raw 69 · 3/8 done · stuck |
$0plan ≈$0.00 | 230 min | 2.1 M | 51 k | 89 % | 3 | 5 |
Bugs are weighted: critical 3, medium 2, minor 1. Every run
uses the frozen prompt-guided.txt. Claude runs on a
subscription plan show computed token cost, not a metered bill.
Peak ctx is the largest context one turn occupied, the
response of that turn included, across all compaction cycles.
count-tool-calls.mjs reads it from the session log of
the scored run. † marks the rows from the retired
claude-code harness: that harness writes its log in another shape,
which the counter does not read, so those cells keep the older
reading, which counts the prompt of the largest turn without the
response.
Every row was checked against the branch or the session log, not against what the model said it did. Weights in the left column; the best cell in each row is highlighted.
| Criterion | glm-flash | ds-flash | luna | sonnet-4.5 | qwen | haiku-4.5 | qwen | gemma (partial) | qwen | Gemma-4-12B (llama.cpp, off) (partial) | bonsai-prism (partial) | bonsai (partial) | bonsai-prism (f16 KV) (partial) | gemma (partial) | qwen3.8 (partial) | gemma-12b (partial) | gemma-12b-low (partial) | bonsai (partial) | bonsai (partial) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Correctness | |||||||||||||||||||
| Bugs remaining weighted, 25 pts | clean 25 | clean 25 | exit-hook regression (trap C) 19 | trap A: TypeError 16 | trap A: TypeError 16 | TypeError: glob(...).then is not a function (trap A) 16 | trap A still throws: glob(...).then is not a function 7 | trap A still throws: glob(...).then is not a function 10 | trap A still throws: glob(...).then is not a function 4 | none in the 3 landed replacements 25 | 2 regressions from a memory rewrite: tree-variation-walker 3/3, config.js #10 and #7 0 | no bugs in the 1 landed commit 25 | rimraf devDep removed while test/helpers/index.js still requires it; root rimraf and tmp never touched 1 | urlsafe-base64 swap left broken, tests guarded around it 4 | none 25 | none, no code landed 25 | none, no code landed 25 | none, no code changed (vacuous) 25 | none, no code changed (vacuous) 25 |
| Unit tests still green per package | 680/680, 1 skip | 710/709, lerna ok | clean · pipeline flake confirmed clean standalone (260/260) | 674/674, 1 skip | all package suites pass (core 71/72, 1 skip) | clean (674/674, 1 skip) | suites green, full suite re-run before every commit | full suite green at HEAD, per-package tap green | mendel-pipeline package suite green; root suite never green after the chalk edits | core 71/71, config 41/41 (one documented flake) | mendel-core and mendel-config fail at HEAD, green on the base | not run | mendel-pipeline cli-printer 15/15 green; the same package loses its test helper on a fresh install | not re-run; worktree left mid-edit | not run | not run | not run | not run | not run |
| Runtime smoke test ignore/exclude/external globs | SYNC OK | SYNC OK | SYNC OK | THREW: TypeError | THREW: TypeError | THREW: TypeError | apply-extra-options.js:6 THREW TypeError | apply-extra-options.js:6 THREW TypeError | apply-extra-options.js:6 THREW TypeError | vacuous pass, apply-extra-options.js untouched | SYNC OK, glob never touched | not touched | SYNC OK, but glob was never touched (vacuous pass) | vacuous pass, apply-extra-options.js untouched | not touched | not touched | not touched | not touched (pack: SYNC OK on base code) | not touched |
| Task completion | |||||||||||||||||||
| All 8 libraries done 20 pts (shared) | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 | 8 / 8 done | 7 / 8 done, rimraf incomplete | 8 / 8 done | 3 / 8 done | 3 / 8 (uuid, xtend, urlsafe-base64) | 1 / 8 (uuid only) | 1 / 8 (chalk) | 2 / 8 done | 0 / 8 | 0 / 8 | 0 / 8 | 0 / 8 | 0 / 8 |
| No code still requires it | none left | none left | clean | none left | none left | clean | 0 stale requires | 2 stale rimraf requires in mendel-requirify tests | 0 stale requires | 5 libraries left, rimraf half done | rimraf ticked done, still required in 2 requirify tests | 27 files left | 22 stale requires across 7 libraries | 23 stale requires in 12 packages | 28 files left | 28 sites left | 28 sites left | 28 files left | 28 files left |
| Gone from every package.json | clean | clean | none left | clean | clean | none left | 1 stale entry, rimraf in legacy mendel-requirify | 1 stale entry, mendel-requirify package.json:26 | 0 stale entries, root cleared | 5 left, root rimraf + tmp still declared | 13 stale entries, root rimraf and tmp among them | 16 files left | 16 stale entries, root rimraf and tmp among them | 15 entries left, root rimraf + tmp still declared | 18 files left | 18 left | 18 left | 18 files left | 18 files left |
| Found the reference the issue missed legacy-packages/mendel-requirify | found + fixed | found + fixed | found + fixed | found + fixed | found + fixed | found + fixed | found, but two pnpm remove tries failed and the model gave up | seen in its own grep, then skipped (trap B) | found and removed (trap B) | grepped and edited it, never committed | not found (trap B) | not reached | not found (trap B) | never found | not reached | grepped it, never acted | found it in the plan, never acted | not reached | not reached |
| Subtotal 20 pts | 20 | 20 | 20 | 20 | 20 | 20 | 17 | 16 | 20 | 7.5 | 6 | 2 | 2 | 4 | 0 | 0 | 0 | 0 | 0 |
| Dependency hygiene | |||||||||||||||||||
| node_modules actually pruned 8 pts | real install 8 | real install 8 | real install 8 | real remove, verified 8 | pnpm remove per package, reinstall verified 8 | install unverified · only root deps via `pnpm remove -w` 4 | real pnpm install, store pruned 8 | pnpm used at root, but tmp cut from 3 package.json by hand 5 | pnpm used, but tmp and shasum cut with sed; pnpm-lock.yaml still lists both 3 | pnpm remove x4, store pruned 8 | pnpm used for uuid only, later edits by hand 4 | not pruned 0 | real pnpm install, lockfile regenerated 8 | real pnpm install, store pruned 8 | not pruned 0 | never ran pnpm 0 | never ran pnpm 0 | not pruned 0 | not pruned 0 |
| Lockfile pruned lines removed | +199/−106 lines | +134/−133 lines | +4/−197 lines | +123/−101 lines | +22/-211 lines | +91/−99 lines | 211 lines removed, 22 added | 206 lines removed, 26 added; --frozen-lockfile fails at HEAD | 183 lines removed, 47 added; --frozen-lockfile fails at HEAD | 121 lines | 107 lines removed, 1 added; xtend, urlsafe-base64 and rimraf still listed | 0 lines | 12 lines removed, 0 added; chalk and pipeline rimraf only | 20 lines removed, 5 added | 0 lines | 0 lines | 0 lines | 0 lines | 0 lines |
| Static checks | |||||||||||||||||||
| Prettier & ESLint clean 5 pts · re-run on the branch | clean, self-checked 2x 5 | clean, self-checked 14x 5 | clean, self-checked 4x 5 | clean, self-checked 5 | clean, self-checked via pnpm test (static) 5 | clean, self-checked after last commit 5 | prettier fails on 9 files clean at base; 0 self-runs, all 16 commits --no-verify so no hook caught it 0 | clean on re-run, never ran either tool itself 5 | clean on re-run, never ran either tool itself 3 | clean on re-run, never ran itself 3 | prettier and eslint clean on re-run, never self-run 3 | prettier warn (mendel-config/legacy) 0 | eslint fails at the tip on a self-made implicit dependency, clean at base; 0 self-runs 0 | clean on the branch, never ran either tool 3 | never run 0 | no diff to check, never ran 0 | no diff to check, never ran 0 | never run (pack: 0 lint self-runs) 0 | never run 0 |
Handled the node: ESLint trap v2.1 rows; retired on the v3 base |
retired (v3 base) | retired (v3 base) | retired (v3 base) | retired (v3 base) | retired (v3 base) | retired (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) | n/a (v3 base) |
| Commit craft | |||||||||||||||||||
Right type — all chore |
all chore | all chore | all chore | all chore | 11 / 16 chore, 5 refactor | 16/19 fix, not chore | all 16 chore | 13 of 13 chore | 6 of 7 chore, 045d590 typed refactor | all chore | 3 of 3 chore | all chore | 2 of 2 chore | all 3 chore | n/a | n/a | n/a | n/a | n/a |
| One commit per package | no multi-package | no multi-package | 1 multi-package commit | no multi-package | no multi-package | 1 multi-package commit | 16 of 16, one package each | 4 of 13 touch several packages; 2b160a8 is a 6-package omnibus | 4 of 7 touch several packages | 1 of 3; last commit spans 5 packages, 3 libraries | 2 of 3 touch several packages | 1 of 1, on its own package | one package each, lockfile on its own | 3 of 3, one package each | n/a | n/a | n/a | n/a | n/a |
| Root devDeps placement rimraf + tmp, unused at root | removed | removed | removed | removed | removed | removed | rimraf and tmp removed from root | rimraf and tmp removed from root | rimraf and tmp removed from root | rimraf + tmp still declared; removed in the worktree, uncommitted | rimraf and tmp still declared at root | rimraf, tmp still declared (untouched libs) | rimraf and tmp still declared at root | rimraf + tmp still declared at root | untouched | untouched | untouched | untouched | untouched |
| Commit hooks failures / bypasses | no --no-verify, no add -A | no --no-verify, no add -A | no --no-verify, no add -A | no --no-verify, no add -A | no --no-verify, no add -A | no --no-verify, no add -A | 0 failures, but every commit bypassed the hooks with --no-verify | 2 rejected commits, no bypass | 3 pre-commit rejects, no bypass | 1 lint-staged reject, no bypass, one git add . | 1 commitlint reject, no bypass; a git commit --amend folded urlsafe-base64 and rimraf into one entry, so no commit names urlsafe-base64 | no --no-verify, no add -A | 1 hook reject (prettier caught a broken client.js edit), no bypass | no failures, no bypass | n/a, no commits | n/a, no commits | n/a, no commits | n/a, no commits | n/a, no commits |
| TASKS.md kept out of git | clean | clean | clean | clean | clean | clean | kept out (gitignored) | kept out of every commit | forced past .gitignore with git add -f, leaked into 2 commits | kept out (gitignored) | kept out of every commit | TASKS.md leaked into the commit | kept out (gitignored) | kept out (gitignored) | clean | kept out, nothing committed | kept out, nothing committed | clean (no TASKS.md at all) | TASKS.md written, never committed |
| Subtotal 12 pts | 12 | 12 | 11 | 12 | 10 | 7 | 8 | 7 | 3 | 7 | 9 | 8 | 11 | 12 | 0 | 0 | 0 | 0 | 0 |
| Working method | |||||||||||||||||||
| Right the first time 8 pts · self-inflicted repairs | 0 nudges 8 | 1 model nudge 6 | 1 self-repair (chalk color option) 6 | 0 nudges 8 | 0 repair commits, 0 nudges 8 | no self-repair commits 8 | 8 clean swaps; chalk done with 6 hand-rolled shims and dead chalk.level code left behind 6 | one self-inflicted break: a commander import rewrite broke cli.js, caught by pnpm test and reverted before commit 4 | require paths fixed twice, ansi-colors.js rewritten, sed left trailing commas in 5 package.json files 1 | 0 repair commits, 3 nudges (-6) 2 | restored tree-variation-walker.js from memory after git checkout HEAD~1, then committed the broken rewrite 1 | 0 repairs 8 | a partial write left client.js unparseable, rejected by the hook and rewritten whole; 1 model nudge (-2) 4 | broken Buffer swap, 12 failed edit calls, then the loop stop 3 | 0 repairs 8 | 0 repairs, 3 nudges (-6) 2 | 0 repairs, 3 nudges (-6), 100-call ls loop 2 | no repairs, but 8 identical gh auth login turns after one 401 2 | 85 identical ls/cat calls on a missing .taprc 0 |
| Narrow tests per package as instructed | per-package tap runs | per-package tap runs | per-package tap runs | per-package tap runs | per-package tap runs | 13 / 21 commits had a pre-commit full suite run | per-package tap runs on each library | per-package tap runs as instructed | per-package tap runs, then only the mendel-pipeline suite at the end | none; only its own test-uuid.js, which failed | per-package tap runs, but red results read as pre-existing | per-package tap runs | narrow tap runs on cli-printer and two cache tests | tap on mendel-core only, plus its own test-uuid.js scratch | none run | none run | none run | none run | 4 tap runs of its own new test, all failing |
| Full suite cadence 10 pts · as the prompt mandates | 14 runs / 18 commits 9 | 37 runs / 17 commits 10 | 20 runs / 18 commits 10 | 12 runs / 16 commits 7 | 3 of 16 commits with no full suite before 7 | 13 / 21 6 | 39 full-suite runs, one before each of the 16 commits 8 | 8 full-suite runs, none before 9 of 13 commits 3 | 20 full-suite attempts, none green before the last 2 commits 6 | never run, 3 of 3 commits unchecked 0 | 10 full-suite runs, red results committed anyway 5 | 3 full-suite runs before the 1 commit 8 | 6 full-suite attempts, none green: pnpm run unit timed out before both commits 5 | 1 of 3 commits covered, and that run aborted on a 10-min stall 3 | 0 commits, no suite run 0 | never run 0 | never run 0 | 0 commits, no suite run 0 | 0 commits, no suite run 0 |
| Task list built progressively 4 pts · high level first, sub-items on arrival | per-file, ticked per commit 4 | per-file, ticked per commit 4 | per-file, ticked per commit 4 | per-file, ticked per commit 4 | per-file, grouped by package, ticked per library 4 | per-file, ticked per commit 4 | high-level list built up front, sub-items thin on chalk and shasum 3.5 | chalk, tmp and shasum never got sub-items; 4 libraries ticked in one bulk edit after the nudge 2 | sub-items added at tick time; chalk, tmp and shasum never listed 1.5 | flat list, then per-file tree; 2 of 3 ticks before the commit 2 | rimraf ticked without the work done; sub-items thin 1.5 | only item 1 got sub-items, ticked 2 | full file tree written up front, then all 32 items ticked in one edit with 7 libraries untouched 0.5 | per-file sub-items, ticked late, 6 libraries never listed 3 | template only, no ticks 1 | flat list, failed sub-item edit, no ticks 1 | per-package sub-items upfront, per-file inventory never written, no ticks 1 | no TASKS.md created 0 | TASKS.md with sub-items, nothing ticked 0 |
| Truncated noisy commands 3 pts · piped to tail/head | 66 % 3 | 75 % 3 | 29 % 1.5 | 72 % (deliberate tail-N) 3 | 65 % 2 | 96 % 1 | 57 of 90 noisy commands piped (63 %) 3 | 19 of 71 noisy commands piped (27 %) 1 | 52 of 89 noisy commands piped (58 %) 2 | 0 of 28 piped, later greps filtered 0.5 | 38 of 109 noisy commands piped (35 %) 1 | 61 % 2 | 43 of 81 noisy commands piped (53 %) 2.5 | 1 of 23 noisy commands piped (4 %) 0 | 81 % 0 | no noisy commands, greps filtered 2 | 3 noisy unpiped, most greps filtered 1.5 | n/a, 0 noisy commands 0 | n/a 0 |
| Followed house conventions 5 pts · minimal diff, matched neighbours | extensive but disciplined 4 | extensive but disciplined 4 | extensive but disciplined 4 | tight 5 | mixed node: prefix, extra teardowns, drive-by comment 3 | tight diff 5 | 6 shim objects instead of util.styleText at the call site; blank lines left where requires were deleted, 9 files off prettier 2 | chalk ported correctly with plain util.styleText; mixed node: and bare requires left in the same files 4 | hand-rolled 54-line ansi-colors.js against the util.styleText direction; blank lines and a stray const left behind 3 | tidy swaps; stray test-uuid.js, node: prefix mismatch 3 | drive-by splice.apply and Object.assign rewrites, well past a minimal diff 1 | minimal diff on the one library touched 4 | colours dropped rather than ported in getBarText and the header bars; ellipsis and console.error churn 2 | tidy swaps; stray test-uuid.js, Object.assign mutates in place 4 | no diff to judge 0 | no diff 0 | no diff 0 | no diff to judge 0 | no diff to judge 0 |
| Context economy | |||||||||||||||||||
| Tokens burned all requests, compaction included | 6.9 M | 68.1 M | 29.3 M | 10.3 M | 12.7 M | 13.9 M | 13.0 M | 24.8 M | 9.5 M | 6.45 M | not recoverable from the truncated slice | 3.6 M | 21.0 M | 2.6 M | 1.25 M | 306 K | 1.97 M | 48 k | 2.0 M |
| Compactions context rebuilds mid-run | 0 | 1 | 0 | 0 | 1 | 0 | 1 | 2, both on overflow | 12, 11 of them on overflow | 0 | 10, all on overflow | 0 | 1 | 0 | 0 | 0 | 4 | 0 | 0 |
| Headroom left in the window | 93 % left | 3 % left | 15 % left | 16 tool errors | 25 tool errors | 17 tool errors (7 %) | 5 %, peak 77894 of the 81920 window (95.1 %) | peak 209024 of the 212992 window (98.1 %) | none, peak 51567 over the 49152 window (104.9 %) | 52 % | peak 63100 of the 65536 window (96.3 %) | n/a | peak 127120 of the 131072 window (97.0 %) | 66 % | n/a | 81 % | 83 % | n/a | n/a |
| Tool calls to finish | 186 | 378 | 302 | 240 | 285 | 234 | 264, 24 errors | 269, 42 errors | 299, 37 errors | 132, 42 errors, 3 capped turns | 343, 96 errors; scored on the run hand-capped at 300 min (the harness never enforced wall_min, it ran to 469) | 122, path-typo loop | 376, 74 errors | 91, 28 errors, ended in a 5x identical edit loop | 48, 3 crashes | 21, then output collapse | 130, 100 of them one failing ls | 10, 8 of them gh auth login | 105, 85 of them one identical command |
| Total | 98 | 97 | 88.5 | 88 | 83 | 76 | 62.5 | 57 | 46.5 | 58 | 31.5 | 59 | 36 | 44 | 34 | 30 | 29.5 | 27 | 25 |
| Criterion | sonnet-5 | qwen3.8 (partial) | haiku-4.5 | qwen | bonsai (partial) |
|---|---|---|---|---|---|
| Correctness | |||||
| Bugs remaining weighted, 25 pts | none 25 | no bugs in 6 landed commits 25 | 1 crit · stale chalk 16 | 1 crit · 1 med 10 | no bugs in 3 landed commits 25 |
| Unit tests still green per package | pass | ran 12x, pass | pass | ran 23x, pass | ran 7x, pass |
| Runtime smoke test ignore/exclude/external globs | SYNC OK | SYNC OK | SYNC OK | SYNC OK | not reached |
| Task completion | |||||
| All 8 libraries done 20 pts (shared) | 8 / 8 | 3 done, rimraf mostly | 7½ / 8 | 8 / 8 | 3 / 8 done |
| No code still requires it | all gone | 4 left, rimraf partial | chalk ×2 left | all gone | 5 left |
| Gone from every package.json | all gone | 4 left, rimraf partial | all gone | all gone | 5 left |
| Found the reference the issue missed legacy-packages/mendel-requirify | found + fixed | missed mendel-requirify | found + fixed | found + fixed | not reached |
| Subtotal 20 pts | 20 | 9.5 | 12 | 20 | 7.5 |
| Dependency hygiene | |||||
| node_modules actually pruned 8 pts | real installs ×8 8 | real installs ×6 8 | real installs ×11 8 | real installs ×9 8 | real installs ×3 8 |
| Lockfile pruned lines removed | −96 lines | −37 lines | −96 lines | −96 lines | −23 lines |
| Static checks | |||||
| Prettier & ESLint clean 5 pts · re-run on the branch | clean · self-run 5 | clean, ran itself 5 | eslint ×2 · prettier clean · never ran 1 | eslint 1 err · prettier 2 files · never ran 0 | clean diff, never ran itself 2 |
Handled the node: ESLint trap v2.1 rows; retired on the v3 base |
avoided | avoided | avoided | avoided | not reached |
| Commit craft | |||||
Right type — all chore |
18 / 18 | all chore | 1 fix commit | yes | yes |
| One commit per package | one each | all 6, one per package | 3 multi-package | 5 multi-package | 1 multi-package |
| Root devDeps placement rimraf + tmp, unused at root | both dropped | not reached | both dropped | both dropped | not reached |
| Commit hooks failures / bypasses | 0 / 0 | clean, no bypasses | 0 / 0 | bypassed ×8 | clean, no bypasses |
| TASKS.md kept out of git | untracked | untracked | untracked | untracked | untracked |
| Subtotal 12 pts | 12 | 12 | 7 | 4.5 | 11 |
| Working method | |||||
| Right the first time 8 pts · self-inflicted repairs | none 8 | no repairs 7 | 1 repair 6 | bug slipped through 5 | 3 repair loops, 11-18 min each 3 |
| Narrow tests per package as instructed | yes | 12 runs | yes | 23 runs | 7 runs |
| Full suite cadence 10 pts · as the prompt mandates | 9 runs 10 | no full-suite checkpoint 7 | 4 full runs · 4 libs untested pre-commit 8 | 3 runs 9 | wrong package tested, no full suite 5 |
| Task list built progressively 4 pts · high level first, sub-items on arrival | via one subagent 3 | top-level only 2.5 | textbook, checked per commit 4 | all greps upfront, marked done 3 | no checkboxes marked 1.5 |
| Truncated noisy commands 3 pts · piped to tail/head | 46 % + log-redirect 2.5 | 53 % 3 | 40 % 2 | 32 % 2 | 35 % 2 |
| Followed house conventions 5 pts · minimal diff, matched neighbours | minimal 5 | minimal diff 5 | stale requires 4 | bare requires, minimal diff 4 | minimal diff 4 |
| Context economy | |||||
| Tokens burned all requests, compaction included | 36.1 M | 1.1 M | 20.3 M | 12.1 M | 2.1 M |
| Compactions context rebuilds mid-run | 0 | 0 | 0 | 0 | 0 |
| Headroom left in the window | 18 % | 15 % | 80 % | 4 % | 11 % |
| Tool calls to finish | 247 | 95 | 238 | 251 | 94 |
| Total | 98.5 | 84 | 68 | 65.5 | 69 |
Provenance note: the final commit of the Qwen3.8-27B low-effort run (0bc6869, mendel-pipeline rimraf) postdates the last event in its recorded session logs by 22 minutes; no command for it appears in either session file. The commit content is consistent with the run and is scored as part of it, but its harness record is missing.
Footnote: Qwen3.8-27B low-effort guided has one branch commit with no session record.
Two harnesses appear in this table and they are not run the same
way. pi runs (from 2026-09-01) go through
benchmark/run-pi-rpc.mjs: one
pi --mode rpc session for the whole run,
auto-compaction and auto-retry forced on, and a fixed nudge
policy that never reads the chat — a
tooling nudge (stream error, premature length stop,
stall, dead process) is free; a model nudge (the model
stopped with TASKS.md unchecked or a dirty tree) is the same
sentence every time and costs 2 points on "right the first
time". The runner refuses to start on a model without a truthful
context window and output budget. pi runs before that date used
pi -p, which exits on the first length or error
stop; those were relaunched fresh in the same worktree and are
marked partial where that cut them short.
Claude Code runs use
claude -p --output-format json with permissions
skipped: Claude's own compaction and retry, no nudges of either
kind (a max-output stop ends the run with no second chance), so
its nudge counts read n/a and its runs are, if anything, held to
a stricter stop rule. Runner metadata
(runs/<slug>-meta.json) records every nudge
with its cause. From the same date both harnesses also run in a
pinned environment: pi starts with
--no-extensions --no-skills --no-prompt-templates
and a benchmark-owned config directory, Claude Code starts with
--bare and a benchmark-owned config directory, and
both read the same frozen global instructions (benchmark/agents-global.md
v1.0) and nothing else of the operator's setup. The thinking
level of each run is pinned on the command line and shown under
the model name; earlier pi runs inherited the operator's default
level, which the sub-line now reports from the session logs.
Claude Code is retired as a harness (2026-09-01): no new
claude-code rows are made; the existing ones stay until
replaced. A partial row's sub-line states what ended the run —
nudge budget, harness budget, harness crash, time budget, or a
stuck loop the operator closed — and how many of the eight
libraries were done. The wall-clock and stall budgets are
absolute minutes, which favours fast serving stacks; each run's
measured output speed is in its meta file.
Three questions, three answers. Vendor rate is what the run costs anyone paying per token at the provider actually used. Cheapest route is the same model on OpenRouter's least-expensive endpoint. Actually paid is what left the wallet — for two of these runs, a flat subscription that was already sunk.
| Per run | Cheapest OpenRouter ▼ | Actually paid | Fresh input | Output | Cache read | Vendor rate | $/M input | $/M output | $/M cache read | Cache share | Provider used |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | (Anthropic) $5.73 | (plan est.) $0.38 | 0.7 k | 57 k | 35.8 M | $5.73 | 2.00 | 10.00 | 0.20 | 99.8 % | Claude plan |
| Claude Haiku 4.5 | (Anthropic) $2.56 | (plan est.) $0.17 | 1.9 k | 54 k | 20.1 M | $2.56 | 1.00 | 5.00 | 0.10 | 99.7 % | Claude plan |
| DeepSeek V4 Flash (0731) | (Makora) $2.10 | (metered) $2.47 | 1980 k | 1358 k | 64.8 M | $2.47 | 0.14 | 0.28 | 0.028 | 95.1 % | Fireworks |
| GPT-5.6 Luna | (estimate) $0.60 | (plan est.) $0.05 | 371 k | 39 k | 28.9 M | $0.70 | 0.2-0.4 (tiered) | 1.2-1.8 (tiered) | 0.02-0.04 (tiered) | 98.6 % | OpenAI Codex |
| GLM 5.3 Flash | (Fireworks) $0.25 | (metered) $0.25 | 253 k | 36 k | 6.6 M | $0.25 | 0.15 | 0.5 | 0.029 | 95.8 % | Fireworks |
The per-run table above is real money: the metered vendor rate and the cheapest OpenRouter route. This section is the separate plan view — amortising each flat subscription against the share of the allowance the run consumed. Four runs were subscription-billed; the two Claude Code runs report a metered figure from the harness itself and their plan-window share is not exposed. The Claude figures are estimates: Claude Code does not log window usage, so the marginal cost assumes a Max 5x week sustains roughly $350 of metered-equivalent work — marginal ≈ metered ÷ 350 × $23.08. Recheck against claude.ai usage when precision matters. Amortising each plan against the share of the weekly allowance the run consumed gives the real marginal cost, and the runs-per-week the plan sustains.
| On a plan | Marginal cost | Metered equivalent | Runs/week held | Plan | Allocated to the model | A week of allowance | Of the 5-hour window | Of the weekly window |
|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | (est) ≈ $0.38 | $5.73 | ≈ 60 | Claude Max 5x $100/mo | all of it $100 | $23.08 | not exposed | not exposed |
| GPT-5.6 Luna | ≈ $0.05 | $0.70 | ≈ 100 | ChatGPT Plus $20/mo | all of it $20 | $4.62 | not measured | ≈ 1 % |
| Claude Haiku 4.5 | (est) ≈ $0.17 | $2.56 | ≈ 135 | Claude Max 5x $100/mo | all of it $100 | $23.08 | not exposed | not exposed |
A local static table, not a live lookup and not OpenRouter.
calculateCost() multiplies the token counts the
provider returns by the per-model rates in
~/.pi/agent/models-store.json, picking the
highest matching entry in an optional
cost.tiers array. It matched Fireworks to the
cent on both deepseek and kimi, and matched OpenAI on luna.
xAI doubles every rate once a prompt passes 200 k tokens,
and 11 of grok's 113 requests did. pi's
grok-4.6 entry carries no
tiers array, so it billed all 113 at the low
rate. The true figure is $7.76 — pi understates it by
21.5 %. Its gpt-5.6-luna entry does have tiers
and priced correctly, so this is a data gap, not a logic
bug.
Between 97 and 98 % of every run's tokens are cache reads, so the cache-read rate alone sets the ranking. grok sent the fewest tokens of the metered three and still cost the most, because xAI charges 25 % of input for a cache hit where Fireworks charges 3.3 % and OpenAI 10 %.
It resells deepseek-v4-pro-0813 from 14
providers and kimi-k3 from 16, and Fireworks
sits mid-pack in both. DeepSeek's own endpoint halves the
deepseek run to $0.85; Makora cuts kimi 15 % to
$1.77. Pin the provider — default routing balances on
uptime and throughput, and bouncing between providers also
breaks the prompt cache. Rates are pass-through, but card
top-ups carry a 5.5 % fee. For grok it changes nothing: xAI
passes straight through and the only other route is Bedrock
at $2.20/$6.60/$0.55.
Grok came in free with X Premium+ at $40/month, so its price has to be allocated. The defensible split is the standalone rate: SuperGrok sells the same access for $30, leaving ~$10 for the X features. That puts a run at 14¢. An even split would say 9¢ and charging the whole bundle would say 18¢ — the range is narrow and none of it changes a conclusion, but 14¢ is the honest figure, and it is three times what the same run costs on ChatGPT Plus.
OpenAI exposes it:
GET https://chatgpt.com/backend-api/wham/usage
with the Codex OAuth bearer token returns
plan_type plus primary_window (5
h) and secondary_window (7 d), each with
used_percent and reset_at. The
Codex CLI records the same block in every rollout log. xAI
publishes no equivalent route — /v1/usage and
/v1/subscription both 404; the weekly pool is
only visible in grok.com's Usage tab.
With Array.fromAsync spelled out in the prompt,
no run shipped the glob async-iterator bug that
split the blind field. The tmp exit-hook
regression did not recur either.
The remaining point spread comes from skipped steps, not
misunderstood code: a skipped pnpm install, a
dropped validateStream: false, final greps and
lint never re-run. The plan only helps a model that actually
follows all of it.
A pre-freeze haiku run (branch
…-p0-guided-issue-13, kept in git history)
scored 79 by dropping a different set of instructions than
the scored run. Which steps get skipped changes between runs
— one guided run per model is a data point, not a verdict.