Controlled bake-off · irae/mendel · issue #13 · guided run

Dependency purge, guided: the structured-plan run

The same task as the blind run, but the prompt now hands every model a numbered plan and names the traps. What is left to measure is instruction-following — how far a structured plan lifts the models that stumbled blind. Scores compare only against the same model's blind run, never across the two tests.

Base tag benchmark-guided-base (per-row base_commit) Branches <model>-guided-issue-13 Libraries 8 Prompt prompt-guided.txt Rubric unchanged

← Blind report  · 

Scoreboard

Ranked by the weighted criteria in the second table. Cost is the metered vendor rate with what was really paid beneath it; local models are free. Peak context is the largest single request, against each model's own window.

     

Prompt v3.0

# Model Score Cost Wall Tokens Peak ctx Window Commits Bugs
1 GLM 5.3 Flashpi · fireworks · think high
98
9 tool errors
$0.25paid $0.25 24 min 6.9 M 72 k 7 % 18 0
2 DeepSeek V4 Flash (0731)pi · fireworks · think high
97
1 model nudge · 7 tool errors · 1 compaction · repetition loop 0.02 on text
$2.10paid $2.47 249 min 68.1 M 996 k 97 % 16 0
3 GPT-5.6 Lunapi · openai-codex · think high
89
$0.60plan ≈$0.05 40 min 29.3 M 232 k 85 % 18 2
4 Claude Sonnet 4.5pi · anthropic · think high
88
16 tool errors
$3.92paid $3.92 32 min 10.3 M 93 k 9 % 16 3
5 Qwen3.6 35B-A3Bpi · llama · local · think high
83
peak_context corrected 2026-09-04: the earlier 77849 was the value after the compaction. · 25 tool errors · 1 compaction
$0plan ≈$0.00 92 min 12.7 M 94 k 79 % 16 5
6 Claude Haiku 4.5pi · anthropic · think high
76
17 tool errors
paid $1.73 27 min 13.9 M 108 k 54 % 19 3
7 Qwen3.6 35B-A3Bpi · llama · local · think off
63
24 tool errors · 1 compaction
$0plan ≈$0.00 89 min 13.0 M 78 k 95.1 % 16 8
8 Gemma 4 26B-A4Bpi · llama · local · think highpartial · 7 of 8
57
7/8 done · stopped without finishing
$0plan ≈$0.00 115 min 24.8 M 209 k 98.1 % 13 16
9 Qwen3.6 35B-A3Bpi · llama · local · think off
47
2 model nudges · 37 tool errors · 12 compactions
$0plan ≈$0.00 96 min 9.5 M 52 k 104.9 % 7 16
10 Gemma-4-12B (llama.cpp, off)pi · llama · local · think offpartial · 3 of 8 · model-nudge budget spent
38
raw 58 · 3/8 done · model-nudge budget spent · repetition loop 0.02 on text
$0plan ≈$0.00 98 min 6.5 M 125 k 48 % 3 6
11 Ternary Bonsai 27B (PrismML GGUF)pi · llama · local · think highpartial · 3 of 8 · hit the 300 min time budget
32
3/8 done · hit the 300 min time budget
$0plan ≈$0.00 300 min 0.0 M 63 k 96.28 % 3 16
12 Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 1 of 8 · hit the 300 min time budget
13
raw 59 · 1/8 done · hit the 300 min time budget
$0plan ≈$0.00 300 min 3.6 M 46 k 80 % 1 1
13 bonsai-prism (f16 KV)pi · llama · local · think highpartial · 1 of 8
13
raw 36 · 1/8 done · stopped without finishing
$0plan ≈$0.00 195 min 21.0 M 127 k 97 % 2 10
Gemma 4 26B-A4Bpi · llama · local · think offpartial · 2 of 8 · repetition loop, stopped by the runner
25
invalid — live loop stop: repetition_loop on 5 identical edit tool calls, 18:12:06Z to 18:12:47Z; 2 of 8 libraries landed · repetition loop 0.22 on tool call
$0plan ≈$0.00 20 min 2.6 M 73 k 34.15 % 3 11
Qwen3.8 27B (MLX 4-bit)pi · mlx · local · think lowpartial · 0 of 8 · tooling-nudge budget spent
0
invalid — three Metal OOM server crashes, tooling budget exhausted; zero commits
$0plan ≈$0.00 261 min 1.3 M 30 k 0 3
Gemma 4 12Bpi · lmstudio · local · think highpartial · 0 of 8 · model-nudge budget spent
0
invalid — repetition loop on the LM Studio MLX thinking-on entry google/gemma-4-12b, pre-fix chat template (research run 2); zero commits
$0plan ≈$0.00 46 min 0.3 M 30 k 19 % 0 4
Gemma 4 12B (low reasoning)pi · lmstudio · local · think lowpartial · 0 of 8 · model-nudge budget spent
0
invalid — repetition loop on the LM Studio MLX thinking-on entry google/gemma-4-12b, pre-fix chat template (research run 2); zero commits · repetition loop 0.02 on tool call
$0plan ≈$0.00 99 min 2.0 M 45 k 28 % 0 6
Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think offpartial · 0 of 8 · tooling-nudge budget spent
0
invalid — gh token invalid on the host (HTTP 401), model looped on interactive gh auth login, tooling budget exhausted; zero commits
$0plan ≈$0.00 84 min 0.0 M 5.3 k 9 % 0 3
Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think offpartial · 0 of 8 · stopped
0
invalid — operator stop: non-terminating identical-command loop (one bash ls/cat of a missing .taprc, 85 times in a row); zero commits · repetition loop 0.02 on tool call
$0plan ≈$0.00 187 min 2.0 M 27 k 47 % 0 3

Prompt v2.1

# Model Score Cost Wall Tokens Peak ctx Window Commits Bugs
1 Claude Sonnet 5claude-code · anthropic
99
1 tool error
$5.73plan ≈$0.38 31 min 36.1 M 180 k † 18 % 18 0
2 Qwen3.8 27B (MLX 4-bit)pi · mlx · local · think lowpartial · 6 of 8 · harness crash
75
raw 84 · 6/8 done · harness crash
$0plan ≈$0.00 154 min 1.1 M 23 k 85 % 6 4
3 Claude Haiku 4.5claude-code · anthropic
68
6 tool errors
$2.56plan ≈$0.17 18 min 20.3 M 160 k † 80 % 12 4
4 Qwen3.6 35B-A3Bpi · llama · local · think high
66
18 tool errors
$0plan ≈$0.00 76 min 12.1 M 94 k 96 % 8 8
5 Ternary Bonsai 27B (MLX 2-bit)pi · mlx · local · think highpartial · 3 of 8 · stuck
38
raw 69 · 3/8 done · stuck
$0plan ≈$0.00 230 min 2.1 M 51 k 89 % 3 5

Bugs are weighted: critical 3, medium 2, minor 1. Every run uses the frozen prompt-guided.txt. Claude runs on a subscription plan show computed token cost, not a metered bill.

Peak ctx is the largest context one turn occupied, the response of that turn included, across all compaction cycles. count-tool-calls.mjs reads it from the session log of the scored run. † marks the rows from the retired claude-code harness: that harness writes its log in another shape, which the counter does not read, so those cells keep the older reading, which counts the prompt of the largest turn without the response.

Criteria

Every row was checked against the branch or the session log, not against what the model said it did. Weights in the left column; the best cell in each row is highlighted.

Prompt v3.0

Criterion glm-flash ds-flash luna sonnet-4.5 qwen haiku-4.5 qwen gemma (partial) qwen Gemma-4-12B (llama.cpp, off) (partial) bonsai-prism (partial) bonsai (partial) bonsai-prism (f16 KV) (partial) gemma (partial) qwen3.8 (partial) gemma-12b (partial) gemma-12b-low (partial) bonsai (partial) bonsai (partial)
Correctness
Bugs remaining weighted, 25 pts clean 25 clean 25 exit-hook regression (trap C) 19 trap A: TypeError 16 trap A: TypeError 16 TypeError: glob(...).then is not a function (trap A) 16 trap A still throws: glob(...).then is not a function 7 trap A still throws: glob(...).then is not a function 10 trap A still throws: glob(...).then is not a function 4 none in the 3 landed replacements 25 2 regressions from a memory rewrite: tree-variation-walker 3/3, config.js #10 and #7 0 no bugs in the 1 landed commit 25 rimraf devDep removed while test/helpers/index.js still requires it; root rimraf and tmp never touched 1 urlsafe-base64 swap left broken, tests guarded around it 4 none 25 none, no code landed 25 none, no code landed 25 none, no code changed (vacuous) 25 none, no code changed (vacuous) 25
Unit tests still green per package 680/680, 1 skip 710/709, lerna ok clean · pipeline flake confirmed clean standalone (260/260) 674/674, 1 skip all package suites pass (core 71/72, 1 skip) clean (674/674, 1 skip) suites green, full suite re-run before every commit full suite green at HEAD, per-package tap green mendel-pipeline package suite green; root suite never green after the chalk edits core 71/71, config 41/41 (one documented flake) mendel-core and mendel-config fail at HEAD, green on the base not run mendel-pipeline cli-printer 15/15 green; the same package loses its test helper on a fresh install not re-run; worktree left mid-edit not run not run not run not run not run
Runtime smoke test ignore/exclude/external globs SYNC OK SYNC OK SYNC OK THREW: TypeError THREW: TypeError THREW: TypeError apply-extra-options.js:6 THREW TypeError apply-extra-options.js:6 THREW TypeError apply-extra-options.js:6 THREW TypeError vacuous pass, apply-extra-options.js untouched SYNC OK, glob never touched not touched SYNC OK, but glob was never touched (vacuous pass) vacuous pass, apply-extra-options.js untouched not touched not touched not touched not touched (pack: SYNC OK on base code) not touched
Task completion
All 8 libraries done 20 pts (shared) 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 8 / 8 done 7 / 8 done, rimraf incomplete 8 / 8 done 3 / 8 done 3 / 8 (uuid, xtend, urlsafe-base64) 1 / 8 (uuid only) 1 / 8 (chalk) 2 / 8 done 0 / 8 0 / 8 0 / 8 0 / 8 0 / 8
No code still requires it none left none left clean none left none left clean 0 stale requires 2 stale rimraf requires in mendel-requirify tests 0 stale requires 5 libraries left, rimraf half done rimraf ticked done, still required in 2 requirify tests 27 files left 22 stale requires across 7 libraries 23 stale requires in 12 packages 28 files left 28 sites left 28 sites left 28 files left 28 files left
Gone from every package.json clean clean none left clean clean none left 1 stale entry, rimraf in legacy mendel-requirify 1 stale entry, mendel-requirify package.json:26 0 stale entries, root cleared 5 left, root rimraf + tmp still declared 13 stale entries, root rimraf and tmp among them 16 files left 16 stale entries, root rimraf and tmp among them 15 entries left, root rimraf + tmp still declared 18 files left 18 left 18 left 18 files left 18 files left
Found the reference the issue missed legacy-packages/mendel-requirify found + fixed found + fixed found + fixed found + fixed found + fixed found + fixed found, but two pnpm remove tries failed and the model gave up seen in its own grep, then skipped (trap B) found and removed (trap B) grepped and edited it, never committed not found (trap B) not reached not found (trap B) never found not reached grepped it, never acted found it in the plan, never acted not reached not reached
Subtotal 20 pts 20 20 20 20 20 20 17 16 20 7.5 6 2 2 4 0 0 0 0 0
Dependency hygiene
node_modules actually pruned 8 pts real install 8 real install 8 real install 8 real remove, verified 8 pnpm remove per package, reinstall verified 8 install unverified · only root deps via `pnpm remove -w` 4 real pnpm install, store pruned 8 pnpm used at root, but tmp cut from 3 package.json by hand 5 pnpm used, but tmp and shasum cut with sed; pnpm-lock.yaml still lists both 3 pnpm remove x4, store pruned 8 pnpm used for uuid only, later edits by hand 4 not pruned 0 real pnpm install, lockfile regenerated 8 real pnpm install, store pruned 8 not pruned 0 never ran pnpm 0 never ran pnpm 0 not pruned 0 not pruned 0
Lockfile pruned lines removed +199/−106 lines +134/−133 lines +4/−197 lines +123/−101 lines +22/-211 lines +91/−99 lines 211 lines removed, 22 added 206 lines removed, 26 added; --frozen-lockfile fails at HEAD 183 lines removed, 47 added; --frozen-lockfile fails at HEAD 121 lines 107 lines removed, 1 added; xtend, urlsafe-base64 and rimraf still listed 0 lines 12 lines removed, 0 added; chalk and pipeline rimraf only 20 lines removed, 5 added 0 lines 0 lines 0 lines 0 lines 0 lines
Static checks
Prettier & ESLint clean 5 pts · re-run on the branch clean, self-checked 2x 5 clean, self-checked 14x 5 clean, self-checked 4x 5 clean, self-checked 5 clean, self-checked via pnpm test (static) 5 clean, self-checked after last commit 5 prettier fails on 9 files clean at base; 0 self-runs, all 16 commits --no-verify so no hook caught it 0 clean on re-run, never ran either tool itself 5 clean on re-run, never ran either tool itself 3 clean on re-run, never ran itself 3 prettier and eslint clean on re-run, never self-run 3 prettier warn (mendel-config/legacy) 0 eslint fails at the tip on a self-made implicit dependency, clean at base; 0 self-runs 0 clean on the branch, never ran either tool 3 never run 0 no diff to check, never ran 0 no diff to check, never ran 0 never run (pack: 0 lint self-runs) 0 never run 0
Handled the node: ESLint trap v2.1 rows; retired on the v3 base retired (v3 base) retired (v3 base) retired (v3 base) retired (v3 base) retired (v3 base) retired (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base) n/a (v3 base)
Commit craft
Right type — all chore all chore all chore all chore all chore 11 / 16 chore, 5 refactor 16/19 fix, not chore all 16 chore 13 of 13 chore 6 of 7 chore, 045d590 typed refactor all chore 3 of 3 chore all chore 2 of 2 chore all 3 chore n/a n/a n/a n/a n/a
One commit per package no multi-package no multi-package 1 multi-package commit no multi-package no multi-package 1 multi-package commit 16 of 16, one package each 4 of 13 touch several packages; 2b160a8 is a 6-package omnibus 4 of 7 touch several packages 1 of 3; last commit spans 5 packages, 3 libraries 2 of 3 touch several packages 1 of 1, on its own package one package each, lockfile on its own 3 of 3, one package each n/a n/a n/a n/a n/a
Root devDeps placement rimraf + tmp, unused at root removed removed removed removed removed removed rimraf and tmp removed from root rimraf and tmp removed from root rimraf and tmp removed from root rimraf + tmp still declared; removed in the worktree, uncommitted rimraf and tmp still declared at root rimraf, tmp still declared (untouched libs) rimraf and tmp still declared at root rimraf + tmp still declared at root untouched untouched untouched untouched untouched
Commit hooks failures / bypasses no --no-verify, no add -A no --no-verify, no add -A no --no-verify, no add -A no --no-verify, no add -A no --no-verify, no add -A no --no-verify, no add -A 0 failures, but every commit bypassed the hooks with --no-verify 2 rejected commits, no bypass 3 pre-commit rejects, no bypass 1 lint-staged reject, no bypass, one git add . 1 commitlint reject, no bypass; a git commit --amend folded urlsafe-base64 and rimraf into one entry, so no commit names urlsafe-base64 no --no-verify, no add -A 1 hook reject (prettier caught a broken client.js edit), no bypass no failures, no bypass n/a, no commits n/a, no commits n/a, no commits n/a, no commits n/a, no commits
TASKS.md kept out of git clean clean clean clean clean clean kept out (gitignored) kept out of every commit forced past .gitignore with git add -f, leaked into 2 commits kept out (gitignored) kept out of every commit TASKS.md leaked into the commit kept out (gitignored) kept out (gitignored) clean kept out, nothing committed kept out, nothing committed clean (no TASKS.md at all) TASKS.md written, never committed
Subtotal 12 pts 12 12 11 12 10 7 8 7 3 7 9 8 11 12 0 0 0 0 0
Working method
Right the first time 8 pts · self-inflicted repairs 0 nudges 8 1 model nudge 6 1 self-repair (chalk color option) 6 0 nudges 8 0 repair commits, 0 nudges 8 no self-repair commits 8 8 clean swaps; chalk done with 6 hand-rolled shims and dead chalk.level code left behind 6 one self-inflicted break: a commander import rewrite broke cli.js, caught by pnpm test and reverted before commit 4 require paths fixed twice, ansi-colors.js rewritten, sed left trailing commas in 5 package.json files 1 0 repair commits, 3 nudges (-6) 2 restored tree-variation-walker.js from memory after git checkout HEAD~1, then committed the broken rewrite 1 0 repairs 8 a partial write left client.js unparseable, rejected by the hook and rewritten whole; 1 model nudge (-2) 4 broken Buffer swap, 12 failed edit calls, then the loop stop 3 0 repairs 8 0 repairs, 3 nudges (-6) 2 0 repairs, 3 nudges (-6), 100-call ls loop 2 no repairs, but 8 identical gh auth login turns after one 401 2 85 identical ls/cat calls on a missing .taprc 0
Narrow tests per package as instructed per-package tap runs per-package tap runs per-package tap runs per-package tap runs per-package tap runs 13 / 21 commits had a pre-commit full suite run per-package tap runs on each library per-package tap runs as instructed per-package tap runs, then only the mendel-pipeline suite at the end none; only its own test-uuid.js, which failed per-package tap runs, but red results read as pre-existing per-package tap runs narrow tap runs on cli-printer and two cache tests tap on mendel-core only, plus its own test-uuid.js scratch none run none run none run none run 4 tap runs of its own new test, all failing
Full suite cadence 10 pts · as the prompt mandates 14 runs / 18 commits 9 37 runs / 17 commits 10 20 runs / 18 commits 10 12 runs / 16 commits 7 3 of 16 commits with no full suite before 7 13 / 21 6 39 full-suite runs, one before each of the 16 commits 8 8 full-suite runs, none before 9 of 13 commits 3 20 full-suite attempts, none green before the last 2 commits 6 never run, 3 of 3 commits unchecked 0 10 full-suite runs, red results committed anyway 5 3 full-suite runs before the 1 commit 8 6 full-suite attempts, none green: pnpm run unit timed out before both commits 5 1 of 3 commits covered, and that run aborted on a 10-min stall 3 0 commits, no suite run 0 never run 0 never run 0 0 commits, no suite run 0 0 commits, no suite run 0
Task list built progressively 4 pts · high level first, sub-items on arrival per-file, ticked per commit 4 per-file, ticked per commit 4 per-file, ticked per commit 4 per-file, ticked per commit 4 per-file, grouped by package, ticked per library 4 per-file, ticked per commit 4 high-level list built up front, sub-items thin on chalk and shasum 3.5 chalk, tmp and shasum never got sub-items; 4 libraries ticked in one bulk edit after the nudge 2 sub-items added at tick time; chalk, tmp and shasum never listed 1.5 flat list, then per-file tree; 2 of 3 ticks before the commit 2 rimraf ticked without the work done; sub-items thin 1.5 only item 1 got sub-items, ticked 2 full file tree written up front, then all 32 items ticked in one edit with 7 libraries untouched 0.5 per-file sub-items, ticked late, 6 libraries never listed 3 template only, no ticks 1 flat list, failed sub-item edit, no ticks 1 per-package sub-items upfront, per-file inventory never written, no ticks 1 no TASKS.md created 0 TASKS.md with sub-items, nothing ticked 0
Truncated noisy commands 3 pts · piped to tail/head 66 % 3 75 % 3 29 % 1.5 72 % (deliberate tail-N) 3 65 % 2 96 % 1 57 of 90 noisy commands piped (63 %) 3 19 of 71 noisy commands piped (27 %) 1 52 of 89 noisy commands piped (58 %) 2 0 of 28 piped, later greps filtered 0.5 38 of 109 noisy commands piped (35 %) 1 61 % 2 43 of 81 noisy commands piped (53 %) 2.5 1 of 23 noisy commands piped (4 %) 0 81 % 0 no noisy commands, greps filtered 2 3 noisy unpiped, most greps filtered 1.5 n/a, 0 noisy commands 0 n/a 0
Followed house conventions 5 pts · minimal diff, matched neighbours extensive but disciplined 4 extensive but disciplined 4 extensive but disciplined 4 tight 5 mixed node: prefix, extra teardowns, drive-by comment 3 tight diff 5 6 shim objects instead of util.styleText at the call site; blank lines left where requires were deleted, 9 files off prettier 2 chalk ported correctly with plain util.styleText; mixed node: and bare requires left in the same files 4 hand-rolled 54-line ansi-colors.js against the util.styleText direction; blank lines and a stray const left behind 3 tidy swaps; stray test-uuid.js, node: prefix mismatch 3 drive-by splice.apply and Object.assign rewrites, well past a minimal diff 1 minimal diff on the one library touched 4 colours dropped rather than ported in getBarText and the header bars; ellipsis and console.error churn 2 tidy swaps; stray test-uuid.js, Object.assign mutates in place 4 no diff to judge 0 no diff 0 no diff 0 no diff to judge 0 no diff to judge 0
Context economy
Tokens burned all requests, compaction included 6.9 M 68.1 M 29.3 M 10.3 M 12.7 M 13.9 M 13.0 M 24.8 M 9.5 M 6.45 M not recoverable from the truncated slice 3.6 M 21.0 M 2.6 M 1.25 M 306 K 1.97 M 48 k 2.0 M
Compactions context rebuilds mid-run 0 1 0 0 1 0 1 2, both on overflow 12, 11 of them on overflow 0 10, all on overflow 0 1 0 0 0 4 0 0
Headroom left in the window 93 % left 3 % left 15 % left 16 tool errors 25 tool errors 17 tool errors (7 %) 5 %, peak 77894 of the 81920 window (95.1 %) peak 209024 of the 212992 window (98.1 %) none, peak 51567 over the 49152 window (104.9 %) 52 % peak 63100 of the 65536 window (96.3 %) n/a peak 127120 of the 131072 window (97.0 %) 66 % n/a 81 % 83 % n/a n/a
Tool calls to finish 186 378 302 240 285 234 264, 24 errors 269, 42 errors 299, 37 errors 132, 42 errors, 3 capped turns 343, 96 errors; scored on the run hand-capped at 300 min (the harness never enforced wall_min, it ran to 469) 122, path-typo loop 376, 74 errors 91, 28 errors, ended in a 5x identical edit loop 48, 3 crashes 21, then output collapse 130, 100 of them one failing ls 10, 8 of them gh auth login 105, 85 of them one identical command
Total 98 97 88.5 88 83 76 62.5 57 46.5 58 31.5 59 36 44 34 30 29.5 27 25

Prompt v2.1

Criterion sonnet-5 qwen3.8 (partial) haiku-4.5 qwen bonsai (partial)
Correctness
Bugs remaining weighted, 25 pts none 25 no bugs in 6 landed commits 25 1 crit · stale chalk 16 1 crit · 1 med 10 no bugs in 3 landed commits 25
Unit tests still green per package pass ran 12x, pass pass ran 23x, pass ran 7x, pass
Runtime smoke test ignore/exclude/external globs SYNC OK SYNC OK SYNC OK SYNC OK not reached
Task completion
All 8 libraries done 20 pts (shared) 8 / 8 3 done, rimraf mostly 7½ / 8 8 / 8 3 / 8 done
No code still requires it all gone 4 left, rimraf partial chalk ×2 left all gone 5 left
Gone from every package.json all gone 4 left, rimraf partial all gone all gone 5 left
Found the reference the issue missed legacy-packages/mendel-requirify found + fixed missed mendel-requirify found + fixed found + fixed not reached
Subtotal 20 pts 20 9.5 12 20 7.5
Dependency hygiene
node_modules actually pruned 8 pts real installs ×8 8 real installs ×6 8 real installs ×11 8 real installs ×9 8 real installs ×3 8
Lockfile pruned lines removed −96 lines −37 lines −96 lines −96 lines −23 lines
Static checks
Prettier & ESLint clean 5 pts · re-run on the branch clean · self-run 5 clean, ran itself 5 eslint ×2 · prettier clean · never ran 1 eslint 1 err · prettier 2 files · never ran 0 clean diff, never ran itself 2
Handled the node: ESLint trap v2.1 rows; retired on the v3 base avoided avoided avoided avoided not reached
Commit craft
Right type — all chore 18 / 18 all chore 1 fix commit yes yes
One commit per package one each all 6, one per package 3 multi-package 5 multi-package 1 multi-package
Root devDeps placement rimraf + tmp, unused at root both dropped not reached both dropped both dropped not reached
Commit hooks failures / bypasses 0 / 0 clean, no bypasses 0 / 0 bypassed ×8 clean, no bypasses
TASKS.md kept out of git untracked untracked untracked untracked untracked
Subtotal 12 pts 12 12 7 4.5 11
Working method
Right the first time 8 pts · self-inflicted repairs none 8 no repairs 7 1 repair 6 bug slipped through 5 3 repair loops, 11-18 min each 3
Narrow tests per package as instructed yes 12 runs yes 23 runs 7 runs
Full suite cadence 10 pts · as the prompt mandates 9 runs 10 no full-suite checkpoint 7 4 full runs · 4 libs untested pre-commit 8 3 runs 9 wrong package tested, no full suite 5
Task list built progressively 4 pts · high level first, sub-items on arrival via one subagent 3 top-level only 2.5 textbook, checked per commit 4 all greps upfront, marked done 3 no checkboxes marked 1.5
Truncated noisy commands 3 pts · piped to tail/head 46 % + log-redirect 2.5 53 % 3 40 % 2 32 % 2 35 % 2
Followed house conventions 5 pts · minimal diff, matched neighbours minimal 5 minimal diff 5 stale requires 4 bare requires, minimal diff 4 minimal diff 4
Context economy
Tokens burned all requests, compaction included 36.1 M 1.1 M 20.3 M 12.1 M 2.1 M
Compactions context rebuilds mid-run 0 0 0 0 0
Headroom left in the window 18 % 15 % 80 % 4 % 11 %
Tool calls to finish 247 95 238 251 94
Total 98.5 84 68 65.5 69

Provenance note: the final commit of the Qwen3.8-27B low-effort run (0bc6869, mendel-pipeline rimraf) postdates the last event in its recorded session logs by 22 minutes; no command for it appears in either session file. The commit content is consistent with the run and is scored as part of it, but its harness record is missing.

Footnote: Qwen3.8-27B low-effort guided has one branch commit with no session record.

Harnesses

Two harnesses appear in this table and they are not run the same way. pi runs (from 2026-09-01) go through benchmark/run-pi-rpc.mjs: one pi --mode rpc session for the whole run, auto-compaction and auto-retry forced on, and a fixed nudge policy that never reads the chat — a tooling nudge (stream error, premature length stop, stall, dead process) is free; a model nudge (the model stopped with TASKS.md unchecked or a dirty tree) is the same sentence every time and costs 2 points on "right the first time". The runner refuses to start on a model without a truthful context window and output budget. pi runs before that date used pi -p, which exits on the first length or error stop; those were relaunched fresh in the same worktree and are marked partial where that cut them short. Claude Code runs use claude -p --output-format json with permissions skipped: Claude's own compaction and retry, no nudges of either kind (a max-output stop ends the run with no second chance), so its nudge counts read n/a and its runs are, if anything, held to a stricter stop rule. Runner metadata (runs/<slug>-meta.json) records every nudge with its cause. From the same date both harnesses also run in a pinned environment: pi starts with --no-extensions --no-skills --no-prompt-templates and a benchmark-owned config directory, Claude Code starts with --bare and a benchmark-owned config directory, and both read the same frozen global instructions (benchmark/agents-global.md v1.0) and nothing else of the operator's setup. The thinking level of each run is pinned on the command line and shown under the model name; earlier pi runs inherited the operator's default level, which the sub-line now reports from the session logs. Claude Code is retired as a harness (2026-09-01): no new claude-code rows are made; the existing ones stay until replaced. A partial row's sub-line states what ended the run — nudge budget, harness budget, harness crash, time budget, or a stuck loop the operator closed — and how many of the eight libraries were done. The wall-clock and stall budgets are absolute minutes, which favours fast serving stacks; each run's measured output speed is in its meta file.

Cost

Three questions, three answers. Vendor rate is what the run costs anyone paying per token at the provider actually used. Cheapest route is the same model on OpenRouter's least-expensive endpoint. Actually paid is what left the wallet — for two of these runs, a flat subscription that was already sunk.

Per run Cheapest OpenRouter Actually paid Fresh input Output Cache read Vendor rate $/M input $/M output $/M cache read Cache share Provider used
Claude Sonnet 5 (Anthropic) $5.73 (plan est.) $0.38 0.7 k 57 k 35.8 M $5.73 2.00 10.00 0.20 99.8 % Claude plan
Claude Haiku 4.5 (Anthropic) $2.56 (plan est.) $0.17 1.9 k 54 k 20.1 M $2.56 1.00 5.00 0.10 99.7 % Claude plan
DeepSeek V4 Flash (0731) (Makora) $2.10 (metered) $2.47 1980 k 1358 k 64.8 M $2.47 0.14 0.28 0.028 95.1 % Fireworks
GPT-5.6 Luna (estimate) $0.60 (plan est.) $0.05 371 k 39 k 28.9 M $0.70 0.2-0.4 (tiered) 1.2-1.8 (tiered) 0.02-0.04 (tiered) 98.6 % OpenAI Codex
GLM 5.3 Flash (Fireworks) $0.25 (metered) $0.25 253 k 36 k 6.6 M $0.25 0.15 0.5 0.029 95.8 % Fireworks

Plan estimated costs

The per-run table above is real money: the metered vendor rate and the cheapest OpenRouter route. This section is the separate plan view — amortising each flat subscription against the share of the allowance the run consumed. Four runs were subscription-billed; the two Claude Code runs report a metered figure from the harness itself and their plan-window share is not exposed. The Claude figures are estimates: Claude Code does not log window usage, so the marginal cost assumes a Max 5x week sustains roughly $350 of metered-equivalent work — marginal ≈ metered ÷ 350 × $23.08. Recheck against claude.ai usage when precision matters. Amortising each plan against the share of the weekly allowance the run consumed gives the real marginal cost, and the runs-per-week the plan sustains.

On a plan Marginal cost Metered equivalent Runs/week held Plan Allocated to the model A week of allowance Of the 5-hour window Of the weekly window
Claude Sonnet 5 (est) ≈ $0.38 $5.73 ≈ 60 Claude Max 5x $100/mo all of it $100 $23.08 not exposed not exposed
GPT-5.6 Luna ≈ $0.05 $0.70 ≈ 100 ChatGPT Plus $20/mo all of it $20 $4.62 not measured ≈ 1 %
Claude Haiku 4.5 (est) ≈ $0.17 $2.56 ≈ 135 Claude Max 5x $100/mo all of it $100 $23.08 not exposed not exposed

How pi prices a run

A local static table, not a live lookup and not OpenRouter. calculateCost() multiplies the token counts the provider returns by the per-model rates in ~/.pi/agent/models-store.json, picking the highest matching entry in an optional cost.tiers array. It matched Fireworks to the cent on both deepseek and kimi, and matched OpenAI on luna.

Where pi is wrong about grok

xAI doubles every rate once a prompt passes 200 k tokens, and 11 of grok's 113 requests did. pi's grok-4.6 entry carries no tiers array, so it billed all 113 at the low rate. The true figure is $7.76 — pi understates it by 21.5 %. Its gpt-5.6-luna entry does have tiers and priced correctly, so this is a data gap, not a logic bug.

Cache reads decide everything

Between 97 and 98 % of every run's tokens are cache reads, so the cache-read rate alone sets the ranking. grok sent the fewest tokens of the metered three and still cost the most, because xAI charges 25 % of input for a cache hit where Fireworks charges 3.3 % and OpenAI 10 %.

OpenRouter is worth it for two of the four

It resells deepseek-v4-pro-0813 from 14 providers and kimi-k3 from 16, and Fireworks sits mid-pack in both. DeepSeek's own endpoint halves the deepseek run to $0.85; Makora cuts kimi 15 % to $1.77. Pin the provider — default routing balances on uptime and throughput, and bouncing between providers also breaks the prompt cache. Rates are pass-through, but card top-ups carry a 5.5 % fee. For grok it changes nothing: xAI passes straight through and the only other route is Bedrock at $2.20/$6.60/$0.55.

Splitting a bundled plan

Grok came in free with X Premium+ at $40/month, so its price has to be allocated. The defensible split is the standalone rate: SuperGrok sells the same access for $30, leaving ~$10 for the X features. That puts a run at 14¢. An even split would say 9¢ and charging the whole bundle would say 18¢ — the range is narrow and none of it changes a conclusion, but 14¢ is the honest figure, and it is three times what the same run costs on ChatGPT Plus.

Reading your own plan usage

OpenAI exposes it: GET https://chatgpt.com/backend-api/wham/usage with the Codex OAuth bearer token returns plan_type plus primary_window (5 h) and secondary_window (7 d), each with used_percent and reset_at. The Codex CLI records the same block in every rollout log. xAI publishes no equivalent route — /v1/usage and /v1/subscription both 404; the weekly pool is only visible in grok.com's Usage tab.

What decided it

Disclosed traps mostly worked

With Array.fromAsync spelled out in the prompt, no run shipped the glob async-iterator bug that split the blind field. The tmp exit-hook regression did not recur either.

Compliance became the score

The remaining point spread comes from skipped steps, not misunderstood code: a skipped pnpm install, a dropped validateStream: false, final greps and lint never re-run. The plan only helps a model that actually follows all of it.

Run-to-run variance is visible

A pre-freeze haiku run (branch …-p0-guided-issue-13, kept in git history) scored 79 by dropping a different set of instructions than the scored run. Which steps get skipped changes between runs — one guided run per model is a data point, not a verdict.