Skip to content

Mendel — M1 Max 32 GB

The agentic tier of the quality flow, after EvalPlus: one real repo task with known traps, scored on a 100-point rubric — from the open-source Mendel project, where the task, the rubric, and the raw results live. Method and house rules: Mendel in the methodology.

Mendel is two tests on the same task. The blind test gives a terse prompt and asks whether the model finds the traps by itself. The guided test hands every model the same structured plan with the traps disclosed, and measures instruction-following. Strong API models run blind only; local and weak models run both, so each pair shows the lift. Scores never compare across the two tests.

The full reports are hosted here, generated from the Mendel data:

The tables below are drawn from the mirrored result files in benchmarks/mendel/ (npm run docs:tables). They show only the current prompt version of each test (blind v1.1, guided v3.0); rows from older prompt versions live in historical and in the hosted reports, one scoreboard per version.

Each row names the serving path and the harness window. The KV cache type of a run is in its config note in the Mendel report.

One Qwen3.6-35B-A3B score below, the guided 83 at thinking high, is pending a re-run at low priority: it ran on a 120K harness window, and at wired limit 25000 the model serves -c 98304. The score stays as a record of what the model did; the other Qwen3.6 rows ran on windows the machine serves today.

Local models — blind test

configscoreworst defect
Qwen3.8-27B Q4_K_Mbartowski,llama-servermtp/3f16xhigh93/100medium
Qwen3.8-27B Q4_K_Mbartowski,llama-servermtp/3f16medium87/100minor
Qwen3.8-27B IQ3_S-mtpISTA-DASLab,llama-serverf16xhigh80.5/100critical
Qwen3.8-27B IQ3_S-mtpISTA-DASLab,llama-servermtp/3f16medium76.5/100critical
Qwen3.8-27B Q4_K_Mbartowski,llama-servermtp/3f16medium76/100critical
Qwen3.8-27B IQ3_S-mtpISTA-DASLab,llama-serverf16low66/100 (partial)critical
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-servermtp/3q8_0on63/100critical
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-servermtp/3q8_0off50.5/100critical
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-serverf16on50/100critical
Gemma-4-26B-A4B UD-Q4_K_XLunsloth,llama-servermtp/2f16on47.5/100critical
Ternary-Bonsai-27B 2-bitprism-ml,mlx_lm.serverf16on37.5/100 (partial)medium
Qwen3.8-27B AD-IQ3_SAtomicChat,llama-servermtp/3f16medium37.5/100 (partial)medium
Qwen3.8-27B 4-bitmlx-community,mlx_lm.serverf16low12.5/100 (partial)minor
Ternary-Bonsai-27B Q2_g64prism-ml,prism-llamaq4_0+biason12.5/100critical
Gemma-4-26B-A4B UD-Q4_K_XLunsloth,llama-servermtp/2f16off12.5/100 (partial)critical

Run notes for the two partials are in the comparison page's Mendel section: both closed early on mlx_lm.server failures or the time budget, not on the rubric.

Three Gemma-4-12B runs are marked invalid and are not listed above. They ran the retired LM Studio entry google/gemma-4-12b with thinking on and its pre-fix chat template, fell into a repetition loop, and committed nothing. They measure that serving combination, not the model — the evidence is on the Gemma-12B data page.

Two Ternary-Bonsai-27B guided runs with thinking off (2026-09-06) are invalid and are not listed above. The first hit a dead gh token and looped on login. The second ran 85 identical shell calls in a row against a missing file and committed nothing in three hours, and the operator stopped it. Both rows are in the guided CSV with their stop reasons. A third attempt is scheduled, with the runner's live loop stop in place.

Three scored rows above are scheduled for a re-run: Qwen3.8-27B GGUF blind (87), Gemma-4-26B-A4B GGUF blind (47.5) and Qwen3.6-35B-A3B guided (83). They compacted under the harness's old reserve of 16384 tokens; since 2026-09-06 the harness reserves 8192, the answer budget. Each keeps its row until the fresh one lands. The 87 ran at effort medium, which Qwen3.8 is no longer run at; its fresh row is the 4-bit build at effort xhigh, 93 above, complete on a 65536 window at the 8192 reserve. The rows at xhigh (80.5) and low (66) above are the ISTA 3-bit build at the 8192 reserve, on a 147456 window. Every row from 2026-09-09 on records its sampling: temperature 1.0, top_p 0.95, the server's default.

Cloud reference — blind test

modelharnessscore
kimi-k3pi93.5/100
grok-4.6pi92.5/100
gpt-5.6-solpi92/100
claude-opus-5pi90.5/100
deepseek-v4-flash-0731pi84.5/100
gpt-5.6-lunapi83.5/100
deepseek-v4-pro-0813pi79/100
glm-5p3-flashpi75/100
Claude Sonnet 4.5pi43.5/100
claude-haiku-4.5pi34/100

Guided test

Two local models have guided rows on the current prompt so far, alongside the cloud anchors. More local guided runs are queued on the same frozen prompt; a blind-guided pair can land at different times.

configharnessscore
glm-5p3-flashpi98/100
deepseek-v4-flash-0731pi97/100
gpt-5.6-lunapi88.5/100
claude-sonnet-4.5pi88/100
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-servermtp/3q8_0onpi83/100
Claude Haiku 4.5pi76/100
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-servermtp/3q8_0offpi62.5/100
Gemma-4-26B-A4B UD-Q4_K_XLunsloth,llama-servermtp/2f16onpi57/100 (partial)
Qwen3.6-35B-A3B UD-Q4_K_XLunsloth,llama-servermtp/3q8_0offpi46.5/100
Gemma-4-12B Q4_K_XLunsloth,llama-serverf16offpi37.5/100 (partial)
Ternary-Bonsai-27B Q2_g64prism-ml,prism-llamaq4_0+biasonpi31.5/100 (partial)
Gemma-4-26B-A4B UD-Q4_K_XLunsloth,llama-servermtp/2f16offpi25/100 (partial)
Ternary-Bonsai-27B 2-bitprism-ml,mlx_lm.serverf16onpi12.5/100 (partial)
Ternary-Bonsai-27B Q2_g64prism-ml,prism-llamaf16onpi12.5/100 (partial)