Skip to content

Historical data — M1 Max 32 GB

What this page is

Superseded measurements, moved off the current pages as they were replaced. Newest first: the top section is the most recent supersession; the bottom is the oldest era. Each section says what replaced it and why. The per-run findings and conclusions live in the repo's benchmark findings index.

DO NOT USE THESE NUMBERS

Everything on this page is superseded, wrong, or both. It is kept only to show what changed and why.

  • Retired memory limit. Many rows were measured at iogpu.wired_limit_mb=27000, which made the machine too slow for normal use, and some at 24000, which the current 25000 replaced. A context maximum here is either too high or too low for the standing limit.
  • Wrong axis. Several tables measure allocated context, which is storage, not speed. The depth sweeps replaced this with decode speed against used context — the number that decides whether a config is usable.
  • Deflated quality scores. Early EvalPlus passes used a fixed output budget that was too small, so reasoning ran out of tokens and empty completions scored as failures. Two models were badly understated.
  • Mixed eras in one table. Some rows here are current and some are not, and they are not always labeled.

For numbers you can act on, go to the comparison page. Full raw archives, with their eras labeled, live in the benchmarks pages.

Qwen3.6-35B-A3B rows at the retired 24000 limit (superseded 2026-09-10)

Measured 2026-09-07 at iogpu.wired_limit_mb=24000, the standing value until the limit moved to 25000. Replaced on the current pages by the 25000 measurements of 2026-09-06 and 2026-09-07, which serve larger windows on every arm. Kept here because the 24000 numbers were the site's current rows for three days.

configlargest -c that servescreepwired at the deepest row
GGUF, MTP, f16 KV33792 (33920 OOMs)no ceiling found to 32818, 56.0 tok/s there; 67.7 at 4K24.9 GB
GGUF, MTP, q8_0 KV40960 (49920 passes a one-token probe, then OOMs on the first real step)36.7 at 4K, 19.7 at 32818, stopped there on the memory-compression rule with zero swap; the stop was under review24.8 GB
MLX 4-bit, unquantized KV (2026-08-29)no -c53.3 at 4K, 42.0 at 37K, then a Metal OOM before 41K18.7 GB

What replaced them: at 25000 the f16 arm serves -c 40960 (69.1 at 4K, 52.6 at 41K, no ceiling found), the q8_0 arm serves -c 98304 (36.5 at 4K, 9.24 at 82K, speed floor at 98K), and MLX reaches 41K at 37.4 tok/s in 24.6 GB.

The 22000 "in use" wired limit (retired 2026-09-06)

From 2026-08-29 the setup page carried a second standing value: iogpu.wired_limit_mb=22000 for when the owner worked beside a run. Below about 24000 the sysctl gates cleanly, so the machine stayed responsive, at the cost of context depth (Qwen3.6-35B MLX capped near 13K instead of about 35K). Retired on 2026-09-06: at the model sizes under test this machine is a model server, driven from another machine, and a shared-use setting is not a case we evaluate. No published number was measured at 22000. The single standing value is 24000.

q8_0 KV curves and allocation tables on the report pages (superseded 2026-09-06)

The report pages carried the q8_0 KV decode curves and the allocation-only context tables next to the f16 curves that replaced them. They moved here on 2026-09-06. The KV pick is f16 on Qwen3.8 and Gemma-26B; on Qwen3.6 q8_0 stays because f16 does not load, but the published -c 98304 OOMs at load under the 24000 limit, so its allocation table is wrong too.

Gemma-4-26B-A4B, llama+MTP q8_0 KV, limit 25000, 2026-08-28:

depthtok/s
4K23.5
16K11.2
24.5K7.97, under the 8 tok/s floor

RSS at floor depth (24.5K, 32K alloc): 15.4 GB.

Gemma-4-26B-A4B, context allocation at q8_0 KV, n-max 2, limit 24000 (allocation only, nothing measured at depth):

-cslotsresultRSS
262,1441 (f16 KV)Metal OOM at load
262,1441OK, 62.4/53.3 tok/s, 256-token check19.3 GB
327,6802×160KOK, 67.6/52.7 tok/s20.1 GB
376,8322×184KOK, 58.4/56.5 tok/s20.4 GB
385,0242×188KMetal OOM

Qwen3.8-27B, llama+MTP q8_0 KV, limit 25000, 2026-08-28:

depthtok/s
4-8K14.1 / 12.8
16K8.6
24.5K7.3, under the 8 tok/s floor

RSS at floor depth (19K, 32K alloc): 18.9 GB.

Qwen3.6-35B-A3B, llama+MTP q8_0 KV at -c 98304, fast sweep, limit 25000, 2026-08-28:

depthtok/s
4K44.5
16K30.1
33K18.8
49K13.5
65K / 82K / 90K10.7 / 8.8 / 8.1

Qwen3.6-35B-A3B, context ramp at q8_0 KV, n-max 3, limit 24000, 256-token check only:

-cresulttok/sRSS
98,304OK62.0 / 67.722.9 GB
106,496Metal OOM
131,072Metal OOM

The slow creep of 2026-09-04 found 49152 the largest -c that loads for it, with compaction from 16K; the current row is 8K clean at 43.8 tok/s.

Two slot rows at q8_0 KV, allocation only (superseded 2026-09-05)

Both rows carried an allocation as their context and no measured depth. On 2026-09-05 one slot of each was measured at f16 KV, with the other slots loaded and idle, at the largest -c that serves a real completion.

rowoldnow
Gemma-4-12B GGUF MTP, 4 slots, thinking offq8_0 KV, -c 1048576, 4x256k, 33.7 tok/s shallow, 16.9 GBf16 KV, -c 655360, 4x49k gated by mem, 42.9 to 27.7 tok/s, 25.1 GB
Gemma-4-26B-A4B GGUF MTP, 2 slotsq8_0 KV, -c 376832, 2x184k, nothing measuredf16 KV, -c 202752, 2x82k gated by mem, 66.6 to 33.6 tok/s, 25.3 GB

The Gemma-12B sweep ran with free memory near zero and swap already in use at session start; the row says so. Evidence: hardware/m1-max-32gb/benchmarks/bench10/results.md.

Three GGUF rows at the fast sweep and the old KV type (superseded 2026-09-05)

The KV pick of 2026-09-04 chose the cache type per model with a short creep of both types, then ran the slow creep at the pick with the largest -c the machine loads. These rows were measured before that, with the fast sweep of 2026-08-28 and, for two of them, at q8_0 KV. All three published -c values OOM at load under the 24000 limit.

rowoldnow
Qwen3.8-27B GGUF MTP, effort mediumq8_0 KV, -c 32768, 19k gated by speed, 14.1 to 8 tok/s, 18.9 GBf16 KV, -c 49152, 49k gated by mem (OOM at load above it), 20.0 to 15.0 tok/s, 23.5 GB
Gemma-4-26B-A4B GGUF MTPq8_0 KV, -c 262144, 24k gated by speed, 23.5 to 8 tok/s, 15.4 GBf16 KV, -c 212992, 197k gated by mem (OOM at load above it), 60.3 to 17.3 tok/s, 25.6 GB
Qwen3.6-35B-A3B GGUF MTP, thinking onq8_0 KV, -c 98304, 90k gated by speed, 44 to 8.1 tok/s, 22.8 GBq8_0 KV, -c 49152, 8k gated by mem, 36.4 to 43.8 tok/s, 25.0 GB

The Qwen3.6 change is not a KV change. f16 does not load for it, so q8_0 stays; the slow creep's memory columns showed compaction from 16K that the fast sweep could not see, and the published -c never loaded. Evidence: hardware/m1-max-32gb/benchmarks/bench9/results.md.

Gemma-4-12B figures measured on the retired LM Studio entry (superseded 2026-09-04)

Every number in this section was measured on the LM Studio entry google/gemma-4-12b. That entry always thinks, ships Google's pre-fix chat template, produced all three invalid Mendel rows, and is gone from the model store. Its readings were copied onto the thinking-off row of the site, which is a different entry. Do not use them for either entry. The current curves for both backends are on the Gemma-12B report; the evidence behind the retirement is on its data page.

Headline figures withdrawn: 29.3 tok/s at 65K used tokens, a compression-onset ceiling of 65-74K, a shallow reading of 35.4 tok/s, and 8.1 GB at max context.

Ceiling confirmation sweep (--parallel 4, watcher at 20 s, wired limit 24000, 2026-08-30):

depthdecode tok/swatcher state
41,09531.05clean
49,11230.25clean
57,07729.25clean
65,09429.29 — read as the last clean stepclean
74,09927.95compression/swap onset inside this step

Shallow sweep (--parallel 4, STEP_SLEEP=25, wired limit 24000, 2026-08-30):

depthdecode tok/s
4,17535.41
8,29234.89
16,46533.88
24,63833.08
33,07132.19

RSS 8.1 GB at 33,071 tokens.

EvalPlus, thinking on: 0.622 base / 0.610 plus / 63% completion, with 61 of 164 completions empty. It is a completion-rate failure wearing a quality number's clothes — 102 of the 103 answers it delivered pass. It is not a score for any current configuration, and no thinking-on score exists for the model today.

Mendel rows from old prompt versions (superseded 2026-09-02)

Replaced by fresh rows on blind prompt v1.1 and guided v3.0, run from the new moving base tags (benchmark-blind-base, benchmark-guided-base), which include the tap crash fix. The site tables now count only the current prompt version; the full per-version archive stays in the hosted Mendel reports. Never compare these rows with the current ones — different prompt versions never share a table.

modeltestconfigscorestatus
Qwen3.8-27Bblind v1.0mlx 4-bit, effort medium, pi80/100partial — ~4h time budget, 3/8 libraries
Ternary Bonsai-27Bblind v1.0mlx 2-bit, pi58/100partial — mlx_lm.server tool-parser crash
Qwen3.6-35B-A3Bblind v1.0llama-server41.5/100complete
Gemma-4-26B-A4Bblind v1.0llama-server38/100partial
Qwen3.6-35B-A3Bguided v2.1llama-server65.5/100complete — all 8 libraries in 75.6 min
Ternary Bonsai-27Bguided v2.1mlx 2-bit, pi69/100partial — stuck 45+ min on a self-made bug, closed at 3/8
Qwen3.8-27Bguided v2.1mlx 4-bit, pi84/100partial

The run narratives for these rows lived on the comparison page; their findings stay in the benchmark findings index and the hosted reports' defect ledgers.

Bonsai MLX depth rows 44-48K, transient dip (superseded 2026-08-30)

Three rows from an early depth pass read far under the curve. The watched slow-creep re-test recovered to ~18 tok/s at greater depths (50-58K), so the dip was a transient system episode (memory pressure or background load), not a property of the config. The current Bonsai report shows the clean curve.

depthdecode tok/s
44K12.10
46K11.89
48K11.33

Gemma-4-12B LM Studio ceiling, old criterion (superseded 2026-08-30)

Old rows read "170K, 29.7 tok/s" — the deepest point LM Studio's auto-fit loader let a request reach before failing clean, not a compression/swap ceiling. The revised criterion (context creep) defines the ceiling as the onset of memory compression/swap in the watcher log, with tok/s taken from the last clean step before onset. A confirmation sweep under the new criterion found onset between 65K and 74K used tokens (65,094 tokens clean at 29.29 tok/s; 74,099 tokens shows compression bursts up to 114,012 pages). That replacement is itself superseded: it was measured on the retired entry — see the 2026-09-04 section at the top of this page. Current figures are on the comparison page and the Gemma-12B report.

Fast-sweep memory ceilings, pre slow-creep rule (limit 25000, superseded 2026-08-29)

The first depth sweeps for these three MLX configs used a fast sweep — no pause between depth steps. The slow-creep rule (25 s pause per step, the measurement rules) replaced them with a re-test at limit 24000 on 2026-08-29. Shallow-depth rows (below the lowest row here) did not change and stay on the current pages.

modelfast-sweep ceilingfast-sweep last stableslow-creep re-test
Gemma-4-26B-A4B, MLX82-98K82K, 20.6 tok/s70-72K; last stable 70K, 12.83 tok/s
Ternary-Bonsai-27B, MLX57-61K57K, 18.2 tok/s58-60K; last stable 58K, 17.27 tok/s
Qwen3.8-27B, MLX~32K (OOM, server thread died)28.7K, 14.2 tok/s, RSS 14.3 GB28-30K; last stable 28K, 15.29 tok/s

Deflated EvalPlus scores (fixed 3072-token budget, corrected 2026-08-28/29)

The first quality pass capped output at 3072 tokens. Reasoning exhausted the cap, and the empty completions scored as hard failures. These are lower bounds, not measurements.

model / configdeflated basedeflated plusemptycorrected base/plus/completion
Qwen3.6-35B-A3B, llama+MTP, thinking on0.6100.61062/164 (~38%)0.939 / 0.921 / 97%
Ternary Bonsai-27B, mlx 2-bit, thinking on0.6400.63449/164 (~30%)0.915 / 0.884 / 97%
Qwen3.8-27B, mlx 4-bit, effort medium0.9700.9393/164 (~2%)0.982 / 0.939 / 100%

The lesson generalizes: treat any single-pass score with a fixed output budget as a lower bound until the budget is calibrated from measured reasoning length (the EvalPlus method).

Decode speed (best server-usable config per model, retired 2026-08-28)

Cards moved out of the comparison page when it was slimmed to the current picture, 2026-08-28.

modelconfigpy tok/sjs tok/s
Gemma-4-26B-A4B (MoE)llama-server + MTP n=271.969.3
Qwen3.6-35B-A3B (MoE)llama-server + MTP n=368.273.5
Gemma-4-12Bllama-server + MTP n=335.035.6
Ternary Bonsai-27Bmlx_lm.server (limit 25000: 24.5)24.524.5
Qwen3.8-27Bmlx_lm.server19.719.6
Qwen3.8-27Bllama-server + MTP n=316.915.7

All values are thinking-on where the model supports it. Thinking-off (sub-agent mode): Gemma-26B 74.8/71.6 at n=2, Gemma-12B 45.2/31.3 at n=4. Qwen's true fastest, MLX + MTP at 20.2/22.5, is CLI-only and cannot back a harness. Gemma-26B's numbers are f16 KV at 32K; its 256K config needs q8 KV now, which drops js to ~53 tok/s because draft acceptance falls under q8.

Max context — single session (retired 2026-08-28)

modelconfigmax contexttok/s at it
Gemma-4-26B-A4B (MoE)llama-server, 1 slot, q8_0 KV256K (model limit)62.4 py / 53.3 js
Gemma-4-12Bllama-server, 1 slot256K (model limit)45.2
Qwen3.6-35B-A3B (MoE)llama-server, 1 slot, q8_0 KV96K (memory, limit 25000)62.0
Ternary Bonsai-27Bmlx_lm.server, bounded prompt cache49K (memory, limit 25000)18.8
Qwen3.8-27Bllama-server, 1 slot, q8_0 KVre-probe pending at limit 25000
Qwen3.8-27Bmlx_lm.server28K OK; ceiling <33K (limit 25000)14.2 at 28K

Context limits are mode-independent, because KV is preallocated. The tok/s column here comes from short thinking-off probes; see each report for thinking-on speeds.

Multi-session, concurrent agents (retired 2026-08-28)

modelconfigsessionscontext each
Gemma-4-12Bllama-server --parallel 4 -c 1048576, q8_0 KV4256K
Gemma-4-26B-A4B (MoE)llama-server --parallel 2 -c 376832, q8_0 KV2184K
Qwen3.6-35B-A3B (MoE)untested at limit 25000 (at 24000: OOM at 2×20K)
Qwen3.8-27Bre-probe pending at limit 25000
Ternary Bonsai-27Bprism fork --parallel 2 -c 98304, q4_0 KV248K — 9.8 tok/s each concurrent, 10.0 GB RSS

Decision: parallel serving runs on llama-server only. MLX has no slots. Its only concurrency is one server process per agent, each with its own full weight copy and its own port to wire into the harness. That was measured — two Bonsai instances at 14.0 tok/s each, 14.9 GB — but ruled out as not worth the operational fiddling. Bonsai regains a multi-session story when a brew llama.cpp release loads its ternary GGUF, with a projected ~300K total context to split across slots.

Qwen3.6-35B-A3B — context ramp at the retired 27000 limit (oldest era)

Moved off the report page. f16 KV, MTP n-max 3. Under the current 25000 limit this config reaches 96K with q8_0 KV, not 139K.

-cresulttok/sRSS
49,152OK66.322.8 GB
98,304OK67.223.7 GB
131,072OK67.824.3 GB
139,264OK — maximum at 2700065.524.5 GB
147,456Metal OOM
196,608Metal OOM

With q8_0 KV the 27000 limit reached 208K on one slot and 2×96K on two. Those configs are retired.

Qwen3.8-27B — context ramp and slot layouts at the retired 27000 limit (oldest era)

Moved off the report page. f16 KV, MTP n-max 3. The current maxima at 25000 have not been re-probed, so no replacement table exists yet — the depth floor at ~19K makes big allocations pointless for this model anyway.

-cresultRSS
49,152OK21.2 GB
65,536OK22.0 GB
98,304OK — validated with a 4K-token prompt24.1 GB
106,496Metal OOM
114,688Metal OOM
131,072Metal OOM
agentsflagscontext per agentRSS
1--parallel 1 -c 9830496K24.1 GB
2--parallel 2 -c 9011244K24.2 GB

MTP stayed active in both. The next 8K step OOMed in both: -c 106496 with one slot, -c 98304 with two. The 160K single and 2×72K configs from this era are withdrawn.


Raw data, with eras labeled, in the benchmarks pages.