Local inference field note / September 2026
A 125B model runs at full speed on a stock 64GB Mac
A 125B mixture-of-experts model, stock settings, one flag.
I ran DwarfStar’s own benchmark on my 64GB M4 Max, three times per setting. Qwen3.8-Flash-Next holds 41 tokens a second from 2K to 64K context without touching a system setting. Here is what that takes, and how other Macs should differ.
- generation, flat from 2K to 64K context, at the memory limit macOS ships with, 3 of 3 runs
- 41 t/s
- tokens per second of prefill across the same sweep, per-step medians
- 579–621
- from process start to first token with the model warm; 23.7 s cold (separate probe, raised limit)
- 3.7 s
On this page
The short answer
Qwen3.8-Flash-Next is a 125B mixture-of-experts model with 6B parameters active per token, and since part three it has been the model behind my daily assistant on a 64GB Mac. Those notes measured it the way I use it: a long-running server, speculative decoding, real agent turns. Honest numbers, and impossible to hold up against anyone else’s machine. So this time I ran the benchmark that ships with DwarfStar , Salvatore Sanfilippo’s inference engine: the same text and the same 32 steps from 2K to 64K context that upstream asks every machine to run. The results are in PR 1142 and, with the per-run logs, in my bench repo .
On a Mac Studio M4 Max with 64 GB, the 2-bit Qwen reads a prompt at 579 to 621 tokens per second and writes at 40.7 to 41.5, per-step medians, and the line is flat: the 64K step is within 2 percent of the 4K step on both. It does that at the GPU memory limit macOS ships with, with one flag, --prefill-chunk 2048, in 3 of 3 runs, without swapping. That is what I mean by full speed: the stock limit gives the same numbers as a raised one, within 1 percent at every step, and from 4K to 64K the speed holds within about 1.5 percent.
The same Mac also runs DeepSeek V4 Flash, an 81 GiB model file, by streaming its experts from the SSD. It writes at 12.9 to 15.8 tokens a second and reads at about 135, except when it suddenly does not. It runs. I would not live with it.
Have a 64GB Mac? Run the same command and send me your numbers. How to run it
Why this benchmark and not my daily numbers
ds4-bench is DwarfStar’s own instrument. It feeds a fixed Italian novel into the model 2,048 tokens at a time, and after each step it measures how fast that step was read (prefill) and how fast the model then writes 128 tokens (generation). Every machine in upstream’s performance notes ran the same sweep, so a number from it means the same thing on an M5 Max, on the build it was recorded with, as on my M4 Max.
The decode speed in part three, 43 to 46 tokens a second, came from a server with --mtp speculative decoding on Ivan Fioravanti’s fork, on real Hermes turns. ds4-bench does not accept --mtp and measures plain decoding, so 41 here is not a regression. It is a different measurement, and the first reproducible ds4-bench measurement of this machine.
The results
Qwen repeats tightly. Of the 288 steps in the nine valid Qwen runs, every step except the two outliers below lies between 571 and 622 tokens per second of prefill and 40.1 and 41.6 of generation. The outliers are generation dips: 35.7 at 2K in the very first run of the day, and a single 38.6 at 12K in one evening run. Neither repeated.
ds4-bench sweep, median of three valid runs per line, shaded from the slowest to the fastest run at each step. Qwen’s band is almost invisible because the runs agree. DeepSeek’s prefill band is the story: every run dips somewhere, never at the same depth. Qwen ran at the macOS default memory limit; DeepSeek at the raised one.Swipe to compare columns →
| Model and setting (valid runs) | 2K | 16K | 32K | 64K |
|---|---|---|---|---|
| Qwen, default limit, --prefill-chunk 2048 (3 of 3) | 621.3 / 41.54 | 590.1 / 41.14 | 583.4 / 41.10 | 579.4 / 40.70 |
| Qwen, default limit, official command (1 probe run) | 615.8 / 41.29 | 590.3 / 41.01 | 581.3 / 40.74 | 577.4 / 40.63 |
| Qwen, raised limit, official command (3 of 3) | 616.4 / 41.24 | 590.6 / 40.98 | 581.8 / 40.81 | 579.5 / 40.34 |
| Qwen, raised limit, --prefill-chunk 2048 (2 of 4) | 621.3 / 41.15 | 590.0 / 40.90 | 583.4 / 40.84 | 579.8 / 40.48 |
| DeepSeek, raised limit, --ssd-streaming (3 of 4) | 156.0 / 12.85 | 138.2 / 15.59 | 138.0 / 14.83 | 123.1 / 14.39 |
Neither the prefill chunk nor the memory limit changes Qwen’s speed on this machine: the default chunk and 2048, at the raised or the default limit, agree within about 1 percent at every step. So for a 64GB Mac, --prefill-chunk 2048 is free insurance, and the next section is why you want the insurance.
One thing makes the flat line possible. The Qwen file is 137 GiB, more than twice the Mac’s memory, but ds4 keeps only 41.7 GiB of it resident and reads the rest, rows of the model’s large lookup table, from the SSD as tokens need them. In the raised-limit Qwen runs, disk reads averaged about 160 MB/s, with peaks near 2.2 GB/s.
It runs at stock settings
I expected to need a raised GPU memory limit. ds4 prints its memory plan at startup, and the official command plans 52.05 GiB, more than macOS gives the GPU by default. Measured, stock settings work: at the default limit, which Metal reports as a 51.84 GiB working set on this Mac, --prefill-chunk 2048 ran 3 of 3 runs at full speed. The step medians matched the raised-limit runs within 0.3 percent for prefill and 1 percent for generation, and ds4-bench itself swapped nothing.
| GiB | |
|---|---|
| Qwen, official command (chunk 8192) | 52.05 GiB: model 41.72, buffers 8.24, KV 2.09 |
| Qwen, --prefill-chunk 2048 | 46.11 GiB: model 41.72, buffers 2.31, KV 2.09 |
| DeepSeek, --ssd-streaming | 48.36 GiB: 8.20 resident, 35.42 expert cache, 3.38 prefill reserve, 0.86 KV, 0.50 buffers |
| macOS default limit on this Mac | 51.84 GiB (55,662,788,608 bytes) |
| Raised limit, iogpu.wired_limit_mb=57344 | 56.00 GiB |
Chunk 2048 plans 46.11 GiB, leaves 5.7 GiB of room and cost nothing measurable here, so that is the setting for a stock 64GB Mac. The unmodified command also completed one run at the default limit, but with 0.21 GiB less budget than it planned for, so I would not rely on it. Why it fits anyway, I have not traced. My untested guess is that Metal’s working-set figure is a recommendation rather than a hard allocation limit.
DeepSeek V4 Flash, streamed from SSD
DeepSeek V4 Flash at 2 bits is an 81 GiB file, too big to hold in 64 GB. With --ssd-streaming, ds4 keeps 8.2 GiB of shared weights resident, caches 35.4 GiB of experts, and reads the rest from the SSD on demand, at up to 6.1 GB/s in these runs. All four runs completed the sweep; one was excluded by my swap rule, below.
- Generation is steady at 12.9 to 15.8 tokens a second across the sweep, per-step medians.
- Prefill is not. The median is about 135 tokens a second, but every run drops somewhere, never at the same depth. The worst was the day’s first DeepSeek run, which fell to 47 to 50 tokens a second from 57K to 63K. The one other run in the same position in the order, the one excluded below, did not show it. In the first run, the drop coincided with the memory compressor growing from about 2.3 to 5.4 GiB, and with pageouts. That is a correlation in one run, not a cause.
- It squeezes the rest of the Mac. In the excluded run the system swapped in 403 MiB, about two thirds of it in a burst three to six minutes in, while ds4 itself swapped nothing. In that run, streaming pushed other processes out of memory.
It is a real result: a model bigger than the machine’s memory, answering at a readable pace. But a prefill that can collapse to a third of its speed on a long prompt is not what I want under an agent that reads web pages all day. Qwen stays my daily model.
Cold start, the part 64 GB changes
The official sweep does not time model loading, and on a 64GB machine loading is where the page cache fights you. So I added a separate probe, not part of the official format, run at the raised limit: start a fresh process, read a 2,048-token prompt, write one token. “Cold” is right after sudo purge, so no part of the model file is cached; “warm” is an immediate repeat. Three pairs each.
Swipe to compare columns →
| Model and cache | Start to first token | Of which loading | Prefill | First decode |
|---|---|---|---|---|
| Qwen, cold (after sudo purge) | 23.7 s (23.6 to 24.6) | 20.3 s (20.2 to 21.2) | 3.3 s | 26 ms |
| Qwen, warm | 3.7 s (3.7 to 3.7) | 0.3 s (0.3 to 0.3) | 3.3 s | 26 ms |
| DeepSeek streamed, cold | 18.2 s (18.0 to 19.1) | 4.6 s (4.5 to 4.7) | 12.9 s | 597 ms |
| DeepSeek streamed, warm | 13.8 s (13.6 to 14.0) | 0.3 s (0.3 to 0.3) | 12.8 s | 630 ms |
“Loading” is derived: the total minus prefill minus the first token. Qwen’s cold penalty, about 20 seconds, is most likely reading 41.7 GiB of weights from the SSD; that is inference from the arithmetic, not a trace. DeepSeek starts faster cold because only 8.2 GiB is resident, then pays for it on every request: 12.9 seconds to read 2,048 tokens against Qwen’s 3.3, and about 0.6 seconds for its first decoded token against 26 milliseconds.
For a daily server the lesson is simple. Qwen’s load is paid once, when the server starts. DeepSeek’s streaming cost is paid on every turn. I did not measure first-token time inside a long-running server, which is what an agent user actually feels; there the model is always warm.
How other Macs should differ
Three properties of the machine decide these numbers, and a fourth decides whether a model runs at all. The first two are the usual rule of thumb for local inference, not something this benchmark measured:
- Generation is mostly bound by memory bandwidth. Every new token reads the model’s active weights from memory once, so tokens per second tracks how fast memory can be read.
- Prefill is mostly bound by GPU compute. Reading a prompt processes thousands of tokens against the same weights at once, so it tracks GPU cores and how fast each one is.
- Capacity decides resident or streamed. If the working set fits in unified memory, the model runs resident. If not, as with DeepSeek here, it streams from the SSD, and then the SSD’s read speed sets the pace. Apple does not publish SSD speeds on these pages; streaming DeepSeek peaked at 6.1 GB/s of reads on this Mac, measured.
- An Ultra is two Max dies joined. Apple builds M1 Ultra by connecting two M1 Max dies with UltraFusion, which doubles both: 800 GB/s against 400, and 64 GPU cores against 32.
Swipe to compare columns →
| Chip | Bandwidth | GPU cores | Memory | Apple source |
|---|---|---|---|---|
| M1 Max | 400 GB/s | 24 or 32 | 32 or 64 GB | MacBook Pro 2021 specs |
| M1 Ultra, two M1 Max dies | 800 GB/s | 64 | up to 128 GB | Newsroom, March 2022 |
| M4 Max, 32-core GPU | 410 GB/s | 32 | 36 GB | Mac Studio 2025 specs |
| M4 Max, 40-core GPU (this Mac) | 546 GB/s | 40 | 48 to 128 GB | Mac Studio 2025 specs |
| M3 Ultra | 819 GB/s | 60 or 80 | 96 or 256 GB | Mac Studio 2025 specs |
M1 to M4, one data point. PR 1068 ran the same Qwen model on an M1 Max 64 GB with the same --prefill-chunk 2048. At 16K:
Swipe to compare columns →
| M1 Max | M4 Max, 40-core | Ratio | |
|---|---|---|---|
| Prefill at 16K | ~274 t/s | 590.1 t/s | 2.15x |
| Generation at 16K (steady) | ~24.3 t/s | 41.2 t/s | 1.70x |
| Memory bandwidth (Apple spec) | 400 GB/s | 546 GB/s | 1.37x |
Generation is 1.7 times faster while the bandwidth is 1.37 times higher, and prefill is 2.15 times faster. That is consistent with bandwidth mattering for generation and compute for prefill, but it cannot isolate either: the two runs are 25 engine commits apart, the M1 run stopped at 16K, and its GPU core count, 24 or 32, is not stated. Both M1 Max variants have the same 400 GB/s. Apple’s own claim for the M4 Max GPU is up to 1.9 times the M1 Max.
Max against Ultra. By the rule of thumb, a resident model on an Ultra should generate up to about twice as fast as on the Max of the same generation, and prefill should scale with the doubled GPU. In the 2025 Mac Studio the M3 Ultra, a generation older, still has 1.5 times the M4 Max’s bandwidth, 819 GB/s against 546. I have not measured an Ultra. The only 64 GB Ultra data point I found upstream, the closed PR 350 , streamed DeepSeek on an M1 Ultra at about 5 tokens a second, slower than this M4 Max at 12.9 to 15.8, on an older engine commit with a smaller manual cache. A streamed model is limited by the SSD and the cache, and twice the bandwidth does not help much there.
Capacity changes the question. Upstream’s 128 GB rows, an M5 Max and a DGX Spark, hold the same DeepSeek model entirely in memory. On 64 GB it can only be streamed. Those rows are not a faster version of this one; they are a different way of running the model.
The details
Everything below is for people who want to check the numbers or run the same thing. If you only wanted the result, you have it.
Method
The engine was upstream antirez/ds4 at 0aaea5a, the tip of main on 28 September, built with plain make in a separate clean checkout, with zero warnings. Both model files came from upstream’s own download_model.sh and were checked against their Hugging Face sha256 before the first run. The command was the documented one, run from the checkout:
./ds4-bench -m MODEL \ --prompt-file speed-bench/promessi_sposi.txt \ --ctx-start 2048 --ctx-max 65536 --step-incr 2048 \ --gen-tokens 128 [--prefill-chunk 2048 | --ssd-streaming]
- A quiet machine. My daily model server was stopped and all six of Hermes’s scheduled jobs paused for the window; the desktop chat app was quit. Still running: two Claude Code sessions, one of them supervising the runs, an idle agent gateway and macOS’s own services. Load average was 1.3 to 1.6 at the start.
- One machine, three runs per setting, in alternating order (forward, reversed, forward), with two minutes of cooldown after each. A Qwen sweep takes about 25 minutes.
- A validity rule set before the first run: exit code 0, all 32 rows present, and at most 100 MiB swapped in system-wide during the run. System-wide on purpose, so it catches a machine that pages at all, not just the benchmark.
- Sampling every five seconds: memory counters, disk throughput, swap and thermal state, with
/usr/bin/time -laround each run for the process’s own swaps.
The full protocol, including the replacement rule written before the replacement batch, is in the repo’s METHODOLOGY.md .
What is measured and what is not
| Claim | Basis |
|---|---|
| Prefill and generation speeds, planned memory, swap counts, first-token times | Measured: raw CSVs, engine reports, time -l |
| The default GPU limit on this Mac is 51.84 GiB | Measured: Metal's recommendedMaxWorkingSetSize at iogpu.wired_limit_mb=0 |
| --prefill-chunk 2048 runs at full speed at the default limit | Measured: 3 of 3 runs |
| The official command also fits at the default limit | Measured once: 1 probe run, 0.21 GiB over the limit on paper |
| Replay beyond ~30K does not affect the timed numbers | Derived: the timers in ds4_bench.c bracket only prefill and generation |
| The excluded runs' swap-ins came from other processes | Derived: time -l shows 0 swaps for ds4-bench while the system counter rose |
| Qwen's 20 s cold penalty is reading its weights from SSD | Inferred: 41.7 GiB at ~2 GB/s is about 20 s; not traced |
| Memory bandwidth, GPU core counts, Ultra = two Max dies | Apple's tech-specs and newsroom pages |
| Generation tracks bandwidth and prefill tracks GPU compute | The usual rule of thumb; consistent with, not proven by, the M1 Max comparison |
| DeepSeek's prefill drops come from memory pressure | Hypothesis: correlated with compressor growth in one run only |
What went wrong
Three runs failed my own swap rule and are excluded: two Qwen chunk-2048 runs, at 137 and 168 MiB, and the DeepSeek run at 403. In all three, ds4-bench itself swapped nothing, and their numbers match the valid runs, so the rule threw out good data. I kept it anyway: a rule you relax after seeing the results is not a rule. I ran one replacement per affected setting, under a note written before the batch, and stopped there. That is why one row says 2 of 4.
One rule did change after I saw the data, and it is the one worth flagging. The charts and the upstream submission use one representative run per setting. My first rule picked the run closest to the median generation speed, and for DeepSeek that chose the run with the deepest prefill drop, which would have misrepresented prefill. I switched to the run closest to the medians of both prefill and generation at every step, for every setting, and the methodology says so.
Who did what: this benchmark was done with Claude Code , running on the same Mac. It made the run plan, wrote the driver and the probes, ran the batches, did the analysis and wrote it up, including this note and the upstream PR. I directed the work, set the quiet-machine windows, reviewed the results and approved every change to the machine. Every published number was re-derived from the raw CSVs by a separate check. The mistakes are its as much as mine.
Run it on your 64GB Mac
If you have a 64GB Apple Silicon Mac, I would like your numbers. The benchmark alone is a handful of commands:
git clone https://github.com/antirez/ds4 && cd ds4 git checkout 0aaea5a && make ds4-bench ./download_model.sh qwen38-q2 # 137 GiB ./ds4-bench -m gguf/Qwen3.8-Flash-Next-Q2.gguf \ --prompt-file speed-bench/promessi_sposi.txt \ --ctx-start 2048 --ctx-max 65536 --step-incr 2048 \ --gen-tokens 128 --prefill-chunk 2048
The bench repo has the driver that quiets the machine, samples memory and applies the validity rule, and every per-run CSV, engine report and memory sample behind this note is in its 2026-09-28 results . The data is CC BY 4.0 and the scripts are MIT. Stop anything else that holds GPU memory first, then send me your numbers on X , or upstream the way PR 1068 and PR 1142 did.
Sources
- 64gb-mac-llm-bench: raw results, driver and methodology
- PR 1142: this data submitted to DwarfStar
- PR 1068: slycrel's M1 Max 64 GB Qwen3.8 run
- PR 350: ozgursoy's M1 Ultra 64 GB streaming run
- DwarfStar performance notes: the 128 GB rows
- Apple: Mac Studio (2025) tech specs, M4 Max and M3 Ultra
- Apple: MacBook Pro (16-inch, 2021) tech specs, M1 Max
- Apple newsroom: M1 Ultra, two M1 Max dies joined by UltraFusion
- Apple newsroom: M4 Pro and M4 Max
- DwarfStar, Salvatore Sanfilippo's inference engine
- Qwen3.8-Flash-Next model card: 125B, 6B active, mixture of experts
- Upstream's Qwen3.8-Flash-Next model files
- Upstream's DeepSeek V4 model files
- Part three: A 2-bit 125B model became my daily assistant on a 64GB M4 Max
Measured on a Mac Studio M4 Max with 64 GB, macOS 26.6.2, on 28 September 2026, using upstream DwarfStar 0aaea5a.