Skip to content

Local inference field note / September 2026

A 125B model runs at full speed on a stock 64GB Mac

A 125B mixture-of-experts model, stock settings, one flag.

I ran DwarfStar’s own benchmark on my 64GB M4 Max, three times per setting. Qwen3.8-Flash-Next holds 41 tokens a second from 2K to 64K context without touching a system setting. Here is what that takes, and how other Macs should differ.

generation, flat from 2K to 64K context, at the memory limit macOS ships with, 3 of 3 runs
41 t/s
tokens per second of prefill across the same sweep, per-step medians
579–621
from process start to first token with the model warm; 23.7 s cold (separate probe, raised limit)
3.7 s
On this page
  1. The short answer
  2. Why this benchmark
  3. The results
  4. Stock settings
  5. DeepSeek, streamed
  6. Cold start
  7. How other Macs differ
  8. The details
  9. Method
  10. Measured or inferred
  11. What went wrong
  12. Run it yourself

The short answer

Qwen3.8-Flash-Next is a 125B mixture-of-experts model with 6B parameters active per token, and since part three it has been the model behind my daily assistant on a 64GB Mac. Those notes measured it the way I use it: a long-running server, speculative decoding, real agent turns. Honest numbers, and impossible to hold up against anyone else’s machine. So this time I ran the benchmark that ships with DwarfStar , Salvatore Sanfilippo’s inference engine: the same text and the same 32 steps from 2K to 64K context that upstream asks every machine to run. The results are in PR 1142 and, with the per-run logs, in my bench repo .

On a Mac Studio M4 Max with 64 GB, the 2-bit Qwen reads a prompt at 579 to 621 tokens per second and writes at 40.7 to 41.5, per-step medians, and the line is flat: the 64K step is within 2 percent of the 4K step on both. It does that at the GPU memory limit macOS ships with, with one flag, --prefill-chunk 2048, in 3 of 3 runs, without swapping. That is what I mean by full speed: the stock limit gives the same numbers as a raised one, within 1 percent at every step, and from 4K to 64K the speed holds within about 1.5 percent.

The same Mac also runs DeepSeek V4 Flash, an 81 GiB model file, by streaming its experts from the SSD. It writes at 12.9 to 15.8 tokens a second and reads at about 135, except when it suddenly does not. It runs. I would not live with it.

Have a 64GB Mac? Run the same command and send me your numbers. How to run it

Why this benchmark and not my daily numbers

ds4-bench is DwarfStar’s own instrument. It feeds a fixed Italian novel into the model 2,048 tokens at a time, and after each step it measures how fast that step was read (prefill) and how fast the model then writes 128 tokens (generation). Every machine in upstream’s performance notes ran the same sweep, so a number from it means the same thing on an M5 Max, on the build it was recorded with, as on my M4 Max.

The decode speed in part three, 43 to 46 tokens a second, came from a server with --mtp speculative decoding on Ivan Fioravanti’s fork, on real Hermes turns. ds4-bench does not accept --mtp and measures plain decoding, so 41 here is not a regression. It is a different measurement, and the first reproducible ds4-bench measurement of this machine.

The results

Qwen repeats tightly. Of the 288 steps in the nine valid Qwen runs, every step except the two outliers below lies between 571 and 622 tokens per second of prefill and 40.1 and 41.6 of generation. The outliers are generation dips: 35.7 at 2K in the very first run of the day, and a single 38.6 at 12K in one evening run. Neither repeated.

Prefill and generation speed from 2K to 64K context. Qwen3.8 Flash Next holds about 580 to 620 tokens per second prefill and 41 generation; DeepSeek V4 Flash streamed from SSD holds 12.9 to 15.8 generation (per-step medians), with prefill drops in every run.Qwen3.8 Flash Next Q2, stock limit, chunk 2048DeepSeek V4 Flash Q2, streamed from SSDPrefill, tokens per second0200400600Qwen3.8 Flash Next, stock limitDeepSeek V4 Flash, streamedone run fell to 47Generation, tokens per second02040Qwen3.8 Flash Next, stock limitDeepSeek V4 Flash, streamed2K16K32K48K64Kcontext, tokens
The official ds4-bench sweep, median of three valid runs per line, shaded from the slowest to the fastest run at each step. Qwen’s band is almost invisible because the runs agree. DeepSeek’s prefill band is the story: every run dips somewhere, never at the same depth. Qwen ran at the macOS default memory limit; DeepSeek at the raised one.

Swipe to compare columns →

ds4-bench, prefill / generation tokens per second, median of valid runs
Model and setting (valid runs)2K16K32K64K
Qwen, default limit, --prefill-chunk 2048 (3 of 3)621.3 / 41.54590.1 / 41.14583.4 / 41.10579.4 / 40.70
Qwen, default limit, official command (1 probe run)615.8 / 41.29590.3 / 41.01581.3 / 40.74577.4 / 40.63
Qwen, raised limit, official command (3 of 3)616.4 / 41.24590.6 / 40.98581.8 / 40.81579.5 / 40.34
Qwen, raised limit, --prefill-chunk 2048 (2 of 4)621.3 / 41.15590.0 / 40.90583.4 / 40.84579.8 / 40.48
DeepSeek, raised limit, --ssd-streaming (3 of 4)156.0 / 12.85138.2 / 15.59138.0 / 14.83123.1 / 14.39

Neither the prefill chunk nor the memory limit changes Qwen’s speed on this machine: the default chunk and 2048, at the raised or the default limit, agree within about 1 percent at every step. So for a 64GB Mac, --prefill-chunk 2048 is free insurance, and the next section is why you want the insurance.

One thing makes the flat line possible. The Qwen file is 137 GiB, more than twice the Mac’s memory, but ds4 keeps only 41.7 GiB of it resident and reads the rest, rows of the model’s large lookup table, from the SSD as tokens need them. In the raised-limit Qwen runs, disk reads averaged about 160 MB/s, with peaks near 2.2 GB/s.

It runs at stock settings

I expected to need a raised GPU memory limit. ds4 prints its memory plan at startup, and the official command plans 52.05 GiB, more than macOS gives the GPU by default. Measured, stock settings work: at the default limit, which Metal reports as a 51.84 GiB working set on this Mac, --prefill-chunk 2048 ran 3 of 3 runs at full speed. The step medians matched the raised-limit runs within 0.3 percent for prefill and 1 percent for generation, and ds4-bench itself swapped nothing.

What ds4 plans to hold against the GPU limit (components rounded, so they may not sum exactly)
GiB
Qwen, official command (chunk 8192)52.05 GiB: model 41.72, buffers 8.24, KV 2.09
Qwen, --prefill-chunk 204846.11 GiB: model 41.72, buffers 2.31, KV 2.09
DeepSeek, --ssd-streaming48.36 GiB: 8.20 resident, 35.42 expert cache, 3.38 prefill reserve, 0.86 KV, 0.50 buffers
macOS default limit on this Mac51.84 GiB (55,662,788,608 bytes)
Raised limit, iogpu.wired_limit_mb=5734456.00 GiB

Chunk 2048 plans 46.11 GiB, leaves 5.7 GiB of room and cost nothing measurable here, so that is the setting for a stock 64GB Mac. The unmodified command also completed one run at the default limit, but with 0.21 GiB less budget than it planned for, so I would not rely on it. Why it fits anyway, I have not traced. My untested guess is that Metal’s working-set figure is a recommendation rather than a hard allocation limit.

DeepSeek V4 Flash, streamed from SSD

DeepSeek V4 Flash at 2 bits is an 81 GiB file, too big to hold in 64 GB. With --ssd-streaming, ds4 keeps 8.2 GiB of shared weights resident, caches 35.4 GiB of experts, and reads the rest from the SSD on demand, at up to 6.1 GB/s in these runs. All four runs completed the sweep; one was excluded by my swap rule, below.

  • Generation is steady at 12.9 to 15.8 tokens a second across the sweep, per-step medians.
  • Prefill is not. The median is about 135 tokens a second, but every run drops somewhere, never at the same depth. The worst was the day’s first DeepSeek run, which fell to 47 to 50 tokens a second from 57K to 63K. The one other run in the same position in the order, the one excluded below, did not show it. In the first run, the drop coincided with the memory compressor growing from about 2.3 to 5.4 GiB, and with pageouts. That is a correlation in one run, not a cause.
  • It squeezes the rest of the Mac. In the excluded run the system swapped in 403 MiB, about two thirds of it in a burst three to six minutes in, while ds4 itself swapped nothing. In that run, streaming pushed other processes out of memory.

It is a real result: a model bigger than the machine’s memory, answering at a readable pace. But a prefill that can collapse to a third of its speed on a long prompt is not what I want under an agent that reads web pages all day. Qwen stays my daily model.

Cold start, the part 64 GB changes

The official sweep does not time model loading, and on a 64GB machine loading is where the page cache fights you. So I added a separate probe, not part of the official format, run at the raised limit: start a fresh process, read a 2,048-token prompt, write one token. “Cold” is right after sudo purge, so no part of the model file is cached; “warm” is an immediate repeat. Three pairs each.

Swipe to compare columns →

Process start to first token, 2,048-token prompt, median (min to max)
Model and cacheStart to first tokenOf which loadingPrefillFirst decode
Qwen, cold (after sudo purge)23.7 s (23.6 to 24.6)20.3 s (20.2 to 21.2)3.3 s26 ms
Qwen, warm3.7 s (3.7 to 3.7)0.3 s (0.3 to 0.3)3.3 s26 ms
DeepSeek streamed, cold18.2 s (18.0 to 19.1)4.6 s (4.5 to 4.7)12.9 s597 ms
DeepSeek streamed, warm13.8 s (13.6 to 14.0)0.3 s (0.3 to 0.3)12.8 s630 ms

“Loading” is derived: the total minus prefill minus the first token. Qwen’s cold penalty, about 20 seconds, is most likely reading 41.7 GiB of weights from the SSD; that is inference from the arithmetic, not a trace. DeepSeek starts faster cold because only 8.2 GiB is resident, then pays for it on every request: 12.9 seconds to read 2,048 tokens against Qwen’s 3.3, and about 0.6 seconds for its first decoded token against 26 milliseconds.

For a daily server the lesson is simple. Qwen’s load is paid once, when the server starts. DeepSeek’s streaming cost is paid on every turn. I did not measure first-token time inside a long-running server, which is what an agent user actually feels; there the model is always warm.

How other Macs should differ

Three properties of the machine decide these numbers, and a fourth decides whether a model runs at all. The first two are the usual rule of thumb for local inference, not something this benchmark measured:

  • Generation is mostly bound by memory bandwidth. Every new token reads the model’s active weights from memory once, so tokens per second tracks how fast memory can be read.
  • Prefill is mostly bound by GPU compute. Reading a prompt processes thousands of tokens against the same weights at once, so it tracks GPU cores and how fast each one is.
  • Capacity decides resident or streamed. If the working set fits in unified memory, the model runs resident. If not, as with DeepSeek here, it streams from the SSD, and then the SSD’s read speed sets the pace. Apple does not publish SSD speeds on these pages; streaming DeepSeek peaked at 6.1 GB/s of reads on this Mac, measured.
  • An Ultra is two Max dies joined. Apple builds M1 Ultra by connecting two M1 Max dies with UltraFusion, which doubles both: 800 GB/s against 400, and 64 GPU cores against 32.

Swipe to compare columns →

Apple's published specs, not measured here
ChipBandwidthGPU coresMemoryApple source
M1 Max400 GB/s24 or 3232 or 64 GBMacBook Pro 2021 specs
M1 Ultra, two M1 Max dies800 GB/s64up to 128 GBNewsroom, March 2022
M4 Max, 32-core GPU410 GB/s3236 GBMac Studio 2025 specs
M4 Max, 40-core GPU (this Mac)546 GB/s4048 to 128 GBMac Studio 2025 specs
M3 Ultra819 GB/s60 or 8096 or 256 GBMac Studio 2025 specs

M1 to M4, one data point. PR 1068 ran the same Qwen model on an M1 Max 64 GB with the same --prefill-chunk 2048. At 16K:

Swipe to compare columns →

M1 Max 64 GB (PR 1068) against this M4 Max, 16K context
M1 MaxM4 Max, 40-coreRatio
Prefill at 16K~274 t/s590.1 t/s2.15x
Generation at 16K (steady)~24.3 t/s41.2 t/s1.70x
Memory bandwidth (Apple spec)400 GB/s546 GB/s1.37x

Generation is 1.7 times faster while the bandwidth is 1.37 times higher, and prefill is 2.15 times faster. That is consistent with bandwidth mattering for generation and compute for prefill, but it cannot isolate either: the two runs are 25 engine commits apart, the M1 run stopped at 16K, and its GPU core count, 24 or 32, is not stated. Both M1 Max variants have the same 400 GB/s. Apple’s own claim for the M4 Max GPU is up to 1.9 times the M1 Max.

Max against Ultra. By the rule of thumb, a resident model on an Ultra should generate up to about twice as fast as on the Max of the same generation, and prefill should scale with the doubled GPU. In the 2025 Mac Studio the M3 Ultra, a generation older, still has 1.5 times the M4 Max’s bandwidth, 819 GB/s against 546. I have not measured an Ultra. The only 64 GB Ultra data point I found upstream, the closed PR 350 , streamed DeepSeek on an M1 Ultra at about 5 tokens a second, slower than this M4 Max at 12.9 to 15.8, on an older engine commit with a smaller manual cache. A streamed model is limited by the SSD and the cache, and twice the bandwidth does not help much there.

Capacity changes the question. Upstream’s 128 GB rows, an M5 Max and a DGX Spark, hold the same DeepSeek model entirely in memory. On 64 GB it can only be streamed. Those rows are not a faster version of this one; they are a different way of running the model.

The details

Everything below is for people who want to check the numbers or run the same thing. If you only wanted the result, you have it.

Method

The engine was upstream antirez/ds4 at 0aaea5a, the tip of main on 28 September, built with plain make in a separate clean checkout, with zero warnings. Both model files came from upstream’s own download_model.sh and were checked against their Hugging Face sha256 before the first run. The command was the documented one, run from the checkout:

./ds4-bench -m MODEL \
  --prompt-file speed-bench/promessi_sposi.txt \
  --ctx-start 2048 --ctx-max 65536 --step-incr 2048 \
  --gen-tokens 128 [--prefill-chunk 2048 | --ssd-streaming]
  • A quiet machine. My daily model server was stopped and all six of Hermes’s scheduled jobs paused for the window; the desktop chat app was quit. Still running: two Claude Code sessions, one of them supervising the runs, an idle agent gateway and macOS’s own services. Load average was 1.3 to 1.6 at the start.
  • One machine, three runs per setting, in alternating order (forward, reversed, forward), with two minutes of cooldown after each. A Qwen sweep takes about 25 minutes.
  • A validity rule set before the first run: exit code 0, all 32 rows present, and at most 100 MiB swapped in system-wide during the run. System-wide on purpose, so it catches a machine that pages at all, not just the benchmark.
  • Sampling every five seconds: memory counters, disk throughput, swap and thermal state, with /usr/bin/time -l around each run for the process’s own swaps.

The full protocol, including the replacement rule written before the replacement batch, is in the repo’s METHODOLOGY.md .

What is measured and what is not

How each claim in this note is known
ClaimBasis
Prefill and generation speeds, planned memory, swap counts, first-token timesMeasured: raw CSVs, engine reports, time -l
The default GPU limit on this Mac is 51.84 GiBMeasured: Metal's recommendedMaxWorkingSetSize at iogpu.wired_limit_mb=0
--prefill-chunk 2048 runs at full speed at the default limitMeasured: 3 of 3 runs
The official command also fits at the default limitMeasured once: 1 probe run, 0.21 GiB over the limit on paper
Replay beyond ~30K does not affect the timed numbersDerived: the timers in ds4_bench.c bracket only prefill and generation
The excluded runs' swap-ins came from other processesDerived: time -l shows 0 swaps for ds4-bench while the system counter rose
Qwen's 20 s cold penalty is reading its weights from SSDInferred: 41.7 GiB at ~2 GB/s is about 20 s; not traced
Memory bandwidth, GPU core counts, Ultra = two Max diesApple's tech-specs and newsroom pages
Generation tracks bandwidth and prefill tracks GPU computeThe usual rule of thumb; consistent with, not proven by, the M1 Max comparison
DeepSeek's prefill drops come from memory pressureHypothesis: correlated with compressor growth in one run only

What went wrong

Three runs failed my own swap rule and are excluded: two Qwen chunk-2048 runs, at 137 and 168 MiB, and the DeepSeek run at 403. In all three, ds4-bench itself swapped nothing, and their numbers match the valid runs, so the rule threw out good data. I kept it anyway: a rule you relax after seeing the results is not a rule. I ran one replacement per affected setting, under a note written before the batch, and stopped there. That is why one row says 2 of 4.

One rule did change after I saw the data, and it is the one worth flagging. The charts and the upstream submission use one representative run per setting. My first rule picked the run closest to the median generation speed, and for DeepSeek that chose the run with the deepest prefill drop, which would have misrepresented prefill. I switched to the run closest to the medians of both prefill and generation at every step, for every setting, and the methodology says so.

Who did what: this benchmark was done with Claude Code , running on the same Mac. It made the run plan, wrote the driver and the probes, ran the batches, did the analysis and wrote it up, including this note and the upstream PR. I directed the work, set the quiet-machine windows, reviewed the results and approved every change to the machine. Every published number was re-derived from the raw CSVs by a separate check. The mistakes are its as much as mine.

Run it on your 64GB Mac

If you have a 64GB Apple Silicon Mac, I would like your numbers. The benchmark alone is a handful of commands:

git clone https://github.com/antirez/ds4 && cd ds4
git checkout 0aaea5a && make ds4-bench
./download_model.sh qwen38-q2     # 137 GiB
./ds4-bench -m gguf/Qwen3.8-Flash-Next-Q2.gguf \
  --prompt-file speed-bench/promessi_sposi.txt \
  --ctx-start 2048 --ctx-max 65536 --step-incr 2048 \
  --gen-tokens 128 --prefill-chunk 2048

The bench repo has the driver that quiets the machine, samples memory and applies the validity rule, and every per-run CSV, engine report and memory sample behind this note is in its 2026-09-28 results . The data is CC BY 4.0 and the scripts are MIT. Stop anything else that holds GPU memory first, then send me your numbers on X , or upstream the way PR 1068 and PR 1142 did.

Sources

Measured on a Mac Studio M4 Max with 64 GB, macOS 26.6.2, on 28 September 2026, using upstream DwarfStar 0aaea5a.