Local inference field note / October 2026
I tested 14 open models against Jev on my Mac. Two kept pace.
I tested Jev through its API, then ran fourteen open model builds on my M4 Max with 64 GB. On my own extraction check, two local models couldn’t be told apart from Jev on this 300-line sample. One is the model I already run every day, with no decision training. The other is Kev 9B, a dedicated decision model in 31 GiB.
Ranking quality on my extraction check (1.0 is perfect)
- Qwen3.8 Flash-Next, the model my daily server already runs, with no training
- 0.97
- Jev, tested through its API
- 0.95
- Kev 9B, a dedicated decision model in 31 GiB
- 0.93
On this page
The short answer
- Two local models couldn’t be told apart from Jev on my task. Qwen3.8 Flash-Next , the model my daily server already runs, was given a plain lettered-answer prompt with no training. Kev 9B is a small dedicated decision model. On 300 lines, neither can be told apart from Jev.
- For Flash-Next, that call is borderline. It scored higher than Jev, but the statistical test only just includes zero. Another run could tip it either way.
- Kev 9B is the practical one. It needs 31 GiB. Flash-Next needs about 44 GiB, most of the machine.
- Jev and Flash-Next both missed all ten day/month swaps that still made a valid date.
- A dozen-line script still beats all of them on this task, because the check is exact string matching.
What I tested
A decision model takes a piece of text and a typed question, and returns a probability for each answer. Jev is TypeSafe’s hosted one. The open ones run locally.
I used two tests:
- My own task. 300 synthetic bank-statement lines, each paired with an extracted record. Half the records are faithful. Half carry one injected error, such as a swapped digit or a wrong date. The question is whether the record matches its line. It is the same set I used for the Laya note .
- A public benchmark. Typed Decisions has 400 cases and 2,000 questions across four business workflows. Its labels come from a roughly 4B-class teacher model, so it measures agreement with that teacher.
Ranking quality here means AUC. It is how often a faithful line gets a higher probability than a corrupted one. 0.5 is a coin flip; 1.0 is perfect.
My task: who can tell a good record from a bad one
Most local models can. Flash-Next matches Jev, and Kev 9B is close enough that 300 lines can’t separate them.
Swipe to compare columns →
| Model | Where it runs | Ranking quality, four field questions | Memory (GiB) |
|---|---|---|---|
| Qwen3.8 Flash-Next (2-bit), no training | my Mac | 0.970 | ~44 (planned) |
| Jev | TypeSafe's API | 0.949 | – |
| Kev 9B | my Mac | 0.932 | 31 (17 steady) |
| StartLux-Decision 27B, 8-bit | my Mac | 0.893 | 31 (over 10 requests) |
| Decider 4B | my Mac | 0.878 | 13 (PyTorch MPS runtime) |
| Clef 27B, 8-bit | my Mac | 0.854 | 31 |
| Clef-flash, 8-bit | my Mac | 0.851 | 13 |
| Laya typed-decisions | my Mac | 0.665 | 1 |
StartLux-Decision by StartLux Labs, CC BY-NC 4.0.
Chart data
| Model | Ranking quality (95% interval) | Time per call | Memory |
|---|---|---|---|
| Qwen3.8 Flash-Next (2-bit), no training | 0.970 (0.951–0.985) | 0.71 s, per question (one chat request per question) | 44 GiB planned |
| Jev | 0.949 (0.918–0.975) | 0.27 s, from Singapore, includes network | hosted |
| Kev 9B | 0.932 (0.898–0.962) | 0.23 s, re-measured 2026-10-06 with the MLX cache capped; same as 2026-10-05 (232.8 ms) | 31 GiB |
| StartLux-Decision 27B, 8-bit | 0.893 (0.854–0.925) | 0.87 s | 31 GiB |
| Decider 4B | 0.878 (0.836–0.917) | 0.19 s, PyTorch MPS, empty_cache after every request (2026-10-06); 2026-10-05 without it: 180.2 ms | 13 GiB |
| Clef 27B, 8-bit | 0.854 (0.810–0.892) | 1.14 s, re-measured 2026-10-06 with the MLX cache capped; same as 2026-10-05 (1138.2 ms) | 31 GiB |
| Clef-flash, 8-bit | 0.851 (0.806–0.891) | 0.33 s, re-measured 2026-10-06 with the MLX cache capped; same as 2026-10-05 (325.1 ms) | 13 GiB |
| Laya typed-decisions | 0.665 (0.602–0.722) | 16 ms, GPU run 2026-09-21 with the daily server loaded and idle | 1.1 GiB MLX peak |
StartLux-Decision by StartLux Labs, CC BY-NC 4.0.
Flash-Next and Kev 9B couldn’t be told apart from Jev on this sample. StartLux 27B is measurably behind it.
Flash-Next needs no decision training. I letter each answer option (A for yes, B for no) and read the probability of each letter from the first token it would write. That is the simplest readout there is, and it was enough.
Where every model fails
Swapped dates are the blind spot. On the date question, Jev and Flash-Next caught none of ten day/month swaps that still made a valid date. Only Clef-flash caught any, three of them. It also flagged 59 of 150 correct dates.
They did better on swaps that produced a month above 12. Jev caught 9 of 14, Flash-Next 5.
Jev takes “exactly” literally. My question asked whether each record matches its line exactly. Jev flagged 20 of 150 faithful lines, mostly because the record drops a thousands comma. Flash-Next flagged 7. If you want formatting differences ignored, say so in the question.
The public benchmark: a different winner
On Typed Decisions, StartLux leads.
Swipe to compare columns →
| Model | Accuracy | Notes |
|---|---|---|
| StartLux-Decision 27B, 8-bit | 0.804 | above the benchmark's 0.735 ceiling; trained on its train split |
| StartLux-Decision 4B | 0.798 | above the ceiling; trained on its train split |
| StartLux-Decision 0.8B | 0.758 | above the ceiling; trained on its train split |
| Jev | 0.733 | |
| Qwen3.8 Flash-Next (2-bit), no training | 0.728 | |
| Clef 27B, 8-bit | 0.722 | |
| Clef-flash, 8-bit | 0.704 | |
| Kev 9B | 0.689 | |
| Decider 4B | 0.680 |
StartLux-Decision by StartLux Labs, CC BY-NC 4.0.
My Mac came within 0.002 of StartLux’s published scores, with the 27B as an 8-bit conversion against its published bf16. The 0.735 ceiling is the teacher’s own self-agreement. The benchmark’s card says scores well above it mean a model is learning the teacher’s quirks. I asked StartLux’s maker whether this benchmark was in its training. They told me on X that its training split was and its test split was not. Its larger models’ scores sit in the same band as the models the benchmark’s card lists as fine-tuned on that split, which may explain much of the lead. My test lines are new and unpublished, so none of these models could have trained on them; on those lines, StartLux 27B sits behind Jev and Kev 9B.
Jev scored 0.73 here, close to the 0.727 on the benchmark’s leaderboard. Flash-Next again couldn’t be told apart from it.
Speed
Jev answers in about a quarter of a second from Singapore. Most small local models are faster; Clef-flash is not. Flash-Next and the 27B models are slower.
| Model | Time per call, broad question |
|---|---|
| StartLux-Decision 0.8B | 29 ms |
| Kev 9B | 0.23 s |
| Jev, from Singapore | 0.27 s |
| Clef-flash, 8-bit | 0.33 s |
| Qwen3.8 Flash-Next (2-bit), per question | 0.71 s |
| StartLux-Decision 27B, 8-bit | 0.87 s |
| Clef 27B, 8-bit | 1.14 s |
StartLux-Decision by StartLux Labs, CC BY-NC 4.0.
What I’d use
- For exact fields, code. A script that compares strings scored 0.99.
- For fuzzy fields on this Mac, Kev 9B. It couldn’t be told apart from Jev, needs 31 GiB at peak, and is a dedicated decision model.
- For batch work, the model I already run. It also couldn’t be told apart from Jev, and needs nothing new beyond a small test patch. But it takes most of the machine.
- For a sidecar beside the daily model, nothing yet. The models small enough to sit beside it, the 0.8B builds (StartLux 0.8B measured 4.5 GiB), topped out at about 0.81 on my task.StartLux-Decision by StartLux Labs, CC BY-NC 4.0.
The details
Method, caveats and what was measured
- Runs: every local run used pinned revisions and one harness. StartLux , Decider and Kev used the makers’ own code. Clef used community 8-bit ports, and Flash-Next used my own readout. The Typed Decisions scorer reproduces the card’s uniform-baseline KL and Brier exactly. Large models ran in a quiet window, with my daily server stopped and scheduled jobs paused.
- Flash-Next: a 36-line test patch to the ds4 server exposed first-token probabilities. The daily build was not touched. Its 44 GiB is the server’s own memory plan, not a measurement. The memory tool cannot see mapped model files. It is the model from my daily setup .
- The Clef models ran as community 8-bit MLX conversions of Cloudflare’s Clef and Clef-flash . I checked the smaller one against Cloudflare’s own code on 50 cases: they picked the same answer on 214 of 220 questions, and every disagreement was a near-tie. On the first run both peaked at about 54 GiB, and the machine swapped. MLX’s buffer cache was the cause: its default limit here is 61 GiB. A rerun with the cache capped at 2 GiB peaked at 31 GiB for the 27B and 13 GiB for Clef-flash, with identical answers and the same speed. The table shows the capped figures.
- Memory, measured on equal terms: the Kev, Decider and Clef figures are peak footprints with the model’s cache capped or cleared: MLX capped at 2 GiB, or PyTorch’s cache emptied after every request for Decider. Flash-Next’s figure is the server’s own plan, and Laya’s is an MLX peak from its September run. Kev 9B’s 31 GiB did not change under the cap. It is a one-time peak, most likely at load; it then runs in about 17 GiB. StartLux 27B was measured over only 10 requests, which is likely before any cache growth (my inference, not re-measured with a cap).
- StartLux ran with its own MLX server. The 27B was converted to 8-bit, as its docs describe. StartLux-Decision by StartLux Labs, CC BY-NC 4.0.
- My test set has four mislabelled lines and fourteen easy date swaps; see the Laya note . Dropping them leaves the top four unchanged (Flash-Next, Jev, Kev 9B, StartLux 27B). Lower down, Clef 27B moves above Decider 4B .
- Not tested: CLM-8B, which needs a vLLM backend; Tev1, which has no licence yet; OpenAI’s Decisions API, which is in limited preview.
- Stats: on 300 lines, a ranking-quality score is good to roughly ±0.02 to ±0.04, depending on the model (±0.06 for Laya). Gaps were tested with paired bootstraps: neither Flash-Next nor Kev 9B could be told apart from Jev; Jev is ahead of StartLux 27B.
- Credits: harness and analysis built with Claude Code .