Skip to content

Local inference field note / October 2026

I tested 14 open models against Jev on my Mac. Two kept pace.

I tested Jev through its API, then ran fourteen open model builds on my M4 Max with 64 GB. On my own extraction check, two local models couldn’t be told apart from Jev on this 300-line sample. One is the model I already run every day, with no decision training. The other is Kev 9B, a dedicated decision model in 31 GiB.

Ranking quality on my extraction check (1.0 is perfect)

Qwen3.8 Flash-Next, the model my daily server already runs, with no training
0.97
Jev, tested through its API
0.95
Kev 9B, a dedicated decision model in 31 GiB
0.93
On this page
  1. The short answer
  2. What I tested
  3. My task
  4. Where every model fails
  5. The public benchmark
  6. Speed
  7. What I'd use
  8. The details

The short answer

  • Two local models couldn’t be told apart from Jev on my task. Qwen3.8 Flash-Next , the model my daily server already runs, was given a plain lettered-answer prompt with no training. Kev 9B is a small dedicated decision model. On 300 lines, neither can be told apart from Jev.
  • For Flash-Next, that call is borderline. It scored higher than Jev, but the statistical test only just includes zero. Another run could tip it either way.
  • Kev 9B is the practical one. It needs 31 GiB. Flash-Next needs about 44 GiB, most of the machine.
  • Jev and Flash-Next both missed all ten day/month swaps that still made a valid date.
  • A dozen-line script still beats all of them on this task, because the check is exact string matching.

What I tested

A decision model takes a piece of text and a typed question, and returns a probability for each answer. Jev is TypeSafe’s hosted one. The open ones run locally.

I used two tests:

  1. My own task. 300 synthetic bank-statement lines, each paired with an extracted record. Half the records are faithful. Half carry one injected error, such as a swapped digit or a wrong date. The question is whether the record matches its line. It is the same set I used for the Laya note .
  2. A public benchmark. Typed Decisions has 400 cases and 2,000 questions across four business workflows. Its labels come from a roughly 4B-class teacher model, so it measures agreement with that teacher.

Ranking quality here means AUC. It is how often a faithful line gets a higher probability than a corrupted one. 0.5 is a coin flip; 1.0 is perfect.

My task: who can tell a good record from a bad one

Most local models can. Flash-Next matches Jev, and Kev 9B is close enough that 300 lines can’t separate them.

Swipe to compare columns →

Ranking quality on my extraction check, four field questions, 300 lines
ModelWhere it runsRanking quality, four field questionsMemory (GiB)
Qwen3.8 Flash-Next (2-bit), no trainingmy Mac0.970~44 (planned)
JevTypeSafe's API0.949–
Kev 9Bmy Mac0.93231 (17 steady)
StartLux-Decision 27B, 8-bitmy Mac0.89331 (over 10 requests)
Decider 4Bmy Mac0.87813 (PyTorch MPS runtime)
Clef 27B, 8-bitmy Mac0.85431
Clef-flash, 8-bitmy Mac0.85113
Laya typed-decisionsmy Mac0.6651

StartLux-Decision by StartLux Labs, CC BY-NC 4.0.

local; area = memoryJev, hosted; time includes the network from Singapore0.60.70.80.91.010 ms30 ms100 ms300 ms1 stime per call, log scale →↑ ranking qualityJev’s scoreFlash-NextJevKev 9BStartLux 27BDecider 4BClef 27BClef-flashLaya, timed another day
Flash-Next and Kev 9B couldn’t be told apart from Jev on this sample; the rest sit below it. Whiskers: per-model 95% intervals. Whether a model can be told apart from Jev was judged by paired comparison (see details).
Chart data
ModelRanking quality (95% interval)Time per callMemory
Qwen3.8 Flash-Next (2-bit), no training0.970 (0.951–0.985)0.71 s, per question (one chat request per question)44 GiB planned
Jev0.949 (0.918–0.975)0.27 s, from Singapore, includes networkhosted
Kev 9B0.932 (0.898–0.962)0.23 s, re-measured 2026-10-06 with the MLX cache capped; same as 2026-10-05 (232.8 ms)31 GiB
StartLux-Decision 27B, 8-bit0.893 (0.854–0.925)0.87 s31 GiB
Decider 4B0.878 (0.836–0.917)0.19 s, PyTorch MPS, empty_cache after every request (2026-10-06); 2026-10-05 without it: 180.2 ms13 GiB
Clef 27B, 8-bit0.854 (0.810–0.892)1.14 s, re-measured 2026-10-06 with the MLX cache capped; same as 2026-10-05 (1138.2 ms)31 GiB
Clef-flash, 8-bit0.851 (0.806–0.891)0.33 s, re-measured 2026-10-06 with the MLX cache capped; same as 2026-10-05 (325.1 ms)13 GiB
Laya typed-decisions0.665 (0.602–0.722)16 ms, GPU run 2026-09-21 with the daily server loaded and idle1.1 GiB MLX peak

StartLux-Decision by StartLux Labs, CC BY-NC 4.0.

Flash-Next and Kev 9B couldn’t be told apart from Jev on this sample. StartLux 27B is measurably behind it.

Flash-Next needs no decision training. I letter each answer option (A for yes, B for no) and read the probability of each letter from the first token it would write. That is the simplest readout there is, and it was enough.

Where every model fails

Swapped dates are the blind spot. On the date question, Jev and Flash-Next caught none of ten day/month swaps that still made a valid date. Only Clef-flash caught any, three of them. It also flagged 59 of 150 correct dates.

They did better on swaps that produced a month above 12. Jev caught 9 of 14, Flash-Next 5.

Jev takes “exactly” literally. My question asked whether each record matches its line exactly. Jev flagged 20 of 150 faithful lines, mostly because the record drops a thousands comma. Flash-Next flagged 7. If you want formatting differences ignored, say so in the question.

The public benchmark: a different winner

On Typed Decisions, StartLux leads.

Swipe to compare columns →

Accuracy on the public Typed Decisions test set, 400 cases, 2,000 questions
ModelAccuracyNotes
StartLux-Decision 27B, 8-bit0.804above the benchmark's 0.735 ceiling; trained on its train split
StartLux-Decision 4B0.798above the ceiling; trained on its train split
StartLux-Decision 0.8B0.758above the ceiling; trained on its train split
Jev0.733
Qwen3.8 Flash-Next (2-bit), no training0.728
Clef 27B, 8-bit0.722
Clef-flash, 8-bit0.704
Kev 9B0.689
Decider 4B0.680

StartLux-Decision by StartLux Labs, CC BY-NC 4.0.

My Mac came within 0.002 of StartLux’s published scores, with the 27B as an 8-bit conversion against its published bf16. The 0.735 ceiling is the teacher’s own self-agreement. The benchmark’s card says scores well above it mean a model is learning the teacher’s quirks. I asked StartLux’s maker whether this benchmark was in its training. They told me on X that its training split was and its test split was not. Its larger models’ scores sit in the same band as the models the benchmark’s card lists as fine-tuned on that split, which may explain much of the lead. My test lines are new and unpublished, so none of these models could have trained on them; on those lines, StartLux 27B sits behind Jev and Kev 9B.

Jev scored 0.73 here, close to the 0.727 on the benchmark’s leaderboard. Flash-Next again couldn’t be told apart from it.

Speed

Jev answers in about a quarter of a second from Singapore. Most small local models are faster; Clef-flash is not. Flash-Next and the 27B models are slower.

Median time per call, broad question
ModelTime per call, broad question
StartLux-Decision 0.8B29 ms
Kev 9B0.23 s
Jev, from Singapore0.27 s
Clef-flash, 8-bit0.33 s
Qwen3.8 Flash-Next (2-bit), per question0.71 s
StartLux-Decision 27B, 8-bit0.87 s
Clef 27B, 8-bit1.14 s

StartLux-Decision by StartLux Labs, CC BY-NC 4.0.

What I’d use

  • For exact fields, code. A script that compares strings scored 0.99.
  • For fuzzy fields on this Mac, Kev 9B. It couldn’t be told apart from Jev, needs 31 GiB at peak, and is a dedicated decision model.
  • For batch work, the model I already run. It also couldn’t be told apart from Jev, and needs nothing new beyond a small test patch. But it takes most of the machine.
  • For a sidecar beside the daily model, nothing yet. The models small enough to sit beside it, the 0.8B builds (StartLux 0.8B measured 4.5 GiB), topped out at about 0.81 on my task.StartLux-Decision by StartLux Labs, CC BY-NC 4.0.

The details

Method, caveats and what was measured
  • Runs: every local run used pinned revisions and one harness. StartLux , Decider and Kev used the makers’ own code. Clef used community 8-bit ports, and Flash-Next used my own readout. The Typed Decisions scorer reproduces the card’s uniform-baseline KL and Brier exactly. Large models ran in a quiet window, with my daily server stopped and scheduled jobs paused.
  • Flash-Next: a 36-line test patch to the ds4 server exposed first-token probabilities. The daily build was not touched. Its 44 GiB is the server’s own memory plan, not a measurement. The memory tool cannot see mapped model files. It is the model from my daily setup .
  • The Clef models ran as community 8-bit MLX conversions of Cloudflare’s Clef and Clef-flash . I checked the smaller one against Cloudflare’s own code on 50 cases: they picked the same answer on 214 of 220 questions, and every disagreement was a near-tie. On the first run both peaked at about 54 GiB, and the machine swapped. MLX’s buffer cache was the cause: its default limit here is 61 GiB. A rerun with the cache capped at 2 GiB peaked at 31 GiB for the 27B and 13 GiB for Clef-flash, with identical answers and the same speed. The table shows the capped figures.
  • Memory, measured on equal terms: the Kev, Decider and Clef figures are peak footprints with the model’s cache capped or cleared: MLX capped at 2 GiB, or PyTorch’s cache emptied after every request for Decider. Flash-Next’s figure is the server’s own plan, and Laya’s is an MLX peak from its September run. Kev 9B’s 31 GiB did not change under the cap. It is a one-time peak, most likely at load; it then runs in about 17 GiB. StartLux 27B was measured over only 10 requests, which is likely before any cache growth (my inference, not re-measured with a cap).
  • StartLux ran with its own MLX server. The 27B was converted to 8-bit, as its docs describe. StartLux-Decision by StartLux Labs, CC BY-NC 4.0.
  • My test set has four mislabelled lines and fourteen easy date swaps; see the Laya note . Dropping them leaves the top four unchanged (Flash-Next, Jev, Kev 9B, StartLux 27B). Lower down, Clef 27B moves above Decider 4B .
  • Not tested: CLM-8B, which needs a vLLM backend; Tev1, which has no licence yet; OpenAI’s Decisions API, which is in limited preview.
  • Stats: on 300 lines, a ranking-quality score is good to roughly ±0.02 to ±0.04, depending on the model (±0.06 for Laya). Gaps were tested with paired bootstraps: neither Flash-Next nor Kev 9B could be told apart from Jev; Jev is ahead of StartLux 27B.
  • Credits: harness and analysis built with Claude Code .