Local inference field note / October 2026
I tested Jev against a free local model. Ten lines of code beat both.
One narrow task, 300 synthetic statement lines. Jev, tested through its API, was accurate. Laya, an open model I ran locally, was fast but too flat to use. A short script beat both, because it knows my generator’s exact line layout.
Accuracy on the same 300 lines, one broad question
- the short comparator, which reads my generator's fixed line layout
- 0.99
- Jev, tested through its API
- 0.89
- Laya, run locally on the Mac: about 17 times faster than Jev per single-question call from Singapore
- 0.59
On this page
The short answer
Jev, an AI model that returns decisions instead of chatting, was everywhere in September. It is TypeSafe AI’s hosted model. Laya is an open-weight model of the same kind, and an independent port runs it on Apple Silicon. I put both on one narrow job: 300 synthetic statement lines, each with a record extracted from it. Does the record match the line?
- Jev was accurate. With one broad question it was right on 89 percent of lines. With four per-field questions it raised no false alarm on any of the 150 faithful lines.
- Laya was fast, local and free, and too flat to use. It answered in 8 to 16 milliseconds on my M4 Max, in about a gigabyte of memory. But it barely told good lines from bad: it scored 0.59. Its own card calls it “a fast base to specialise, not a zero-shot decision engine”, and this checkpoint was trained for four other workflows.
- Ten lines of code beat both. A script of about a dozen lines scored 0.99, because it knows my generator’s exact column layout. Neither model had that advantage. For exact numbers in a known format, plain code is the right tool, and TypeSafe’s own documentation says so too.
What I tested
A language model writes text. A decision model does not. You give it some text and a typed question, such as pick one, score this, or true or false. It returns the answer with a probability, in one pass.
- Jev is TypeSafe AI’s hosted, pay-per-use model, released in limited early access on 15 September . I tested it through its API on 1 October, pinned to version
jev-1.13.0. - Laya is an open-weight model from Convai Innovations, licensed Apache 2.0. In this note “Laya” means its typed-decisions checkpoint , unless I say otherwise. I ran it on my Mac on 21 September.
- laya-mlx is an independent port of Laya to Apple’s MLX framework, first published on 19 September. Its README says plainly that it is not an official Convai release.
The task: given a statement line and a record extracted from it, does the record match the line? A document-extraction job would ask that thousands of times, which is exactly the kind of question a decision model promises to handle in bulk.
The 300 lines are synthetic and seeded. Half are faithful. The other half carry one injected error each, of eight kinds: swapped digits, a shifted decimal point, an amount off by one cent, the balance copied in as the amount, a wrong day, a day and month swapped, debit and credit flipped, or the wrong merchant.
Both models got the same lines, the same input text and the same scoring code. Each was asked two ways: one broad question (“is this extraction accurate?”), and four true-or-false questions, one per field.
Results: Jev reads the line, Laya barely does
Jev separated faithful lines from corrupted ones; Laya barely did. The clearest single measure is AUC: how often a random faithful line gets a higher probability than a random corrupted one. 0.5 is a coin flip and 1.0 is perfect. On the broad question, Jev scored 0.95 against Laya’s 0.66.
Swipe to compare columns →
| Run | Accuracy | AUC | Mean P, faithful | Mean P, one error |
|---|---|---|---|---|
| One broad question, typed-decisions 421M | 0.593 | 0.660 | 0.535 | 0.480 |
| One broad question, multilingual 322M | 0.503 | 0.646 | 0.944 | 0.892 |
| One broad question, Jev | 0.887 | 0.951 | 0.629 | 0.110 |
| Four per-field questions, typed-decisions 421M | 0.619 | 0.665 | 0.513 | 0.488 |
| Four per-field questions, Jev | 0.965 | 0.949 | 0.954 | 0.737 |
Per-field accuracy is over 1,200 answers, and always answering “yes” would score 0.875 there. For the per-field runs the means and AUC use each line’s average over its four answers.
The gap is not noise. On the lines where the two disagreed on the broad question, Jev was right on 96 and Laya on 8. Jev’s per-field probabilities are also well calibrated, and asking the same thing as a choice question gave much the same result.
My own generator mislabelled four of the 300 lines. Dropping them moves no figure by more than 0.016; the details are below.
Why Laya was flat
Laya was unsure rather than wrong. Its average probability for faithful lines sat only about 0.05 above the one for corrupted lines. That gap is small but real: its 178 right answers out of 300 are well above chance. Asked field by field, it leaned heavily towards “no”, as the per-field details show.
The other checkpoint I tried, the multilingual one, said “accurate” to almost everything. At the default threshold that scores 0.50, though its ranking still carried some signal.
And Laya was asked something it was not trained for. Its base card calls it “a fast base to specialise, not a zero-shot decision engine”. The checkpoint I used is that base fine-tuned for four synthetic workflows: agent traces, customer service, invoices and security incidents. Invoice reconciliation is a near neighbour of my task, but the specialisation did not carry over. This is a specialist tested outside its domain.
Where Jev falls short
- Speed. A Jev call took 271 ms at the median from Singapore, against Laya’s 8 to 16 ms on the Mac. About 46 ms of that is the trip to the nearest network edge. The API reports no server timing, so I can’t split the rest. The gap shrinks with more questions per call: four questions took Jev the same time as one, and Munro found the API faster than Laya above three or four.
- Dates, like Laya. A day and month swapped into another valid date got past it: the per-field questions caught 0 of 10. TypeSafe’s own list of known weak spots notes that this version treats dates as text rather than as quantities to compare.
- Literal reading. The broad question asks for an “exact” match, and Jev flagged 20 of the 150 faithful lines. Most of them had a thousands comma in the amount. But every comma amount in my set is a deposit, so I cannot tell the comma from the column (untested). TypeSafe’s weak-spots page lists literal reading too.
- The same question twice is not quite the same answer. I ran the broad question twice, and 6 of 300 decisions flipped. All six sat right at the 0.5 threshold. The API has no temperature or seed setting to stop it.
- It leaves the machine. Anything you would not send to a hosted API is ruled out, whatever the score. Laya runs on the Mac.
Ten lines of code beat both
Plain code was the best checker here. A script of about a dozen lines compares the record with the line field by field: the amount, the date, the column and the description. It scored 0.99 on the broad question. Its only “errors” were the four lines my generator mislabelled.
It had an advantage neither model had: it knows the fixed column layout my generator writes, so it can cut the fields out of the line by position. A job that holds the raw line in a known format can do the same. A statement in a layout it has never seen would break it, while the models read the line as text.
So when the source line is available in a known format, comparing amounts, dates and columns in code is exact and free. TypeSafe’s documentation agrees: its guide to building with Jev recommends keeping deterministic rules in code, and its weak-spots page recommends doing any mathematical logic in code. A decision model earns its place on the fuzzy checks a script cannot settle, such as a merchant name written three ways. These 300 lines barely test that.
What’s next
Laya’s score here is a zero-shot floor, not a verdict. Next I want to fine-tune it on my Mac for something fuzzier: spotting low-effort “slop” writing, with my local 125B model as the teacher, measured against my own hand labels.
Details
Everything below is for people who want to check the numbers. If you only wanted the verdict, you have it.
A. Full results
The checks behind “the gap is not noise”, the calibration errors, the choice-question variant and the comparator’s exact scores.
| Check | Result |
|---|---|
| Jev against Laya, broad question, lines where they disagree | 96 to 8, p ≈ 3e-20 |
| Laya, broad question, right answers | 178 of 300, p ≈ 0.0015 against chance |
| Calibration error (ECE), broad question: Laya / Jev | 0.026 / 0.163 |
| Calibration error (ECE), per field: Laya / Jev | 0.375 / 0.037 |
| Jev asked a choice instead of true-or-false, broad question | 0.833 accuracy, AUC 0.944 |
| Jev asked a choice instead of true-or-false, per field | 0.958 accuracy, 8 of 150 faithful lines flagged |
| Jev, the broad question run twice | 6 of 300 decisions flipped, all between 0.45 and 0.55; mean probability shift 0.018, largest 0.12 |
| Comparator accuracy, broad / per field | 0.987 / 0.997 |
| Always answering yes, per field | 0.875 |
| Dropping the four mislabelled lines | every accuracy and AUC moves by 0.016 or less; Laya 0.593 → 0.595, Jev 0.887 → 0.895 |
B. Per field
Swipe to compare columns →
| Question | Laya false alarms | Laya caught | Jev false alarms | Jev caught |
|---|---|---|---|---|
| Is the amount right? | 95/150 (98/154 real) | 72/74 (69/70 real) | 0/150 (0/154 real) | 68/74 (68/70 real) |
| Is the date right? | 25/150 | 7/37 | 0/150 | 22/37 |
| Is the description right? | 42/150 | 13/22 | 0/150 | 22/22 |
| Is it in the right column (debit or credit)? | 32/150 | 2/17 | 0/150 | 15/17 |
Bracketed counts treat the four mislabelled lines as faithful. Laya’s amount question caught 69 of 70 real amount errors. It also called 98 of 154 correct amounts wrong. Its date question missed every day and month swap, and its debit-or-credit question caught 2 of 17 flips. Laya’s base card warns that a true-or-false question can follow its option labels and return a confident “no”; that may be the cause here (untested).
Jev did much better on impossible dates than on plausible ones. On the 14 swaps that produced a month above 12, its broad question caught all 14 and its per-field questions 9. On the 10 plausible swaps, its broad question caught 2 and its per-field questions none. Its broad-question false alarms cluster on amounts with a thousands comma (“1,234.56” on the line, written without the comma in the record), which in my set are all deposits: 14 of the 24 comma deposits were flagged, against 5 of 123 withdrawals.
Fine-tuning Laya is the obvious next step. The official notebook targets Kaggle’s free pair of T4 GPUs and takes about four to five hours. laya-mlx only runs inference, so on a Mac the training would have to happen in PyTorch on the GPU and be converted afterwards.
C. Speed and footprint
Swipe to compare columns →
| Checkpoint | Load | Peak memory | Per decision, p50 / p95 |
|---|---|---|---|
| typed-decisions, 421M | 0.56 s | 1.06 GiB | 15.5 / 15.9 ms |
| multilingual, 322M | 0.74 s | 0.77 GiB | 8.3 / 8.6 ms |
That is up to about 2 ms slower than the laya-mlx README’s M3 Max figures (13.4 and 7.4 ms at the median), on my longer inputs, so not a like-for-like comparison.
Swipe to compare columns →
| Setting | Median per call | Per question | Tail |
|---|---|---|---|
| Batch size 16 (the port's default) | 239 ms | about 12 ms | 7.8 s at p95 |
| Batch size 20 | 308 ms | 15.4 ms | 1.7 s |
The README’s own 50-question benchmark uses a batch size of 64 (146.8 questions a second on an M3 Max), which I did not try. My guess at the long tail is that Metal compiles a new kernel for every new padded length (unverified).
Jev’s 95th percentile from Singapore was 343 ms. After the ~46 ms trip to the edge, about 225 ms remains: TypeSafe’s servers plus whatever distance lies behind that edge. Four questions per call took the same time as one, with about 22 percent more input.
On a 64GB Mac, memory matters more than speed. My daily assistant, a 125B model on DwarfStar, was already about 12 GB into swap at a 160K context when I ran this. Whether Laya’s gigabyte pushed it further, I can’t say: I did not sample swap before the run.
D. Method
- Data: 300 synthetic lines from
make_synth.py, seed 20260921, generated alternating faithful and corrupted, then shuffled. Per-field labels come from the injected error. No real data was used. - Questions: each asked as a true-or-false question (a noul), counted as true at 0.5 or above. Both models saw identical question objects and state text. The longest Laya input was 175 tokens for the broad question (212 with the multilingual tokenizer) and 639 per field, of a 1,024-token budget.
- Laya: Mac Studio, M4 Max, 64 GB, on 21 September; laya-mlx 0.1.0, MLX 0.32.2, fp16. Typed-decisions has 421M parameters and the multilingual checkpoint 322M. The daily DwarfStar server was running and idle. The broad runs were done on the CPU and repeated on the GPU with the same accuracy; the per-field run is CPU-only. Probabilities and calibration are the CPU runs’; for typed-decisions the GPU’s fp16 numbers differ in the second decimal (ECE 0.031, mean probability 0.544 on faithful lines).
- Jev: tested through its API on 1 October, in one seven-minute window from Singapore (Cloudflare edge SIN), pinned to
jev-1.13.0rather than thejev-latestalias. Raw HTTP, one request per line, sequential, with five unscored warm-up calls per pass. The API has no temperature or seed parameter. All 1,500 scored calls reported the same model, with no retries or errors. The five passes used about 697,000 input tokens. According to Wikipedia, TypeSafe reports Jev is 40 to 200 times faster than frontier LLMs on its own workflows; I did not test that. - Credit: the test harness, the runs and the analysis were done with Claude Code ; I directed and reviewed.
E. A flaw in my own test set
Checking Jev’s misses turned up two flaws in my generator. An audit script found both; the comparator independently confirms the first.
- Four digit swaps change nothing. They swapped two identical digits, so four “corrupted” lines are actually faithful. Jev’s amount question called all four amounts correct, which is right; its broad question called three of the four accurate. Laya’s amount question flagged three of them.
- 14 of the 24 day and month swaps are easier than intended. They produce a month above 12, an impossible date.
Dropping the four lines changes no conclusion. I left the generator as it was, so the figures here stay reproducible from the same seed. The generator, harness and raw results are not published yet.
F. How this fits other public tests
These are broader than mine, and none of them checks a record against its source like this.
Swipe to compare columns →
| Test | Scope | Headline | Caveat |
|---|---|---|---|
| Harry Munro, Jev against Laya | eight typed-decision tasks | Jev 92.9%, Laya 65.3%, Laya typed-decisions 71.1%; Jev 136 ms median, about 50 ms of it network | no task targets Jev's documented weak spots: counting, arithmetic, dates, multi-hop |
| priorbench/jev | 400 items of its own, through OpenRouter from Western Europe | 95.9% zero-shot | a fixed overhead of about 430 ms per call it cannot separate from the gateway |
| The typed-decisions benchmark and Convai's cards | an independent benchmark (LocalLLaMA), not affiliated with TypeSafe or Convai | base Laya 0.362, fine-tuned Laya 0.766, Jev 0.727 measured by the benchmark's maintainers through TypeSafe's API | the 0.766 checkpoint was trained on the benchmark's training split, and above its 0.735 teacher self-agreement the benchmark says a model learns the teacher's quirks |
| fly2abhishek/jev-field-tests | three invoice calls | an invoice total $100 too high was judged to match its line items with probability 0.90 | that task needs addition; mine only needs comparison |
Links: Munro , priorbench , the typed-decisions benchmark , fly2abhishek . My Jev numbers come from one connection, one seven-minute window and one model version, on synthetic lines from one generator. 300 lines give about ±0.02 of standard error on Jev’s accuracy.
G. Measured or reported
| Claim | Basis |
|---|---|
| Laya accuracy, probabilities, calibration, per-field counts | Measured on my Mac, 21 September |
| Jev accuracy, AUC, calibration, catches, latency, tokens, repeatability | Measured through TypeSafe's API, jev-1.13.0, 1 October, from Singapore |
| The ten-line comparator's scores | Measured on the same 300 lines |
| Significance and AUC | Computed from the per-line results: Laya 178 of 300, p ≈ 0.0015; Jev against Laya, 96 lines to 8, p ≈ 3e-20 |
| Laya load time, memory, latency, batching | Measured on my Mac, 21 September |
| Jev's server time against network time | Derived: wall clock minus TCP round trip. TypeSafe reports no server timing |
| Laya's 0.362 base and 0.766 fine-tuned scores, Jev's 0.727 | Reported: Convai's model cards and the independent typed-decisions benchmark |
| laya-mlx's M3 Max latencies | Reported: the laya-mlx README |
| Jev's speed against frontier LLMs, its release date | Reported: TypeSafe via Wikipedia. Not tested, beyond the latency measured here |
| Other public Jev and Laya results | Reported: each repository named in the text |
| Four mislabelled lines, 14 easy date swaps | Measured: an audit script found both; the comparator confirms the four lines |
| Why batching has a slow tail | Guess: per-shape Metal compilation. Not verified |
| Why Jev false-flags comma amounts | Hypothesis: the broad question's wording. Confounded with the deposit column; not tested |
| Whether the sidecar pushed the Mac further into swap | Unknown: swap was not sampled before the run |
| Why Laya's amount question leans to 'no' | Hypothesis: the noul option-label bias Laya's base card documents. Not tested |
Sources
- TypeSafe docs: models
- TypeSafe docs: jev-1.13 known weak spots
- TypeSafe docs: how to build with System One
- Jev (AI model), Wikipedia
- TechCrunch on Jev, 18 September 2026
- The typed-decisions benchmark (LocalLLaMA), independent of TypeSafe and Convai
- Laya typed-decisions model card, Convai Innovations
- Laya base model card: "a fast base to specialise"
- Laya source and fine-tuning notebook
- laya-mlx, the independent MLX port
- Harry Munro's Jev and Laya benchmark (commit 0f977d2)
- priorbench/jev (commit 92e0a55)
- fly2abhishek/jev-field-tests, followup/exact-results.json
- Part three: A 2-bit 125B model became my daily assistant on a 64GB M4 Max
Laya measured on a Mac Studio M4 Max with 64 GB on 21 September 2026, using laya-mlx 0.1.0 and MLX 0.32.2. Jev tested through its API, jev-1.13.0, on 1 October 2026.