Skip to content

Local inference field note / October 2026

I tested Jev against a free local model. Ten lines of code beat both.

One narrow task, 300 synthetic statement lines. Jev, tested through its API, was accurate. Laya, an open model I ran locally, was fast but too flat to use. A short script beat both, because it knows my generator’s exact line layout.

Accuracy on the same 300 lines, one broad question

the short comparator, which reads my generator's fixed line layout
0.99
Jev, tested through its API
0.89
Laya, run locally on the Mac: about 17 times faster than Jev per single-question call from Singapore
0.59
On this page
  1. The short answer
  2. What I tested
  3. Results
  4. Why Laya was flat
  5. Where Jev falls short
  6. Ten lines of code
  7. What's next
  8. Details

The short answer

Jev, an AI model that returns decisions instead of chatting, was everywhere in September. It is TypeSafe AI’s hosted model. Laya is an open-weight model of the same kind, and an independent port runs it on Apple Silicon. I put both on one narrow job: 300 synthetic statement lines, each with a record extracted from it. Does the record match the line?

  • Jev was accurate. With one broad question it was right on 89 percent of lines. With four per-field questions it raised no false alarm on any of the 150 faithful lines.
  • Laya was fast, local and free, and too flat to use. It answered in 8 to 16 milliseconds on my M4 Max, in about a gigabyte of memory. But it barely told good lines from bad: it scored 0.59. Its own card calls it “a fast base to specialise, not a zero-shot decision engine”, and this checkpoint was trained for four other workflows.
  • Ten lines of code beat both. A script of about a dozen lines scored 0.99, because it knows my generator’s exact column layout. Neither model had that advantage. For exact numbers in a known format, plain code is the right tool, and TypeSafe’s own documentation says so too.

What I tested

A language model writes text. A decision model does not. You give it some text and a typed question, such as pick one, score this, or true or false. It returns the answer with a probability, in one pass.

  • Jev is TypeSafe AI’s hosted, pay-per-use model, released in limited early access on 15 September . I tested it through its API on 1 October, pinned to version jev-1.13.0.
  • Laya is an open-weight model from Convai Innovations, licensed Apache 2.0. In this note “Laya” means its typed-decisions checkpoint , unless I say otherwise. I ran it on my Mac on 21 September.
  • laya-mlx is an independent port of Laya to Apple’s MLX framework, first published on 19 September. Its README says plainly that it is not an official Convai release.

The task: given a statement line and a record extracted from it, does the record match the line? A document-extraction job would ask that thousands of times, which is exactly the kind of question a decision model promises to handle in bulk.

The 300 lines are synthetic and seeded. Half are faithful. The other half carry one injected error each, of eight kinds: swapped digits, a shifted decimal point, an amount off by one cent, the balance copied in as the amount, a wrong day, a day and month swapped, debit and credit flipped, or the wrong merchant.

Both models got the same lines, the same input text and the same scoring code. Each was asked two ways: one broad question (“is this extraction accurate?”), and four true-or-false questions, one per field.

Results: Jev reads the line, Laya barely does

Jev separated faithful lines from corrupted ones; Laya barely did. The clearest single measure is AUC: how often a random faithful line gets a higher probability than a random corrupted one. 0.5 is a coin flip and 1.0 is perfect. On the broad question, Jev scored 0.95 against Laya’s 0.66.

Swipe to compare columns →

Accuracy at the 0.5 threshold, AUC, and mean probability of 'accurate', 300 lines, zero-shot
RunAccuracyAUCMean P, faithfulMean P, one error
One broad question, typed-decisions 421M0.5930.6600.5350.480
One broad question, multilingual 322M0.5030.6460.9440.892
One broad question, Jev0.8870.9510.6290.110
Four per-field questions, typed-decisions 421M0.6190.6650.5130.488
Four per-field questions, Jev0.9650.9490.9540.737

Per-field accuracy is over 1,200 answers, and always answering “yes” would score 0.875 there. For the per-field runs the means and AUC use each line’s average over its four answers.

Mean probability that an extraction is accurate, for faithful and corrupted lines, in five runs. Laya's three runs put the two groups 0.025 to 0.055 apart; Jev's broad question puts them 0.52 apart and its per-field questions 0.22.150 faithful lines150 lines with one error0.000.250.500.751.00mean P(accurate)Broad question, typed-decisionsgap 0.055Broad question, multilingualgap 0.052Broad question, Jevgap 0.519Per-field questions, typed-decisionsgap 0.025Per-field questions, Jevgap 0.217
A useful checker puts the blue dots near 1 and the amber dots near 0. Only Jev’s runs pull them apart; Laya’s stay within about 0.05. Jev’s per-field gap looks smaller because each line averages four answers, and three of them are still correct on a corrupted line.

The gap is not noise. On the lines where the two disagreed on the broad question, Jev was right on 96 and Laya on 8. Jev’s per-field probabilities are also well calibrated, and asking the same thing as a choice question gave much the same result.

My own generator mislabelled four of the 300 lines. Dropping them moves no figure by more than 0.016; the details are below.

Why Laya was flat

Laya was unsure rather than wrong. Its average probability for faithful lines sat only about 0.05 above the one for corrupted lines. That gap is small but real: its 178 right answers out of 300 are well above chance. Asked field by field, it leaned heavily towards “no”, as the per-field details show.

The other checkpoint I tried, the multilingual one, said “accurate” to almost everything. At the default threshold that scores 0.50, though its ranking still carried some signal.

And Laya was asked something it was not trained for. Its base card calls it “a fast base to specialise, not a zero-shot decision engine”. The checkpoint I used is that base fine-tuned for four synthetic workflows: agent traces, customer service, invoices and security incidents. Invoice reconciliation is a near neighbour of my task, but the specialisation did not carry over. This is a specialist tested outside its domain.

Where Jev falls short

  • Speed. A Jev call took 271 ms at the median from Singapore, against Laya’s 8 to 16 ms on the Mac. About 46 ms of that is the trip to the nearest network edge. The API reports no server timing, so I can’t split the rest. The gap shrinks with more questions per call: four questions took Jev the same time as one, and Munro found the API faster than Laya above three or four.
  • Dates, like Laya. A day and month swapped into another valid date got past it: the per-field questions caught 0 of 10. TypeSafe’s own list of known weak spots notes that this version treats dates as text rather than as quantities to compare.
  • Literal reading. The broad question asks for an “exact” match, and Jev flagged 20 of the 150 faithful lines. Most of them had a thousands comma in the amount. But every comma amount in my set is a deposit, so I cannot tell the comma from the column (untested). TypeSafe’s weak-spots page lists literal reading too.
  • The same question twice is not quite the same answer. I ran the broad question twice, and 6 of 300 decisions flipped. All six sat right at the 0.5 threshold. The API has no temperature or seed setting to stop it.
  • It leaves the machine. Anything you would not send to a hosted API is ruled out, whatever the score. Laya runs on the Mac.

Ten lines of code beat both

Plain code was the best checker here. A script of about a dozen lines compares the record with the line field by field: the amount, the date, the column and the description. It scored 0.99 on the broad question. Its only “errors” were the four lines my generator mislabelled.

It had an advantage neither model had: it knows the fixed column layout my generator writes, so it can cut the fields out of the line by position. A job that holds the raw line in a known format can do the same. A statement in a layout it has never seen would break it, while the models read the line as text.

So when the source line is available in a known format, comparing amounts, dates and columns in code is exact and free. TypeSafe’s documentation agrees: its guide to building with Jev recommends keeping deterministic rules in code, and its weak-spots page recommends doing any mathematical logic in code. A decision model earns its place on the fuzzy checks a script cannot settle, such as a merchant name written three ways. These 300 lines barely test that.

What’s next

Laya’s score here is a zero-shot floor, not a verdict. Next I want to fine-tune it on my Mac for something fuzzier: spotting low-effort “slop” writing, with my local 125B model as the teacher, measured against my own hand labels.

Details

Everything below is for people who want to check the numbers. If you only wanted the verdict, you have it.

A. Full results

The checks behind “the gap is not noise”, the calibration errors, the choice-question variant and the comparator’s exact scores.

Supporting checks, same 300 lines
CheckResult
Jev against Laya, broad question, lines where they disagree96 to 8, p ≈ 3e-20
Laya, broad question, right answers178 of 300, p ≈ 0.0015 against chance
Calibration error (ECE), broad question: Laya / Jev0.026 / 0.163
Calibration error (ECE), per field: Laya / Jev0.375 / 0.037
Jev asked a choice instead of true-or-false, broad question0.833 accuracy, AUC 0.944
Jev asked a choice instead of true-or-false, per field0.958 accuracy, 8 of 150 faithful lines flagged
Jev, the broad question run twice6 of 300 decisions flipped, all between 0.45 and 0.55; mean probability shift 0.018, largest 0.12
Comparator accuracy, broad / per field0.987 / 0.997
Always answering yes, per field0.875
Dropping the four mislabelled linesevery accuracy and AUC moves by 0.016 or less; Laya 0.593 → 0.595, Jev 0.887 → 0.895

B. Per field

Swipe to compare columns →

Per-field questions: false alarms on the 150 faithful lines, and errors of that field caught
QuestionLaya false alarmsLaya caughtJev false alarmsJev caught
Is the amount right?95/150 (98/154 real)72/74 (69/70 real)0/150 (0/154 real)68/74 (68/70 real)
Is the date right?25/1507/370/15022/37
Is the description right?42/15013/220/15022/22
Is it in the right column (debit or credit)?32/1502/170/15015/17

Bracketed counts treat the four mislabelled lines as faithful. Laya’s amount question caught 69 of 70 real amount errors. It also called 98 of 154 correct amounts wrong. Its date question missed every day and month swap, and its debit-or-credit question caught 2 of 17 flips. Laya’s base card warns that a true-or-false question can follow its option labels and return a confident “no”; that may be the cause here (untested).

Jev did much better on impossible dates than on plausible ones. On the 14 swaps that produced a month above 12, its broad question caught all 14 and its per-field questions 9. On the 10 plausible swaps, its broad question caught 2 and its per-field questions none. Its broad-question false alarms cluster on amounts with a thousands comma (“1,234.56” on the line, written without the comma in the record), which in my set are all deposits: 14 of the 24 comma deposits were flagged, against 5 of 123 withdrawals.

Fine-tuning Laya is the obvious next step. The official notebook targets Kaggle’s free pair of T4 GPUs and takes about four to five hours. laya-mlx only runs inference, so on a Mac the training would have to happen in PyTorch on the GPU and be converted afterwards.

C. Speed and footprint

Swipe to compare columns →

Laya, one question per call on the M4 Max GPU, fp16
CheckpointLoadPeak memoryPer decision, p50 / p95
typed-decisions, 421M0.56 s1.06 GiB15.5 / 15.9 ms
multilingual, 322M0.74 s0.77 GiB8.3 / 8.6 ms

That is up to about 2 ms slower than the laya-mlx README’s M3 Max figures (13.4 and 7.4 ms at the median), on my longer inputs, so not a like-for-like comparison.

Swipe to compare columns →

Laya batching: 20 questions about one input in one call
SettingMedian per callPer questionTail
Batch size 16 (the port's default)239 msabout 12 ms7.8 s at p95
Batch size 20308 ms15.4 ms1.7 s

The README’s own 50-question benchmark uses a batch size of 64 (146.8 questions a second on an M3 Max), which I did not try. My guess at the long tail is that Metal compiles a new kernel for every new padded length (unverified).

Jev’s 95th percentile from Singapore was 343 ms. After the ~46 ms trip to the edge, about 225 ms remains: TypeSafe’s servers plus whatever distance lies behind that edge. Four questions per call took the same time as one, with about 22 percent more input.

On a 64GB Mac, memory matters more than speed. My daily assistant, a 125B model on DwarfStar, was already about 12 GB into swap at a 160K context when I ran this. Whether Laya’s gigabyte pushed it further, I can’t say: I did not sample swap before the run.

D. Method

  • Data: 300 synthetic lines from make_synth.py, seed 20260921, generated alternating faithful and corrupted, then shuffled. Per-field labels come from the injected error. No real data was used.
  • Questions: each asked as a true-or-false question (a noul), counted as true at 0.5 or above. Both models saw identical question objects and state text. The longest Laya input was 175 tokens for the broad question (212 with the multilingual tokenizer) and 639 per field, of a 1,024-token budget.
  • Laya: Mac Studio, M4 Max, 64 GB, on 21 September; laya-mlx 0.1.0, MLX 0.32.2, fp16. Typed-decisions has 421M parameters and the multilingual checkpoint 322M. The daily DwarfStar server was running and idle. The broad runs were done on the CPU and repeated on the GPU with the same accuracy; the per-field run is CPU-only. Probabilities and calibration are the CPU runs’; for typed-decisions the GPU’s fp16 numbers differ in the second decimal (ECE 0.031, mean probability 0.544 on faithful lines).
  • Jev: tested through its API on 1 October, in one seven-minute window from Singapore (Cloudflare edge SIN), pinned to jev-1.13.0 rather than the jev-latest alias. Raw HTTP, one request per line, sequential, with five unscored warm-up calls per pass. The API has no temperature or seed parameter. All 1,500 scored calls reported the same model, with no retries or errors. The five passes used about 697,000 input tokens. According to Wikipedia, TypeSafe reports Jev is 40 to 200 times faster than frontier LLMs on its own workflows; I did not test that.
  • Credit: the test harness, the runs and the analysis were done with Claude Code ; I directed and reviewed.

E. A flaw in my own test set

Checking Jev’s misses turned up two flaws in my generator. An audit script found both; the comparator independently confirms the first.

  • Four digit swaps change nothing. They swapped two identical digits, so four “corrupted” lines are actually faithful. Jev’s amount question called all four amounts correct, which is right; its broad question called three of the four accurate. Laya’s amount question flagged three of them.
  • 14 of the 24 day and month swaps are easier than intended. They produce a month above 12, an impossible date.

Dropping the four lines changes no conclusion. I left the generator as it was, so the figures here stay reproducible from the same seed. The generator, harness and raw results are not published yet.

F. How this fits other public tests

These are broader than mine, and none of them checks a record against its source like this.

Swipe to compare columns →

Other public Jev and Laya tests
TestScopeHeadlineCaveat
Harry Munro, Jev against Layaeight typed-decision tasksJev 92.9%, Laya 65.3%, Laya typed-decisions 71.1%; Jev 136 ms median, about 50 ms of it networkno task targets Jev's documented weak spots: counting, arithmetic, dates, multi-hop
priorbench/jev400 items of its own, through OpenRouter from Western Europe95.9% zero-shota fixed overhead of about 430 ms per call it cannot separate from the gateway
The typed-decisions benchmark and Convai's cardsan independent benchmark (LocalLLaMA), not affiliated with TypeSafe or Convaibase Laya 0.362, fine-tuned Laya 0.766, Jev 0.727 measured by the benchmark's maintainers through TypeSafe's APIthe 0.766 checkpoint was trained on the benchmark's training split, and above its 0.735 teacher self-agreement the benchmark says a model learns the teacher's quirks
fly2abhishek/jev-field-teststhree invoice callsan invoice total $100 too high was judged to match its line items with probability 0.90that task needs addition; mine only needs comparison

Links: Munro , priorbench , the typed-decisions benchmark , fly2abhishek . My Jev numbers come from one connection, one seven-minute window and one model version, on synthetic lines from one generator. 300 lines give about ±0.02 of standard error on Jev’s accuracy.

G. Measured or reported

How each claim in this note is known
ClaimBasis
Laya accuracy, probabilities, calibration, per-field countsMeasured on my Mac, 21 September
Jev accuracy, AUC, calibration, catches, latency, tokens, repeatabilityMeasured through TypeSafe's API, jev-1.13.0, 1 October, from Singapore
The ten-line comparator's scoresMeasured on the same 300 lines
Significance and AUCComputed from the per-line results: Laya 178 of 300, p ≈ 0.0015; Jev against Laya, 96 lines to 8, p ≈ 3e-20
Laya load time, memory, latency, batchingMeasured on my Mac, 21 September
Jev's server time against network timeDerived: wall clock minus TCP round trip. TypeSafe reports no server timing
Laya's 0.362 base and 0.766 fine-tuned scores, Jev's 0.727Reported: Convai's model cards and the independent typed-decisions benchmark
laya-mlx's M3 Max latenciesReported: the laya-mlx README
Jev's speed against frontier LLMs, its release dateReported: TypeSafe via Wikipedia. Not tested, beyond the latency measured here
Other public Jev and Laya resultsReported: each repository named in the text
Four mislabelled lines, 14 easy date swapsMeasured: an audit script found both; the comparator confirms the four lines
Why batching has a slow tailGuess: per-shape Metal compilation. Not verified
Why Jev false-flags comma amountsHypothesis: the broad question's wording. Confounded with the deposit column; not tested
Whether the sidecar pushed the Mac further into swapUnknown: swap was not sampled before the run
Why Laya's amount question leans to 'no'Hypothesis: the noul option-label bias Laya's base card documents. Not tested

Sources

Laya measured on a Mac Studio M4 Max with 64 GB on 21 September 2026, using laya-mlx 0.1.0 and MLX 0.32.2. Jev tested through its API, jev-1.13.0, on 1 October 2026.