TypeSafe’s Jev is getting attention at the moment and most of the discussion I have seen is about what category it belongs to. It is not a chat model. It is not a classifier you train. It sits somewhere between the two and the positioning confuses people, which is a shame, because the thing it does is genuinely useful and quite narrow.
One thing to get out of the way, since I run an AI orchestration platform and this could be read as a competitive piece. It is not. Jev sits in the decision layer, in front of generation. What I build sits underneath generation, choosing and managing which model serves a request. A system can and arguably should use both. I tested Jev against a small model doing the same classification work because that is the honest alternative an engineer would weigh, not because the two products compete for the same slot.
I ran it over an evening. This post is the review.
Key Takeaways
- Jev is a typed judgment API. You send application state and narrow typed questions, and get back typed answers with probabilities, never prose.
- It scored 100% on baseline and contested judgments, 94% on gateway-shaped policy gates, against 88% and 77% for a small model on the same cases.
- It costs roughly twice as much per call as a small model, and the gap is entirely payload shape, application state costs tokens, not a difference in per-token pricing.
- Its real strength is fan-out: one call answering eight independent questions about the same state in about 751 milliseconds, against roughly 22 seconds for eight sequential small-model calls.
- It cannot generate text at all, and it costs more than a classical ML classifier when the taxonomy is frozen and volume is high, so it is not a universal replacement for either.
What It Is
You send application state and a set of narrow typed questions. A Noul is a boolean. A Choice picks from a set you define. A Score returns a number in a range. Each answer comes back with a probability attached, and your code composes them.
It never writes a reply. There is no prose mode, no fallback to a paragraph, no possibility of it deciding to explain itself instead of answering. TypeSafe describe it as a System One primitive and that framing is accurate: fast, narrow, non-deliberative judgment that software consumes.
The version I tested was jev-1.13.0.
How It Scored
Against a small generative model on shared gold labels, both asked to make the same decisions:
| Suite | n | Jev | Small model | Jev p50 | SLM p50 |
| Baseline judgments | 8 | 100% | 88% | 647 ms | 703 ms |
| Contested judgments | 16 | 100% | 88% | 671 ms | 686 ms |
| Stress / edge cases | 12 | 100% | 75% | 656 ms | 600 ms |
| Gateway gates (routing, policy, DLP) | 17 | 94% | 77% | 651 ms | 587 ms |
And on the five-way department bakeoff it scored 12 out of 12 at 678 milliseconds average, matching Opus on accuracy at about 2% of the cost, and matching a trained classifier on accuracy while needing no training data at all.
On cost, TypeSafe’s usage export shows 242 requests and 100,911 input tokens for the session. At their published $0.042 per million input tokens with output free, that is $0.0000175 per call. The small-model tier in my gateway logs came to $0.0000079 per call. So roughly twice the price of a small model, and about 2% of a frontier call.
Look at why before drawing a conclusion from that 2x. gemma-3-4b’s blended rate measured from my meter is $0.058 per million tokens; Jev’s published input rate is $0.042. Per token they are close. The gap is prompt size: 417 input tokens per Jev call against 128 for the small-model arm, because a typed judgment call carries application state and state costs tokens. Same class of inference, different payload shape.
Jev charges nothing for output, which is worth noting and worth not overstating. It averaged 40 output tokens per call, 9,687 across the session. Priced at its own input rate that would have added $0.000407, so free output saves about 10% and moves the gap from 2.43× to 2.22×. It does not offset a threefold difference in input volume.
The shape is more interesting than the size. The small model’s output rate is double its input rate, so output is about 30% of that arm’s bill despite the answers being short. Jev runs the opposite way: you pay to describe the state and the questions, and the answers cost nothing because they are typed and tiny. That is a pricing structure matched to what the product does, rather than a discount. It also means the economics improve as you ask more questions of the same state, which is exactly the fan-out case below.
Quick link: Five Ways to Make a Decision in an AI System
The Three Things It Is Genuinely Good At
1. Fan-Out, Which Is the Real Product
Take one support ticket and ask eight independent questions about it. Is it billing. Is it technical. Is it urgent. What is the tone. Which department. What severity. Does it contain PII. Is there an actionable request.
| Approach | Latency |
| One Jev call, eight questions | ~751 ms |
| Eight sequential small-model calls | ~22 s (about 29× slower) |
| One small-model mega-prompt, all eight in one JSON blob | ~1.3 s |
The mega-prompt is the pragmatic workaround and it is also where the risk concentrates. You have taken eight independent decisions and made them all dependent on one blob parsing cleanly. One markdown fence, which I did observe, and you lose all eight answers rather than one.
Real gates ask batteries of questions, not single questions. If your routing layer needs six or eight signals before it decides, the difference between one 750-millisecond call and eight sequential ones is the difference between an inline gate and a background job. This, more than accuracy, is the argument for the product.
2. A Contract Rather Than a Parse
Every chat-based classifier has a parse step. Most of the time it works. Some of the time you get prose, or JSON wrapped in fences, or a truncated object. You write a repair layer, you maintain it, and you discover its gaps in production.
A typed API removes the category. You asked for a boolean, you get a boolean. There is nothing to parse, so there is nothing to repair.
This matters most on policy gates, which is exactly where the accuracy gap was widest, 94% against 77%. I wrote seventeen gateway-shaped synthetic cases: intent routing, escalate versus auto-handle, EU region locks, retention exports, budget hard stops, AWS key detection, PAN detection, prompt injection, allow/review/block. The small model’s errors there included treating a compliant EU request as a violation and letting an over-budget request proceed. Those are not classification mistakes in the abstract, they are a compliance failure and a spend failure.
3. Criteria You Edit Rather Than Retrain
The classical ML classifier beat everything on speed and cost, and it needed 105 labelled tickets to do it. Change the taxonomy and you are back in a training cycle.
With a typed judgment API you edit the criterion text and deploy. No corpus, no retraining, no drift in the relationship between your labels and your model. For a decision surface that moves monthly, that is the whole value proposition, and it is why the comparison with classical ML is not really a comparison at all. They suit different rates of change.
I also checked stability, because a threshold is only useful if the distribution underneath it holds still. Five repeats of the same Choice returned the same answer every time, with confidence between 0.91 and 0.93.
Where It Does Not Fit
It cannot generate anything. Twelve breadth tasks, coding, summarisation, formatting, grammar, short copy, chat, reasoning, creative writing. The small-model arm passed all twelve at about 1.6 seconds average. Jev’s coverage is 0%, by design. If the job is to write something, this is the wrong tool and scoring it here would be a category error.
It costs more than local inference. 678 milliseconds and a real API call against 2.1 milliseconds and free. If your taxonomy is frozen and your volume is high, a logistic regression is the correct answer and Jev is an expensive way to reach the same label.
It has documented rough edges and I did not find them. TypeSafe publish known weaknesses for 1.13 around literal reading, counting, dates, noisy state and adversarial content. My stress suite scored it 12 out of 12, which I treat as a twelve-case synthetic result rather than a contradiction of the vendor’s own documentation. Do not read my number as “Jev never fails jagged tasks”. Keep arithmetic and calendar logic in code, send it the semantic question, and filter irrelevant state before you ask.
It was not perfect on the gates either. The 94% miss was an escalate-versus-auto-deny nuance on a high-impact contractor request. Exactly the kind of judgment call where you would want a human in the loop regardless of which system made the first pass.
Quick link: WrangleAI vs LangSmith
Guardrails: The Pattern It Was Built For
I rebuilt TypeSafe’s own guardrails cookbook pattern: a battery of hazard booleans plus a severity score, with the actual routing decision, pass, review, block or escalate, made by a plain function in my code.
Across fifteen reconstructed cases, Jev matched the published gold routes about 80% of the time at roughly 677 milliseconds. The small model landed about 73% at roughly 1.5 seconds. Disagreements clustered on severity edges, such as whether a medical dosage question should be reviewed or blocked, and on the small model over-firing on fiction and on general drug information.
That over-firing pattern is the interesting one. A generative model asked “is this harmful” is answering a vibes question. It has no stable sense of where your threshold sits and it will move that threshold between calls. Decomposing into calibrated signals and putting the threshold in your own code fixes the part that should never have been the model’s job in the first place.
Guardrails are policy over calibrated signals, not a model asked whether something is bad. Jev fits that shape natively. A small model can approximate it if you accept the parse risk and weaker calibration.
The Verdict
Jev is a narrow product that is very good at the narrow thing. It is a decision contract: typed in, typed out, probabilities your code can gate on, criteria you edit rather than retrain, and native fan-out across a battery of questions over shared state.
It is the strongest option I tested for policy gates and multi-judgment packs. It is the wrong option if your taxonomy is frozen and your volume is high, because classical ML will beat it on every axis that matters there. And it is not an option at all if you need text out.
The mistake I would guard against is treating it as a replacement for a chat model or for a classifier. It is neither. In the architecture from part 1 it sits in the gate layer, in front of the generative layer, deciding what gets through and where it goes.
| Jev | Small model | Classical ML | Frontier | |
| Accuracy on policy gates | 94% | 77% | n/a | not tested here |
| Latency, one judgment | ~660 ms | ~600 ms | 2 ms | ~2.3 s |
| Latency, eight judgments | ~750 ms | ~22 s sequential | ~17 ms | slower still |
| Cost per call | $0.0000175 | $0.0000079 | $0 | $0.000799 |
| Typed output | Yes | No, parse required | Yes | No, parse required |
| Change criteria | Edit text | Edit prompt | Retrain | Edit prompt |
| Can generate text | No | Yes | No | Yes |
Where WrangleAI Fits
Everything above is about the gate. Something still has to happen after it.
A typed judgment API tells you a request is safe, in scope, and belongs to the billing queue. It does not answer the request. The moment your gate says “allow and generate”, you are back to the question this post has deliberately set aside: which model, at what price, with what controls.
That is the layer I build. SmartRouter puts a pool of small models behind one API and picks per request, so you are not pinning a model and hoping it stays the right one. When I measured my own gateway traffic, the blended cost per million tokens ranged from $0.023 to $7.61 across the models available to it, and a five-fold spread existed inside the small-model tier alone. Selecting correctly across that spread on every request is not something a team should be doing by hand, and it is not a decision that stays correct.
SmartRouter also carries the gate-layer controls the guardrails section of this post argues for: semantic classification, hard budget stops, DLP redaction and policy rules, applied before a request reaches a model. If you like the typed-judgment pattern but do not want to build the policy plumbing around it, there is a large overlap with what we already do.
Meter is where every cost figure in this post came from: billed spend per call, per model, mapped to cost centres. You cannot make the trade-offs in this series without that data, and most teams do not have it. Jev in the gate, orchestrated models behind it. That is the architecture I would build, and we provide the second half. See the platform.

FAQs
What is a typed judgment API?
A typed judgment API, such as TypeSafe’s Jev, takes application state and narrow typed questions, a boolean, a choice, or a score, and returns typed answers with probabilities attached, rather than writing prose. Code gates on the answer directly, with no parsing step.
Is Jev a replacement for a chat model?
No. Jev cannot generate text at all, scoring 0% on breadth tasks such as coding, summarisation or creative writing by design. It sits in the decision layer in front of generation, not in place of it.
How much does Jev cost compared to a small model?
In this test, Jev cost about $0.0000175 per call against $0.0000079 for a small model, roughly twice as much. The gap came from payload size, since a typed judgment call carries application state, not from a higher per-token rate.
What is Jev best used for?
Jev’s strongest use case is fan-out: asking several independent questions about the same piece of state in a single call. It answered eight questions about one support ticket in about 751 milliseconds, against roughly 22 seconds for eight sequential small-model calls.
Does a typed judgment API replace classical machine learning classifiers?
Not when the taxonomy is stable and volume is high, since classical ML wins on both speed and cost there. A typed judgment API is the better fit when criteria change often, since you edit the criterion text instead of retraining a model.




