
TypeSafe AI's non-LLM model Jev is drawing outsized developer demand for cheap, hallucination-free classification.
The AMW Read
A new non-LLM 'calibrated decision' model from a former OpenAI RLHF co-inventor shows measurable production traction (a documented OpenAI-to-Jev swap-out at Vercel), meaningfully updating the model-landscape baseline without resolving an existing open debate.
TypeSafe AI's non-LLM model Jev is drawing outsized developer demand for cheap, hallucination-free classification.
TypeSafe AI, founded two years ago by former OpenAI researcher Diogo Almeida — who helped build ChatGPT and co-invented reinforcement learning from human feedback — this week released Jev, a transformer-based model that does not generate text. Instead it outputs calibrated probabilities against outcomes the developer defines in advance, which the company says makes it fast, cheap and structurally unable to hallucinate. Output tokens are free; input tokens are billed by the billion rather than the million. Demand was high enough that TypeSafe briefly lost the ability to serve its API.
Jev is being adopted as both a substitute and a supplement for large language models in narrow, high-volume tasks. Vercel engineer Pranit Sharma said swapping Jev in for OpenAI's Luna 5.6 as a command-safety classifier produced results five to 18 times faster with better accuracy; Bryo AI CTO Nikhil Mudholkar found Gemini slightly more accurate on email classification but 10 to 20 times more expensive, with Jev's calibrated confidence scores proving more useful for automating decisions. Earendil CTO Armin Ronacher described a second use case: running Jev alongside LLM agents to catch jailbreaks and monitor traces, and eventually to route workloads between models using Jev's real-time probability estimates. Almeida, who trains Jev exclusively on synthetic data through what he calls reinforcement learning from calibrated decisions, argues the industry has over-indexed on language output when most software automation needs a cheap, well-calibrated decision instead.
For builders, Jev's economics point at a fast-growing niche: classifiers, guardrails, and model-routing layers sitting underneath or beside LLM agent stacks, priced on volume rather than per-generated-token. For investors, the caveat is architectural — outside observers suspect Jev runs on top of an open-weight LLM rather than a wholly new base, and Ronacher himself expects competitors to replicate the approach quickly now that the demand signal is public, meaning any durable edge is more likely to sit in synthetic-data quality and calibration technique than in the model itself.


