
Jev Is a New AI Model That Answers in Probabilities
TypeSafe AI's Jev skips text generation entirely and returns calibrated probability scores instead, priced at $0.042 per million input tokens.
An AI Model That Refuses to Write Anything
Almost every AI model you have used in the last four years has the same job description: predict the next token, then the next one, until a sentence appears. TypeSafe AI has released a model that declines that job entirely. Jev, which the startup put on general API access this week, never produces a word. It reads text and returns a calibrated probability for each outcome you defined in advance. The company calls it a System One model, after the fast intuitive half of the thinking-fast-and-slow split, and the framing is more than marketing — a model with no free-form output has no room to invent a fact.
- Who built it: TypeSafe AI, co-founded by Diogo Almeida, a former OpenAI researcher who worked on ChatGPT and on reinforcement learning from human feedback
- What it outputs: calibrated probability scores over a predefined answer space, not generated text
- Pricing: $42 per billion input tokens, which works out to $0.042 per million, with output tokens free
- Limits: a 64k token budget covering the state document plus the longest question, text input only, closed API with no open weights
What Does a Calibrated Decision Model Actually Do?
The mechanics are simpler than the pitch. You hand Jev a state document — an email, a support ticket, a diff, a transcript — and a set of questions with fixed possible answers. The model evaluates the questions against that state in parallel and returns a confidence figure for each.
The word doing the work is *calibrated*. TypeSafe trains the model with a technique it calls reinforcement learning from calibrated decisions, scoring the model against verifiable ground truth using proper scoring rules such as the Brier score. The target is that when Jev reports 85 percent confidence, it is right about 85 percent of the time. Anyone who has tried to threshold an LLM's self-reported confidence knows how rarely that holds. A number you can actually threshold on is the product.
The practical upshot is that an entire class of jobs that teams currently hand to a general-purpose LLM — classification, routing, safety screening, triage, deciding whether another AI agent's output looks reasonable — stops being a text-generation problem. TypeSafe's documentation lists throughput of 250,000 tokens per second or 1,200 requests per minute, whichever ceiling arrives first.
Why Developers Are Paying Attention
The pricing is the part that made the rounds. Output tokens are free because there are no output tokens, and input is metered per billion rather than per million, which is an unusual enough unit that it reframes the cost conversation. TypeSafe told TechCrunch the model runs roughly 5 to 18 times faster than a frontier general model on classification work and 10 to 20 times cheaper than comparable alternatives. Demand was strong enough on launch day to cause brief API hiccups.
A caution worth stating plainly: those speed and cost comparisons are the company's own figures, published alongside its own evaluations. Independent benchmarks have not landed yet, and TypeSafe's public documentation does not include accuracy numbers. Treat the headline multiples as vendor claims until someone outside the company reproduces them. The architectural claim — that a model with a closed output space cannot hallucinate a fact it was never allowed to emit — is structural rather than empirical, and that one stands on its own.
Where a System One Model Fits
Jev is not a replacement for a reasoning model, and TypeSafe is not pitching it as one. It is a component. The obvious pattern is a cheap, fast, calibrated gate sitting in front of an expensive model: Jev decides whether a request needs the big model, whether a generated answer looks safe to ship, or which of six pipelines an incoming document belongs to. That is the same architectural instinct behind the keyless credential work we covered in Postman Passport's secretless API access — push the boring, high-volume decision down to a purpose-built layer.
It also sits neatly beside the agent tooling that has dominated the last few weeks, from the OpenAI Agents API to Paper2Agent's research agents. Agents generate a lot of small decisions, and small decisions are exactly what a System One model is for. More on model releases and developer tooling in our AI coverage.
The startup raised a $40 million seed round to build this, and the first model is deliberately narrow. That narrowness is the interesting bet: not a model that can do everything slightly better, but one that does a single unglamorous thing with a number you can trust.
Sources: TechCrunch — September 18, 2026; TypeSafe AI model documentation — accessed September 19, 2026.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
