Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Jev Is a New AI Model That Answers in Probabilities

Jev Is a New AI Model That Answers in Probabilities

TypeSafe AI's Jev skips text generation entirely and returns calibrated probability scores instead, priced at $0.042 per million input tokens.

Dr. Nova Chen
Dr. Nova Chen★Sep 19, 2026★6 min read

An AI Model That Refuses to Write Anything

Almost every AI model you have used in the last four years has the same job description: predict the next token, then the next one, until a sentence appears. TypeSafe AI has released a model that declines that job entirely. Jev, which the startup put on general API access this week, never produces a word. It reads text and returns a calibrated probability for each outcome you defined in advance. The company calls it a System One model, after the fast intuitive half of the thinking-fast-and-slow split, and the framing is more than marketing — a model with no free-form output has no room to invent a fact.

  • Who built it: TypeSafe AI, co-founded by Diogo Almeida, a former OpenAI researcher who worked on ChatGPT and on reinforcement learning from human feedback
  • What it outputs: calibrated probability scores over a predefined answer space, not generated text
  • Pricing: $42 per billion input tokens, which works out to $0.042 per million, with output tokens free
  • Limits: a 64k token budget covering the state document plus the longest question, text input only, closed API with no open weights

What Does a Calibrated Decision Model Actually Do?

The mechanics are simpler than the pitch. You hand Jev a state document — an email, a support ticket, a diff, a transcript — and a set of questions with fixed possible answers. The model evaluates the questions against that state in parallel and returns a confidence figure for each.

The word doing the work is *calibrated*. TypeSafe trains the model with a technique it calls reinforcement learning from calibrated decisions, scoring the model against verifiable ground truth using proper scoring rules such as the Brier score. The target is that when Jev reports 85 percent confidence, it is right about 85 percent of the time. Anyone who has tried to threshold an LLM's self-reported confidence knows how rarely that holds. A number you can actually threshold on is the product.

The practical upshot is that an entire class of jobs that teams currently hand to a general-purpose LLM — classification, routing, safety screening, triage, deciding whether another AI agent's output looks reasonable — stops being a text-generation problem. TypeSafe's documentation lists throughput of 250,000 tokens per second or 1,200 requests per minute, whichever ceiling arrives first.

Why Developers Are Paying Attention

The pricing is the part that made the rounds. Output tokens are free because there are no output tokens, and input is metered per billion rather than per million, which is an unusual enough unit that it reframes the cost conversation. TypeSafe told TechCrunch the model runs roughly 5 to 18 times faster than a frontier general model on classification work and 10 to 20 times cheaper than comparable alternatives. Demand was strong enough on launch day to cause brief API hiccups.

A caution worth stating plainly: those speed and cost comparisons are the company's own figures, published alongside its own evaluations. Independent benchmarks have not landed yet, and TypeSafe's public documentation does not include accuracy numbers. Treat the headline multiples as vendor claims until someone outside the company reproduces them. The architectural claim — that a model with a closed output space cannot hallucinate a fact it was never allowed to emit — is structural rather than empirical, and that one stands on its own.

Where a System One Model Fits

Jev is not a replacement for a reasoning model, and TypeSafe is not pitching it as one. It is a component. The obvious pattern is a cheap, fast, calibrated gate sitting in front of an expensive model: Jev decides whether a request needs the big model, whether a generated answer looks safe to ship, or which of six pipelines an incoming document belongs to. That is the same architectural instinct behind the keyless credential work we covered in Postman Passport's secretless API access — push the boring, high-volume decision down to a purpose-built layer.

It also sits neatly beside the agent tooling that has dominated the last few weeks, from the OpenAI Agents API to Paper2Agent's research agents. Agents generate a lot of small decisions, and small decisions are exactly what a System One model is for. More on model releases and developer tooling in our AI coverage.

The startup raised a $40 million seed round to build this, and the first model is deliberately narrow. That narrowness is the interesting bet: not a model that can do everything slightly better, but one that does a single unglamorous thing with a number you can trust.

Sources: TechCrunch — September 18, 2026; TypeSafe AI model documentation — accessed September 19, 2026.

More AI Stories