AI · Analysis

The Model That Refuses to Talk

Why Jev, a "System 1" decision model, might be the missing link between LLMs and software

By Lyle Heartman · October 3, 2026

Jev, a new model from TypeSafe AI, doesn't write a single word. You give it a situation and a fixed list of possible answers, and it gives back a probability for each one. That sounds like a step backward. I think it's the most practical idea in AI this year, because the hard part of putting LLMs inside software was never intelligence. It was getting a clean decision out the other end that code could trust. The best evidence that the idea is right: within weeks, several other teams had built their own versions.

The awkward last step

Every developer who has wired a language model into real software knows the moment. The model has done the hard part: it read the ticket, weighed the context, and reached a conclusion. Then it has to hand that conclusion to code, and code does not want prose. It wants a decision.

So we write prompts that beg. Return only APPROVE or REJECT. Nothing else. Please. We add retries, regex guards, and JSON validators. We are asking a probabilistic system that loves to write to behave like a deterministic function, and we are surprised when it occasionally writes a paragraph instead.

So you get two camps that don't fit together well. Regular software is reliable and stiff. LLMs are flexible and clever, and they'll return whatever string they feel like. Jev was built to sit between the two.

Meet Jev

Jev comes from TypeSafe AI, a lab that spent two years in stealth before launching in mid-September 2026. It was co-founded by CEO Diogo Almeida, who worked at OpenAI on the instruction-following research that became InstructGPT and ChatGPT, alongside Sasha Sheng and Erik Gafni.

TypeSafe calls Jev the first "System One" model. The term borrows from Daniel Kahneman: System 1 is fast, instinctive judgment, while System 2 is slow, deliberate reasoning. Chat and reasoning models chase System 2. Jev aims squarely at System 1, behaving less like a writer and more like a learned decision function. The name itself nods to the Jevons paradox, the economic idea that making something cheaper can drive up how much of it gets used.

The workflow flips the usual prompt. Instead of asking a question and hoping for a well-formatted reply, you hand Jev some unstructured state, describe the decision in natural language, and declare the valid outputs up front. Jev never writes the answer. It scores the options you already gave it.

Three ways to decide

Jev exposes exactly three output types, and each maps to something software already knows how to consume.

Primitive What you define What comes back Example
Choice Up to 255 unordered options The top option plus a probability for every option Route a "charged twice" ticket to billing, fraud, or tech support
Score An ordered scale (poor → excellent) A distribution over levels, collapsed into a continuous score Rate a support reply; mass split between good and excellent lands in between
Null A yes/no question Probabilities for true and false "Is this transaction fraudulent?" → 0.87

The customer-support case shows why this matters. A normal LLM has to produce the literal word billing, spelled and cased exactly as your code expects. Jev never generates that word at all. It returns a distribution over the three departments you defined, so it can't send back a typo or an answer you never offered. That is also why TypeSafe can say Jev cannot hallucinate in the narrow sense: a model that never emits free text cannot invent a tool name.

Why skipping the words is so fast

The distinction sounds cosmetic until you look at the compute. An autoregressive LLM produces output one token at a time: predict a token, append it, run the model again, repeat until done. Even a one-word verdict pays for that loop, plus the parsing afterward.

Jev throws the loop away. It reads the context once and returns a probability distribution over the allowed answers. TypeSafe reports end-to-end latency of roughly 70 to 500 milliseconds, with a 32K context window. On its own workflow benchmarks, it claims Jev is over a hundred times faster and around 440 times cheaper than frontier generative models on decision tasks. Those numbers are vendor-tested, and they only apply to decisions, because Jev cannot write.

Then there's the price. Input costs $42 per billion tokens, about 4.2 cents per million. Output is free, since a handful of probabilities is too cheap to bother metering. For a task that is just "read this and pick one," that is a fraction of what a full chat model costs.

Confidence you can build on

The part I find more interesting than the speed is the confidence numbers. TypeSafe trains Jev with a method it calls RLCD, reinforcement learning for calibrated decisions. The aim is simple: if Jev says it's 90% sure, it should be right about 90% of the time.

Most classifier scores do not behave that way. They rank options, but the number itself is loosely tied to reality. If you can trust the number, you can write code around it.

Here's a simple example.

Say Jev is 98% sure a refund request is legitimate. The system just processes it. At 70%, it asks a person to sign off. Below 50%, it passes the whole thing back to a bigger model to think through properly.

Jev ──▶ [ p ] ──▶ ◆ route
                    │
  p ≥ 0.95          ├──▶ EXECUTE
                    │    extremely confident
                    │
  0.70 – 0.95       ├──▶ ASK A HUMAN
                    │    uncertain
                    │
  p < 0.70          └──▶ ESCALATE
                         to a larger reasoning model

Jev's calibrated probability decides what happens next: act, ask a person, or hand off to a bigger model.

That's how I'd actually use it: right after a chat model in an agent. The chat model does the thinking and produces something messy. Jev turns that into a clear state the rest of the program can act on, so you stop relying on one model to both reason well and format its own output perfectly.

What you'd use it for

Jev works best for small, clear decisions that your code acts on immediately. That covers more ground than it sounds.

  • Routing and triage. Send a ticket to the right team, and check its urgency and tone in the same call.
  • Moderation and safety. Put a quick check in front of a bigger model so the obvious cases never get to it.
  • Fraud and risk. Ask a yes/no question and get back a probability rather than a guess written out as text.
  • Quality scores. Rate a support reply or a draft on a scale you define.
  • Intent and extraction gates. Work out what a user wants, or whether a document deserves the expensive processing step.
  • Running agents. Turn a chat model's rambling into the next concrete action, and use the confidence to decide whether to act, ask someone, or escalate.
  • Anything real-time. People have already used it to play Doom and Minecraft and to place trades, where waiting a second for a model to finish typing isn't an option.

It won't help with anything that needs writing, like summaries, emails or explanations. If you can't list the possible answers ahead of time, Jev isn't the right tool.

What is probably under the hood

TypeSafe has not published a technical report, so the architecture has to be inferred from Jev's behavior and its closest relatives. The best clues come from Fastino AI, which open-sourced GLiNER2 in July 2025 and GLiGuard in May 2026.

GLiNER2 makes the task schema part of the input. Rather than training one fixed classification head, you pass in task definitions and labels, even several tasks at once, and score them all in a single forward pass. GLiGuard applies the same logic to safety, arguing that guardrails are a classification problem that has been needlessly forced into text generation.

Both are built on a BERT-style bidirectional encoder. Unlike an autoregressive LLM, which reads left to right, an encoder lets every token attend to the whole input at once. The recipe works like this:

  1. Feed in the text, the task description, and every candidate label together, each label marked with special tokens.
  2. Let the transformer build a representation for each label that already knows the text, the task, and the competing labels.
  3. Pass each label's representation through a small shared network to get a single score.
  4. Apply softmax when only one answer is allowed, or sigmoid when several can be true.

No response is generated, so output is nearly free. GLiGuard reports up to 16 times higher throughput and 17 times lower latency than the much larger generative guard models it was compared with. But they only really understand the task families they were trained on. Jev may simply be this idea taken much further, with a far broader training distribution aimed at arbitrary decisions and calibration, so it can grasp a brand-new decision from plain language and accept a fresh set of answers at runtime.

The copies showed up fast

If Jev is this cheap, the model is probably small. Small models are easy to copy, and people did.

Two days after Jev launched, Conv AI Innovations released Laia, an open-source model with the same kind of interface. Its design is fully public. It's a 395-million-parameter ModernBERT encoder with a small decision head on top, about 421 million parameters in total. On TypeSafe's own benchmark of 2,000 decisions, the base Laia model scored around 36%. A version fine-tuned on that benchmark hit 76.6%. Jev scored about 72% without any fine-tuning. Laia runs in about 16 milliseconds locally, compared with roughly 117 for Jev through its API, but it trails Jev on other tests and its probabilities are noticeably worse when there are lots of options.

Another project, Vaughn, uses the same 395-million-parameter encoder and even copies Jev's API format. Its focus is the calibration piece. It trains with a Brier score loss on top of the usual one, which punishes the model when its confidence is off, then adjusts the probabilities again on held-out data. It runs in about 18 milliseconds and is the strongest of the open versions so far, though it still loses to Jev on almost everything. The one exception reported was playing Doom.

The bigger players moved too. A little over a week after Jev, Fastino released GLiNER 2.5 Decide, which reportedly beats the other open decision models by a wide margin. On September 29, OpenAI teased a decision API at its Dev Day, and Liquid AI released D1, a decision model that works across languages, something Jev doesn't do yet.

None of these beat Jev outright. What matters is that something very close to it can be built from an ordinary encoder in a few weeks and run on your own laptop.

You can fake it with a normal LLM

Some people skipped the encoder entirely and made regular chat models act like Jev. The trick is simpler than you'd expect.

A project called SystemOne takes an ordinary Qwen model and sets up the prompt so the very next thing it would write is an option number, like 0, 1, 2 or 3. Then it doesn't let the model write anything. It just reads how likely the model thinks each number is, normalizes those, and stops. If you have several questions about the same situation, you can process the shared context once and run all the questions in parallel. In a Doom demo, the author measured about 172 milliseconds between actions this way, against about 600 milliseconds using normal tool calls with the same model and prompt.

OpenJev goes further. Instead of scoring a single number, it scores each full candidate answer. It reads the context once, keeps it cached, then tries each possible answer as a different ending and sees which one the model likes best. That costs a bit more, but it works when the options are long, like full tool calls or commands.

These hacks aren't the same as Jev. The models were trained to predict text, not to make decisions, so their confidence numbers aren't calibrated. Still, it shows how much of the behavior you get just by not letting the model talk.

Why it matters

Jev probably isn't a big architectural breakthrough. Most of the pieces existed before it, and open copies appeared within days. What TypeSafe got right is the diagnosis. LLMs are good at reasoning, and regular software is good at acting reliably. Somewhere in between, a fuzzy judgment has to become an exact decision, and until now we've been asking the chat model to do that job by writing carefully.

A dedicated decision model does it better, faster and much cheaper, and it tells you how sure it is. Whether Jev ends up winning or OpenAI, Liquid or an open-source project does, I expect most serious AI systems to have a layer like this within a year.

A caveat: Jev's speed, cost and accuracy figures come from TypeSafe itself, and the benchmark numbers for the open-source models come from their authors. Treat them as claims until someone tests them independently.

Sources