Let me get something out of the way first, because people get this wrong about me all the time. I don't hate AI. I hate the hype around it.

I hate the version of AI that gets sold on stage by guys in expensive hoodies. The one that's going to cure cancer, replace every job, become a god, and maybe kill us all, depending on which week you catch them. What I care about is the boring stuff. Tools that ship. Tools that solve a real problem for a real business and make somebody some money. That's it. That's the whole philosophy.

So when something comes along that is actually useful and actually boring in the best way, I get a little excited. And a new model called Jev, from a startup called TypeSafe AI, is exactly that kind of thing. But to explain why it matters, I need to talk about lasagna first.

The AGI obsession

The whole industry is chasing AGI, artificial general intelligence. One giant brain that can do everything a human can do. Write your emails, do your taxes, diagnose your rash, plan your vacation, write your code, and have a deep conversation about your feelings on the side.

Here's my problem with that. I don't need my car to know how to bake a lasagna. I need my car to drive. I need it to drive well, reliably, every single time I turn the key. If my car could also bake lasagna, that doesn't make it a better car. It makes it a weirder, heavier, more expensive car with more things that can break.

Tools should do their job and do it really well. The idea that everything should collapse into one enormous general system isn't progress. We've actually tried that before, and it didn't go great.

We already made this mistake with servers

If you were in IT back in the Windows NT days, you remember Microsoft Small Business Server. The pitch was simple: why buy a bunch of servers when you can have one box that does everything? Exchange for email, SharePoint for documents, Active Directory for your users, file sharing, the works. All on one machine.

It sounded great on the sales sheet. In practice it was a nightmare, for two reasons.

The first was security. When you cram every service onto one system, you don't just add their features together. You add their vulnerabilities together. A hole in any one of those services became a way into all of them. Your email server problem was now your domain controller problem. One box, one giant target.

The second was resources. A system that does everything needs enough RAM, CPU and disk to do everything, all the time, even if you're only using ten percent of it. You're paying to keep the whole elephant alive when all you wanted was the trunk.

The fix, which the industry figured out over the next couple of decades, was compartmentalization. Break things apart. Put each service in its own box, or its own VM, or its own container. Give each piece only the access and resources it actually needs. If one thing gets compromised, the damage stays in that one little room.

Giant AI models are Small Business Server all over again. One massive system that can write poetry, generate code, browse the web and take actions, all wrapped together. Every capability you add is another door, and they all open into the same house. And the resource cost is absurd. You're spinning up a model that can explain quantum physics just to figure out whether an email is spam.

The better path looks like the one we already took with infrastructure. Smaller, specialized models that each do one job, walled off from each other. Less to attack, far less to run.

OpenAI's real legacy

I'll say something that might sound harsh. I think OpenAI as a company is heading somewhere I don't care about. But I think its real legacy is going to be the people who left.

A lot of the most interesting work right now is coming from OpenAI alumni who walked out and started building smaller, more practical things. Fine-tuning platforms. Specialized models. Tools meant to slot into real software instead of being a chatbot you talk to. TypeSafe AI is one of those. It was co-founded by Diogo Almeida, who worked at OpenAI during the ChatGPT era, and it came out of stealth this month with $40 million in seed money and Jev as its first model.

What AI is actually good at

Forget WarGames. Forget the computer that decides to launch the missiles. Here's what AI is really, genuinely good at, in my opinion: turning messy stuff into clean text that normal code can use.

Think about a security camera. Writing code that reacts to a simple value is trivially easy. If human_detected is true, sound the alarm. A first-year programming student can write that. What's hard is getting from a grainy video feed to that one clean true or false. That's the part traditional code is terrible at, and that's the part AI is great at.

AI is a bridge. On one side you've got the real world, full of video, audio, badly written emails and people who type "yes" seventeen different ways. On the other side you've got software, which wants everything perfectly formatted and predictable. The most valuable thing AI can do is stand in the middle and translate.

The problem with if/else

Here's a tiny example that every programmer has hit a hundred times.

if (answer === "YES") {
  approveRequest();
}

The user types "yes". Doesn't work. They type "Yes". Doesn't work. They type "yeah", "yep", "y", "sure", "YES " with a trailing space, or "yes please". None of it works.

So you start writing sanitizing code. Lowercase everything. Trim whitespace. Build a list of every possible way someone might say yes. Then someone writes "absolutely" and you're back to square one. A shocking amount of real-world programming is just this: dragging messy human input kicking and screaming into a shape your code can handle.

What you actually want to ask is "did this person mean yes?" and get back a clean answer with a confidence level. If the model is 98% sure they meant yes, go ahead. If it's 60% sure, ask them again or send it to a human. That's a much sturdier way to build software. And that's exactly the gap Jev is aiming at.

So what is Jev?

Jev is not a large language model. That's the first thing to understand. It can't write you a poem. It can't draft an email. It can't tell you a story or explain anything. It doesn't produce text at all.

TypeSafe calls it a "System One" model, and the name comes from psychology. Daniel Kahneman split human thinking into two modes. System 1 is fast, instinctive and automatic. You look at a shirt and you just know it's blue. You don't reason your way there. System 2 is the slow, deliberate stuff, like working through a math problem or deciding whether to take a job.

The chatbots everyone uses, the GPTs and Claudes of the world, are basically trying to be System 2. They think things through, write out their reasoning, and generate an answer word by word. That's powerful, but it's slow and it's expensive.

Jev is the other half. It doesn't think things through. It looks at the input and reacts. You hand it some state, which could be text or JSON, plus a set of questions you've defined ahead of time, and it hands back typed answers with a probability attached to each one. According to its documentation, there are a few kinds of questions you can ask: pick one option from a list, give a score, or answer yes or no. It answers all of them in a single pass.

Here's roughly what that looks like in practice. This is simplified pseudo-code, not the exact SDK, but it gets the idea across:

const result = await jev.evaluate({
  state: incomingSupportTicket,
  questions: {
    isIncident: yesNo("Is this reporting an outage or incident?"),
    severity:   choice(["low", "medium", "high", "critical"]),
  },
});

if (result.isIncident.value && result.isIncident.confidence > 0.95) {
  pageOnCall(result.severity.value);
} else if (result.isIncident.confidence > 0.6) {
  sendToHumanReview(incomingSupportTicket);
}

Look at what's happening there. It's an if statement. A smart one, but an if statement. You're not having a conversation with an AI agent. You're calling a function that happens to understand messy input.

Why the format thing matters so much

If you've ever tried to get a regular LLM to return JSON inside a real application, you know the pain. You ask for a specific structure and most of the time you get it. Then every so often it adds a field you didn't ask for, drops one you did, wraps the whole thing in a friendly sentence, or invents a value that isn't one of your options. Your parser throws an error and your pipeline falls over at 3 a.m.

Newer "structured output" features help, but they're still bolted onto a model whose natural instinct is to generate text. Jev comes at it from the other direction. The answers are defined before the model ever sees the input. If you ask it to pick from four severity levels, it physically cannot come back with a fifth one, or with a paragraph. The output always matches the schema, because there's nothing else it's able to produce.

TypeSafe markets this as "no hallucinations." I'll be honest, that phrase makes me roll my eyes a bit. Sure, a model that never writes text can't invent a fake citation. But it can absolutely give you the wrong answer. It can be 97% confident that a cat video is a security incident. In the real world we have a technical term for that, which is that the model messed up. Calling it something fancier doesn't change the fact that your code still has to plan for it being wrong.

The difference, and it's a real one, is that Jev tells you how sure it is. Every answer comes back with a probability. That means the uncertainty isn't hidden inside a confident-sounding paragraph. It's a number sitting right there, and you get to decide where the line is. Whether those numbers are actually well calibrated in the wild is the big open question, and it's one people are going to have to test for themselves. TypeSafe hasn't published a technical paper, and most of the benchmark numbers so far are its own.

The name is the whole pitch

Jev is named after the Jevons paradox, and once you understand that, you understand the business.

Back in the 1800s, an economist named William Stanley Jevons noticed something strange. As steam engines got more efficient and used less coal per unit of work, total coal consumption didn't go down. It went way up. Cheaper power made steam worth using for thousands of things nobody would have bothered with before.

That's the bet here. Jev doesn't do anything a big LLM couldn't technically do. You could already ask GPT or Claude whether an email is urgent. The problem is that it's slow and expensive enough that you'd never do it on every email, every log line, every chat message, every frame of a game. Make that decision cheap enough and fast enough, and suddenly you'll put it everywhere.

The numbers are silly

And cheap it is. TypeSafe's stated pricing is $42 per billion input tokens. That works out to about four cents per million. Output tokens are free, which makes sense, because the output is just a handful of typed answers. Compare that to regular LLMs, which run anywhere from around twenty cents to several dollars per million tokens depending on the model, and it's not in the same universe.

Speed is the other half. A reasoning model can take anywhere from a few seconds to a few minutes to answer something. Jev is built to answer in well under a second. That's the difference between "a thing you call in a background job" and "a thing you can call inside a loop that runs over and over."

TypeSafe's launch demos lean hard into this. One has Jev playing Doom off structured game state at around ten decisions a second, costing roughly seven dollars an hour. Early users have reported sorting emails at a fraction of a second apiece, and classifying tens of thousands of chat messages for about the cost of a nice dinner. Those are self-reported numbers, so take them with the usual grain of salt, but even if they're off by a lot, the economics still look wild.

This is AI's NoSQL moment

The best comparison I can come up with comes from databases.

For decades, SQL databases ruled everything, and for good reason. Their number one priority is data integrity. Every transaction is correct, every record is where it should be, nothing gets lost. That's exactly what you want for a bank. But that guarantee is expensive, and it makes SQL databases hard to scale to truly huge numbers of users.

Then NoSQL came along and made a trade. It gave up some of that perfect consistency in exchange for speed and the ability to scale out almost endlessly. If your like count on a post is off by one for half a second, who cares? That trade is a big part of why platforms like Twitter could exist at the scale they did.

Jev is making the same kind of trade. It gives up general intelligence. It gives up the ability to write, explain, reason or chat. In exchange it gets speed and cost that general models can't touch. It isn't trying to replace the big models, any more than NoSQL replaced SQL. It's opening up a whole category of things that just weren't practical before.

A flock of birds

Years ago I talked with an executive from a quantum computing company, and he said something that stuck with me. We're used to computers being exact. Two plus two is four, every time, to the last decimal. But a lot of the biggest systems in the world don't actually need that. They need to be pointed in the right direction.

He compared it to a flock of birds migrating. Not a single bird is flying the mathematically perfect heading. Each one is a little off. But the flock as a whole gets where it's going, and it does it efficiently.

That's what cheap, fast, probabilistic decisions make possible. You don't need every single call to be perfect. You need thousands of pretty-good decisions per second, each with an honest confidence score, all adding up to a system that heads the right way. That's a different way of building software, and I think it's where a lot of things are going.

Picture a video game that actually reads you

Here's the use case that gets me most excited. Right now, when you play a game, you pick Easy, Normal or Hard at the start, and that's that. Some games do "dynamic difficulty," but it's usually pretty crude.

Now imagine each enemy on screen checking in several times a second. How fast is this player reacting right now? Are their inputs getting sloppy? Are they getting tired or frustrated, or are they cruising and bored? And then each enemy adjusts, a little more aggressive here, a little slower there, so the game stays right on the edge of challenging for exactly where your head is at that moment.

You could never do that with a big LLM. Too slow, and at millions of players, way too expensive. At fractions of a cent and a response well inside a second, it starts to look possible. Same goes for interfaces that rearrange themselves as you use them, or bots that react to live data frame by frame.

Where it fits, and where it really doesn't

This is the part that matters most if you're actually going to build with it, because the easiest mistake here is to treat Jev like a smaller, cheaper chatbot. It isn't one.

Good fits. Anything that's basically a smart if statement or a routing decision. Is this an incident, and how bad? Is this comment spam, abuse or fine? Which team should this ticket go to? Should this coding agent's request to run a command be auto-approved or kicked to a human? Classifying and tagging huge piles of text. Anywhere you'd currently write a brittle regex, or anywhere you'd love to call an LLM but can't justify the cost or the wait.

Don't use it as a judge. There's a popular pattern where one AI grades another AI's output. Is this code good? Is this answer correct? Jev is the wrong tool for that. Deciding whether a complicated piece of code actually works takes real reasoning. That's System 2 work. A fast instinctive model will give you a fast instinctive answer, and a confidence number on a bad judgment is still a bad judgment.

Don't use it to compress agent memory. Long-running AI agents build up huge conversation histories, and people often squeeze those down so the agent can keep going. Jev is a bad fit here too. Its context window is small, and more importantly it can't write a summary, because it doesn't write anything. Compressing an agent's history well means keeping the reasoning that got it to where it is. Lose that and the agent gets lost.

The 10-second rule. The simplest test I've heard: if a person could look at the data and answer the question in under ten seconds, Jev is probably the right tool. Is this email angry? Is this a refund request? Is there a person in this frame? Quick, gut-level calls. If a person would have to sit and think about it, check some facts, or weigh trade-offs, you want a reasoning model instead.

And the mindset shift is simple. Treat Jev like a library, or a switch statement, or a function. Don't treat it like a worker you hand tasks to. It sits inside your code. Your code stays in charge.

Boring is good

I've been doing this long enough to have a real soft spot for tools that do one thing and do it well. SQLite. Vector databases. The Model Context Protocol, which gives AI tools one standard way to plug into other software. None of them are sexy. None of them will end up in a movie. All of them get used every day to build things that work.

Jev feels like it belongs in that group. It isn't trying to be a god. It isn't going to replace your job. It isn't going to write your novel. It takes messy inputs, makes quick decisions, tells you how sure it is, and costs almost nothing. That's a tool. And frankly, after years of hearing about digital gods, a tool is exactly what I wanted.

If you want to kick the tires, it's in early access right now, and you can reach it through Vercel's AI Gateway and OpenRouter. Go build something boring with it.