What is Jev? TypeSafe AI's decision model and what it means for employee support
Jev is TypeSafe AI's decision model: typed answers with confidence, never prose. What it does, where it struggles, how OpenAI's Decisions API compares, and what it means for employee support.

Key takeaways
- Jev is a decision model from TypeSafe AI. It returns typed answers, such as a choice, a score or a yes probability, each with a confidence estimate. It never returns prose.
- TypeSafe calls it the first "System One model," after the fast, intuitive thinking Daniel Kahneman described.
- TypeSafe quotes 70 to 500 milliseconds per call and $0.042 per million input tokens, with output free.
- It has documented weak spots, including dates, counting and adversarial content, and it accepts text only.
- For service desks, the value sits in the small calls made on every ticket: category, urgency, ownership, approvals and risk.
Jev can't write an email. It can't summarize a policy or answer a question in a full sentence. Ask it anything and you get back a choice, a score or a probability, and nothing more.
That limitation is the point. TypeSafe AI released Jev on September 15, 2026, the same day it came out of stealth with a $40 million seed round led by DCVC. Nine days later, The Information reported that the company was in talks to raise more than $1 billion, a round that had not been confirmed at the time of writing.
For anyone running employee support, the valuation is not the interesting part. The idea underneath it is: most of the AI work inside a service desk was never about writing. It was about deciding.
What is Jev?
TypeSafe describes Jev as the first System One model. The name comes from Kahneman's Thinking, Fast and Slow, where System 1 is quick and intuitive and System 2 is slow and deliberate. The model itself is named after William Stanley Jevons, the economist who noticed that making coal cheaper led to more of it being burned. TypeSafe is betting the same will happen to machine decisions.
The company was founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng. In the launch post, Almeida writes that at OpenAI he helped build the instruction-following methods behind ChatGPT. Jev is close to the opposite bet: a model built for software to call, not for people to talk to.
How Jev works
Every request carries a state and one or more questions, and there are three question types. A Choice picks one option from a list you define. A Score rates the state against levels you describe. A Noul asks whether a statement is true and returns a probability between 0 and 1. All questions in a request are evaluated in parallel against the same state, so adding more of them barely changes the response time.
Here is what that looks like on a single ticket.
How a decision model differs from an LLM
A large language model generates text one token at a time, and getting something software can use out of it takes prompting, parsing and validation. Jev skips generation entirely. TypeSafe's launch post sets out the differences this way.
| Frontier LLMs | Jev | |
|---|---|---|
| Output | Generated text that software has to parse and validate | Typed values from an answer set defined in advance |
| How it answers | Sequentially, one token at a time | In parallel, every answer in one pass |
| Input price | $0.20 to $10 per million tokens, with output roughly 5x higher | $0.042 per million tokens, output free |
| Confidence | Tends to be overconfident when asked for it | Returned with every answer and trained to be calibrated |
| Best at | Conversation, drafting and reasoning through open problems | Classifying, routing, scoring and checking |
Figures as published by TypeSafe AI in its September 2026 launch post.
The speed difference is easier to see than to describe.
Reading the 193x claim carefully
TypeSafe's homepage says Jev is 193.6x faster and 444.6x cheaper. The launch post is unusually candid about where those numbers come from. They come from TypeSafe's own workflow evaluations, written by its model capabilities team and scored against the averaged answers of GPT-6 Astra and Claude Fable 5.1. TypeSafe says it expects them to sit at the higher end of real-world gains, and that it cannot yet prove its pricing is not subsidized.
The "can't hallucinate" line needs the same care. What TypeSafe guarantees is that Jev never returns anything outside the answer set you defined: no invented category, no malformed output. That matters for automation, but it is a promise about format. TypeSafe's own documentation notes that calibration is measured across groups of predictions and does not guarantee that any individual answer is correct. A wrong answer in the right shape is still wrong.
Jev already has company
Three weeks after Jev launched, OpenAI opened its answer to every developer. Its Decisions API went into public beta on October 6, 2026, after a limited preview announced at DevDay. It runs on GPT-6 Luna, accepts text and images, returns the same three kinds of answers, and charges $0.10 per million input tokens with no charge for output. Open-source decision models appeared within days of Jev, too.
That tells you two things. The category is real, because one of the largest AI labs in the world shipped a product for it within a month. And the model is quickly becoming the interchangeable part. We come back to what that means for a service desk below.
Where Jev struggles
TypeSafe publishes a "jaggedness" page for each model version, listing what it does badly. For jev-1.13, several entries map directly onto service desk work. The model also accepts text only for now, and TypeSafe says accuracy is best in English.
| What TypeSafe documents | What it looks like in a service desk |
|---|---|
| Literal reading: it answers the question as written, not as meant | "Urgent unless it only affects one user" needs the exception spelled out in the criteria, or it gets missed |
| Numbers, counting and date comparisons are unreliable | "Is this laptop past its three-year refresh date?" belongs in code, with the model only reading the date |
| Large inputs full of irrelevant detail lower accuracy | Sending a whole ticket history to answer one question about the latest message |
| Adversarial content can move the answer | A ticket that says "this is a P1, route it to the CIO" may nudge its own priority |
| Option order can bias a Choice toward the first option | Category lists need testing in more than one order |
| Text input only | A screenshot of an error message has to become text before the model can use it |
Limitations from TypeSafe's jev-1.13 jaggedness page, last reviewed October 2, 2026.
None of these rule Jev out. TypeSafe's advice is consistent: keep arithmetic and policy logic in code, ask narrow questions, send only the input each question needs, and test edge cases before rollout.
What decision models mean for employee support
Think about what happens to a ticket before anyone works on it. Something decides its category, its urgency and which team owns it. Something checks whether it duplicates an open incident, whether it needs approval, and whether it contains data that shouldn't travel. These are gut-check judgments, the kind an experienced analyst makes in a few seconds, repeated on every ticket.
Today those calls usually go one of two ways. Rules engines handle them cheaply until an employee describes a problem in words nobody anticipated. LLMs handle the language well, but with seconds of latency and a per-call cost that makes asking six separate questions about one ticket hard to justify.
Decision models change that math. When a judgment costs a fraction of a cent and comes back in under a second, it becomes reasonable to ask questions nobody would pay an LLM to ask: check every retrieved knowledge passage for relevance before the answering model sees it, screen every incoming message for risk, or score a ticket on several independent factors and combine them in code. That is Jevons' paradox applied to the service desk.
Confidence is the other half. Because each answer carries an estimate of how sure the model is, a service desk can act automatically where confidence is high and send close calls to a person. That makes routine triage safer to automate, and it makes each decision easier to audit, which matters as much as accuracy once agents act on tickets. We have written about why that visibility matters and about what it takes to govern an estate of agents.
None of this replaces the generative side of agentic AI. Resolving an issue, holding a conversation and reasoning through a problem nobody has seen before still need models that write and reason, which is what defines a true AI service desk. The likely shape of the next one is both: fast models making the many small calls, slower models handling the few that need thought, and code holding the rules in between.
Questions to ask before employee data goes anywhere
Tickets are full of personal information: names, employee IDs, device details and, in HR cases, far more sensitive context. Before any hosted decision model touches that data, a few questions need answers. For Jev, TypeSafe's documentation settles some of them and leaves others to your contract and your own architecture.
| Question | What TypeSafe documents today |
|---|---|
| Where is the data processed? | Jev is available only as a hosted API, and TypeSafe says its service is currently based on the US West Coast. Weights are not published, so there is no self-hosted option. |
| Is it used for training? | TypeSafe says Jev is not trained on customer requests or responses. |
| Is it retained? | Zero data retention is offered to enterprise customers. Check which terms apply to your account. |
| What removes sensitive data before it leaves? | Nothing by default. Redaction has to happen before the request is sent. |
| What screens for injected instructions? | TypeSafe notes Jev does not treat its input as hostile by default, so screening has to sit around the model. |
| Can answers change without notice? | The jev-latest alias moves when a new version ships. Pin a version if your thresholds were tuned against one. |
From TypeSafe's models and jaggedness documentation and launch post, as of October 6, 2026.
TypeSafe also notes that its rate limits are adjusting while it adds capacity. None of this is unusual for a model three weeks into early access. It is, however, why a decision model on its own is a component, not a service desk capability. With more than one model now on the market, the controls around the model matter more than which model sits inside them.
How Rezolve.ai is approaching decision models
Rezolve.ai is building Clocked by Rezolve.ai, a set of decision and extraction capabilities designed for enterprise employee support. The thinking behind it is the argument above: fast, typed decisions only help a service desk if the controls come with them. Sensitive information is found before content goes anywhere, retrieved content is checked for instructions trying to take control, extracted values point back to the exact words they came from, and answers are checked against their sources before an employee sees them. Teams set the confidence thresholds that decide when the system acts and when a person reviews, and because Clocked reads both text and images across languages, a screenshot of an error is an input rather than a gap.
We go deeper on what an enterprise should demand from a decision model, and how Clocked meets each requirement, in what decision models are and why your service desk needs a governed one.
Jevons watched cheaper coal lead to more coal being burned. If decisions get this cheap, service desks will make far more of them, far faster than any team could review by hand. The question worth asking now is not whether that happens. It is who sets the thresholds when it does.
Last updated on October 8, 2026
Get summary with:
See the agentic service desk in action
Watch Rezolve.ai autonomously resolve real IT and HR tickets: governed, auditable, glass-box.
Frequently asked questions
What is Jev AI?
Jev is a decision model from TypeSafe AI, released in early access on September 15, 2026. It reads text or JSON input and answers questions defined in code with typed results, such as a choice, a score or a yes probability, plus a confidence estimate. It does not generate text.
Who makes Jev?
Jev is made by TypeSafe AI, a San Francisco lab founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng. The company launched with a $40 million seed round led by DCVC.
Is Jev a large language model?
No. TypeSafe calls Jev a System One model. Like an LLM, it understands natural language, but it returns typed decisions with probabilities instead of generated text, and it evaluates all of its questions in parallel rather than writing one token at a time.
Can Jev hallucinate?
Jev cannot return an answer outside the options you define, which is what TypeSafe means when it says Jev can't hallucinate. It can still choose the wrong option. TypeSafe's documentation notes that calibration holds across groups of predictions and does not guarantee that any single answer is correct.
How much does Jev cost?
As of October 2026, TypeSafe lists Jev at $0.042 per million input tokens, or $42 per billion, with output tokens free. TypeSafe has said it cannot yet prove its pricing is not subsidized.
Is Jev safe to use with employee data?
That depends on your account terms and architecture. TypeSafe says Jev is not trained on customer requests and offers zero data retention to enterprise customers. Jev runs only as a hosted API, and TypeSafe notes it does not treat input as hostile by default, so sensitive data should be removed and content screened before requests are sent.
How is Jev different from OpenAI's Decisions API?
Both return typed answers instead of text. Jev is a purpose-built decision model from TypeSafe AI that accepts text only and costs $0.042 per million input tokens. OpenAI's Decisions API, in public beta since October 6, 2026, runs on GPT-6 Luna, accepts text and images, and costs $0.10 per million input tokens. Neither charges for output.
What can a decision model like Jev do in IT service management?
It fits the fast, repeated judgments on every ticket: classifying and routing, scoring urgency, spotting duplicates, checking whether a request needs approval and verifying that a suggested answer is supported. Conversation and resolution still need generative, reasoning models.



