Skip to content
DSRPT
Oct 1, 2026 · 6 min read

Jev AI vs LLMs: What a Decision Model Does Differently

Jev is a decision model from TypeSafe AI. It reads some text and a list of questions, then returns typed answers with probabilities, not prose. It is much faster and cheaper than an LLM for routing, triage and moderation, and useless for writing or reasoning. Ollama now runs Jev-style models (Nimble, Tev1) locally, so you can self-host the idea.

Abdulkader Safi
Abdulkader Safi Senior Software Engineer
Share:
Jev AI vs LLMs: What a Decision Model Does Differently

Ask a chat model whether a support ticket is urgent and you get a polite paragraph. You wanted one word and a confidence number.

That gap is why Jev exists. TypeSafe AI launched it on 15 September 2026, and it works differently from GPT, Gemini, Llama, Claude or any other LLM you have used. This post covers what Jev is, where it beats a language model, where it falls flat, and how to run a Jev-style model on your own machine with Ollama.

What Jev is

Jev is a decision model. You give it some text (a ticket, a chat log, a JSON blob) and a set of questions. It gives back typed answers with probabilities. No prose.

There are three question types:

  • Choice: pick from a fixed list of options (up to 255), with a probability for each
  • Score: rate the text on an ordered scale of 2 to 10 levels, like severity or frustration
  • Yes/no: one number between 0 and 1, where the number itself is the answer

TypeSafe calls it a System 1 model. Fast, reflexive, a gut call. General LLMs are System 2: they think in words, step by step, and write out the result. The company was founded by a former OpenAI researcher, which is part of why the launch got attention. Vercel reported that 13% of its paid AI Gateway users tried Jev within 24 hours. If terms like System 1 or agent are new, 13 new AI terms every founder should understand covers the basics.

How Jev differs from an LLM

An LLM is autoregressive. It writes one token, then the next, then the next, until the reply is done. Jev is described as non-autoregressive. It reads the input and the allowed options, then returns the decision in one pass. We covered another model that drops the one-word-at-a-time habit in DiffusionGemma: the AI model that writes in blocks.

That changes four things:

  • Output shape is fixed. You never get "Sure! Here's my analysis" wrapped around your answer, and you never parse JSON that came back broken.
  • Speed. Reported latency is 70 to 500 ms per request. A loop of several questions runs about 10 times a second. A chat model needs seconds for the same loop.
  • Price. $0.042 per million input tokens, output free. One outlet compared that with a frontier LLM at $10 in and $50 out per million tokens.
  • Confidence comes with every answer. TypeSafe trains for calibrated confidence rather than plausible-sounding text.

Side by side:

Job LLM Jev
Output Prose you have to parse Typed choices with probabilities
Best at Writing, code, reasoning Routing, classification, guardrails
Speed per call Seconds Tens to hundreds of milliseconds
Explains itself Yes No
Handles open-ended tasks Yes No

Most of the speed and cost numbers come from TypeSafe. I could not find independent benchmarks, and TypeSafe publishes none of its own. Test it on your own data before you trust any of it.

What Jev cannot do

I would not sell this to a client without saying these out loud:

  • It cannot write or explain anything
  • It cannot do maths or multi-hop reasoning
  • It cannot compare dates
  • It struggles with hostile or adversarial input unless you spell out the criteria
  • It guarantees the shape of the answer, not that the answer is right. It can pick the wrong option from the list
  • It gives no reasoning trace, which hurts any workflow where a human needs to see why

The demos so far (a Minecraft bot, a drone simulator, a Subway Surfers-style game) all run in simulators. Nobody has shown it on real hardware with real safety requirements.

Where Jev fits and where an LLM still wins

Use a decision model when the answer is one of a known set: ticket triage, model routing, tool selection, content moderation, checking whether a retrieved passage answers the question, deciding when to escalate to a human.

Use an LLM when someone has to read the output, or when the task needs reasoning across steps.

My opinion: the good setup is both. Jev sits at the front door and decides which LLM, tool or person gets the job. A cheap fast router in front of an expensive slow model. I have not run this in production yet, so treat it as a design idea to test, not a result.

The self-hosted Jev alternative: Ollama

Ollama now supports Jev-style decision models. They follow TypeSafe's Jev API spec but run on your own machine, so there is no per-token bill and no network hop. If running models on your own hardware is new to you, read Gemma 4 12B: AI that runs on your laptop first.

Three models are available:

  • Nimble: 9B parameters, open source, by Bespoke Labs
  • Tev1: 4B parameters, experimental, by Together AI
  • Tev1 0.8B: under a billion parameters, experimental, also Together AI

You need Ollama 0.35 or later:

ollama --version   # needs 0.35 or newer
ollama pull nimble
# smaller options
ollama pull tev1
ollama pull tev1:0.8b

Then send requests to the new /v1/systemone endpoint with curl or TypeSafe's Python SDK. The idea is the same as hosted Jev: send the text as state, add named questions (choice, yes/no, score), and get answers with probabilities back in one request.

Ollama reports Nimble 9B averaging 91 ms per decision on an M5 Max MacBook Pro. Bespoke Labs also benchmarked Nimble and Tev1 against Jev 1.13 across 13 public datasets and 3,880 decisions. I did not find the score table in Ollama's post, so pull those numbers from Bespoke Labs before you pick a model.

Hosted Jev or self-hosted

Self-hosting wins on:

  • Privacy. Ticket text with customer details never leaves your machine or server
  • Cost. No per-token charge. You pay for hardware
  • Latency. No round trip to someone else's API

Hosted Jev wins on:

  • Being the reference model. The Ollama models are compared against it, not the other way round
  • No hardware to run. No GPU to size or babysit
  • 64k tokens of context per request, which the smaller models may not match

If you want to stay open source, start with Nimble. If you just want to see whether decision models fit your workflow, the hosted API is the faster test.

What to do now

  1. Pick one decision your automations already make with an LLM call. Ticket priority, real lead or spam, which model handles a request.
  2. Log 100 real examples with the answer you would want.
  3. Run them through Nimble on Ollama, then through hosted Jev if you have access.
  4. Compare accuracy, latency and cost against your current LLM call.

If the numbers hold, swap the decision step and keep the LLM for the writing. I would not replace an LLM with Jev. I would put Jev in front of it.

If you want help wiring decision models into a real workflow, talk to DSRPT.

Sources: Ollama announcement, TechTarget, MindStudio, DataNorth

FAQ

Q1:

What is Jev AI?

A1:

Jev is a decision model from TypeSafe AI, launched on 15 September 2026. You give it some text and a set of questions, and it returns typed answers (a choice, a score, or a yes/no probability) with confidence values. It does not write text or hold a conversation.

Q2:

Is Jev an LLM?

A2:

No. TypeSafe describes it as a System 1 model that is neither small nor an LLM. LLMs generate text one token at a time. Jev is described as non-autoregressive and returns its decision in a single pass, so it cannot write, explain or handle open-ended tasks.

Q3:

How much does Jev cost?

A3:

TypeSafe's published price is $0.042 per million input tokens, with output tokens free. That figure comes from TypeSafe and the outlets covering the launch. Check the current pricing page before you plan a budget around it.

Q4:

Is there a self-hosted alternative to Jev?

A4:

Yes. Ollama 0.35 and later supports Jev-style decision models through a /v1/systemone endpoint. The available models are Nimble (9B, by Bespoke Labs), Tev1 (4B, by Together AI) and Tev1 0.8B. They follow the Jev API spec but they are not Jev itself.

Q5:

Which Ollama decision model should I start with?

A5:

Start with Nimble. It is the 9B open-source model, and Ollama reports about 91 ms per decision on an M5 Max MacBook Pro. If your hardware is small, try Tev1 or Tev1 0.8B, but both are marked experimental.
NEWSLETTER

Stay Ahead of the Curve

Get the latest digital marketing insights delivered to your inbox weekly.