articleJev doesn’t write, it decides: I opened it up, built a product, and measured how honest it is
September 28, 2026 · selma kocabıyık
Jev doesn’t write text. It doesn’t count, doesn’t see images, doesn’t go online. The one thing it does is decide: you give it a state and typed questions, and every question comes back with a probability. In this piece I first open it up with an engineer’s eye, then show the product I built with it, and finally measure how honest its percentages are, using my own data.
1. Jev is not a chatbot
You ask ChatGPT a question and it writes a sentence. You ask Jev a question and it gives you a number. The input is called the state: a piece of text, a summary of a situation, a product description. On top of it you write typed questions. There are three types:
- choice: “which of these five labels?” Each label gets a probability, and they sum to 1.
- noul: yes or no. A single number between 0 and 1.
- score: a scale you define, for example low / medium / high.
You ask all of them in one call. In the words of the official docs, the questions are evaluated in parallel and in isolation: the answer to the first question never leaks into the context of the second. Here is an example with a sentence from my own footage (Turkish filler, roughly “um, like, I mean, let me start over”):
- self-criticism0.64
- filler0.21
- transition0.08
- repetition0.05
- teaching0.02
- cut0.74
2. What runs inside: a loop or a single pass?
An LLM writing an answer sits inside a loop. For every token it runs through the whole network once, gives a score to each of the roughly one hundred thousand tokens in its vocabulary, softmax turns those scores into probabilities, one token is picked, appended, and it starts again. A 200-token answer means 200 passes. That is why it is slow and why output tokens cost money.
Plainly, softmax turns the model’s raw scores into percentages that sum to 1. Because the total is fixed, the options compete; if one gets a bigger share, the others shrink.
Jev has no loop. It reads the text once. The options you provide become the entire vocabulary of that question: five labels instead of a hundred thousand tokens, and just once. Confidence isn’t a separate feeling, it’s the shape of the distribution: piled up in one place means high, spread out means low.
Why it’s cheap
There are no output tokens, so output is free. What you pay for is input. Asking 13 questions in one call came out about 12.2× cheaper than asking them one by one, because the text travels once instead of 13 times. On Vercel AI Gateway the price is 0.042 dollars per million input tokens.
3. Its relative from school: the classifier
If you studied computer science or AI, you know its relative: the classifier. Spam or not, a review positive or negative. The BERT models on HuggingFace work like this. First they learn general language from a huge amount of text (pre-training), then they are adjusted on labelled data for specific classes (fine-tuning).
BERT is an encoder: it reads the whole sentence at once. Self-attention in one line: each word decides for itself how much to look at every other word in the sentence when working out its meaning. At the end, each class gets a probability.
Jev looks like the frontier version of that family. The difference: in a classic classifier the classes are fixed during training. With Jev you write the classes at call time, and you don’t train anything.
One confusion to close early: Jev is not an embedding model. An embedding gives you a list of numbers (a vector) and you compute similarity yourself. Jev gives you a decision.
4. What gets rewarded: RLHF, RLVR, RLCD
A model first learns by reading the internet. In a second stage it is steered with reinforcement learning: the model does something, gets a reward if it is good, and shifts towards whatever earns rewards. Like training a dog. Whatever you reward, the model becomes. The only difference between the three methods below is what the treat is given for.
RLHF: human preference
Reinforcement Learning from Human Feedback. Three stages: the model learns by reading, people rank answers and a reward model is trained on those rankings, then the main model goes into a loop scored by the reward model. The human is no longer in the loop; the reward model scores in their place.
The result: polite, fluent, persuasive answers. It is also why chat models always sound confident. They were rewarded for being liked, and confident answers get liked more. Being liked is not being right; on “how sure am I, in percent” they are not honest.
RLVR: an answer key
Reinforcement Learning with Verifiable Rewards. When the answer can be checked exactly, the reward is easy: right or wrong. Maths, code. Reasoning models are trained this way; they are accurate but slow and expensive because they think at length. But “should this sentence be cut in the edit?” has no answer key.
RLCD: an honest percentage
The method TypeSafe uses for Jev is called Reinforcement Learning for Calibrated Decisions. Here the reward goes to whether the stated probability matches reality. If the things it calls 70% come true 7 times out of 10, reward; if it says 90% and gets half right, penalty. This is called calibration: what you call 20% should really happen 1 time in 5.
The practical value of a calibrated percentage is that you can set thresholds: above 85% auto-approve, below 20% auto-reject, everything in between goes to a human. If the percentage isn’t honest, those thresholds mean nothing.

5. “It can’t hallucinate”?
That sentence is half true. The official claim is only about type errors: Jev can’t make a type error. The schema is defined up front and the answer is always one of the labels you gave it. Ask an LLM for JSON and sometimes it comes back broken; that doesn’t happen with Jev.
But it can give a wrong answer that is still valid. The label is valid, the decision is wrong. The structure doesn’t break, the judgement does. So the real question is: how sure is it when it’s wrong? I measured that below.
TypeSafe’s own “jaggedness” doc is open about what it can’t do:
- it doesn’t count
- it doesn’t do arithmetic
- it doesn’t compare dates
- it reads text literally
- performance drops with a large state
- it is open to adversarial text written to fool it
My rule: maths in code, judgement in Jev.
6. Kıyafet Bul: the catalogue in code, the decision in Jev

I left theory and built a product. I pulled 561 outerwear and knitwear products from Trendyol once (Apify, 63 cents); they sit in a file. Kıyafet Bul (“find an outfit”) isn’t a search engine. You write what you want as a normal sentence and the rest goes through this pipeline. Code: github.com/selmakcby/jev-kiyafet-bul
- 1youyou write“a brown coat for the office, up to 2500 lira”
- 2Jevextracts filterstype, colour, length, belted, hooded: all in one call.
- 3codenarrows downfrom 561 products to at most 60 candidates. Price and size come from code, because Jev doesn’t read numbers.
- 4Jevscores candidatesa “does it match the request” score and confidence per product; the reason when it doesn’t.
- 5coderankssorts by score × confidence and lays out the results.


Type “mayo” (swimsuit) and Jev says “not in this catalogue”, and none of the remaining steps run. That is the cleanest example of a smart if: it can say “no” to something that isn’t in the option list.

7. Where would you use this?
Clothing was just one example. The logic is the same everywhere: you need a decision, and you want to know how sure that decision is.
8. I measured the calibration
This is what I really wanted to know: when Jev says 80%, is it really 80%? I took the answer key from my own video. I pulled 598 sentences from the raw footage; I had actually cut 392 of them in the edit. For each sentence I asked Jev “should this sentence be cut in the edit?” and compared its probability with my real decision.

| bucket | sentences | Jev said | actually cut |
|---|---|---|---|
| 0.3–0.4 | 34 | 0.36 | 0.62 |
| 0.4–0.5 | 166 | 0.45 | 0.51 |
| 0.5–0.6 | 216 | 0.54 | 0.67 |
| 0.6–0.7 | 136 | 0.64 | 0.75 |
| 0.7–0.8 | 42 | 0.73 | 0.83 |
| 0.8–0.9 | 4 | 0.84 | 1.00 |
The curve rises overall: the higher Jev’s number, the more often I had cut, so it knows the order. But it is timid: it never says 95%, and the top bucket has only 4 sentences. The observed rates across the buckets are 0.62 / 0.51 / 0.67 / 0.75 / 0.83; in most buckets the real cut rate is higher than what Jev said.
598 requests took 23 seconds, about 1.3 cents in total.
9. The live judge

Then I hooked Jev up to myself. The Mac’s microphone listens in 3-second chunks, mlx-whisper transcribes, each sentence goes to Jev, and a panel labels it: teaching, filler, repetition, self-criticism, transition; plus a “cut” percentage. These are my own editing rules. Code: github.com/selmakcby/jev-canli-yargic
- 1codemicrophone3-second chunks
- 2codewhispertranscribes, about 0.6 s
- 3JevJevlabel + cut percentage, median about 290 ms
- 4codepanelputs the label and percentage on screen
The slow part isn’t Jev, it’s the audio side. In rehearsal it labelled six sentences out of six correctly: “I think I explained that badly” got self-criticism 1.0; saying the same sentence a second time got repetition 0.98.

10. Claude Code + Jev: a job-listing matcher
The last experiment is a job-listing matcher. I told Claude Code I was looking for listings that fit my LinkedIn profile. Claude Code pulled current LinkedIn listings with Apify, asked Jev about each one (does it fit me, what percentage; which role; which level) and wrote the result to Excel. Claude Code doesn’t know Jev, so I installed the TypeSafe skill from their docs; the skill works as its user manual. The keys were already in .env.
Out of 45 listings: 1 “apply for sure”, 16 “could apply”, 28 “pass”. The Jev side cost 0.23 cents. Claude Code fetched the data, Jev made the decisions, Claude Code wrote the Excel.
11. Bonus: Minecraft

Everyone has done this; I tried it for a minute. Planning and movement live in code, reflexes in Jev. Every 600 milliseconds the code writes the state: health, hunger, is it night, is the nearest enemy close, mid or far. Jev decides in one call: fight, flee, eat, gather wood, shelter, explore, wait; plus a danger probability. Code carries it out.

I give the options, Jev makes the decision, code does the doing. That has a consequence: if “gather iron” isn’t in the list, Jev won’t mine iron even when it sees it. The panel logged 2893 decisions, about 7 cents in total. Code: github.com/selmakcby/jev-minecraft-bot


12. Sources
- github.com/selmakcby/jev-kiyafet-bul: Kıyafet Bul.
- github.com/selmakcby/jev-canli-yargic: the live judge.
- github.com/selmakcby/jev-minecraft-bot: the Minecraft bot.
- TypeSafe: introducing System One models and Jev, the launch post.
- docs.typesafe.ai: state, question types, call structure.
- Jev 1.13 jaggedness: what it can’t do, in their own words.
- Vercel AI Gateway: Jev: model name
typesafe-ai/jev, 0.042 dollars per million input tokens, output free.
You can get a Jev key in two places: the TypeSafe console (there is a waitlist) or Vercel AI Gateway.
Where in your work is there a decision being made that you can’t write code for? That’s Jev’s spot. If you try it, tell me in the comments what came out.
I Built My Agents an Office in 3D: Claude Code + the Higgsfield API