selma kocabıyıkselmaAI engineer
article
2026-09-28TR · EN · NL

Jev doesn’t write, it decides: I opened it up, built a product, and measured how honest it is

September 28, 2026 · selma kocabıyık

Jev doesn’t write text. It doesn’t count, doesn’t see images, doesn’t go online. The one thing it does is decide: you give it a state and typed questions, and every question comes back with a probability. In this piece I first open it up with an engineer’s eye, then show the product I built with it, and finally measure how honest its percentages are, using my own data.

1. Jev is not a chatbot

You ask ChatGPT a question and it writes a sentence. You ask Jev a question and it gives you a number. The input is called the state: a piece of text, a summary of a situation, a product description. On top of it you write typed questions. There are three types:

  • choice: “which of these five labels?” Each label gets a probability, and they sum to 1.
  • noul: yes or no. A single number between 0 and 1.
  • score: a scale you define, for example low / medium / high.

You ask all of them in one call. In the words of the official docs, the questions are evaluated in parallel and in isolation: the answer to the first question never leaks into the context of the second. Here is an example with a sentence from my own footage (Turkish filler, roughly “um, like, I mean, let me start over”):

state · input“Şey, hani, yani, baştan alayım.” (“Um, like, I mean, let me start over.”)three questions · one call
choice · what kind of sentence is this?
  • self-criticism0.64
  • filler0.21
  • transition0.08
  • repetition0.05
  • teaching0.02
0.64 + 0.21 + 0.08 + 0.05 + 0.02 = 1.00 sum
noul · should it be cut in the edit?
  • cut0.74
score · energy of the delivery
lowmediumhigh
One sentence, three questions, one call: self-criticism 0.64 · cut 0.74 · energy medium.

2. What runs inside: a loop or a single pass?

An LLM writing an answer sits inside a loop. For every token it runs through the whole network once, gives a score to each of the roughly one hundred thousand tokens in its vocabulary, softmax turns those scores into probabilities, one token is picked, appended, and it starts again. A 200-token answer means 200 passes. That is why it is slow and why output tokens cost money.

Plainly, softmax turns the model’s raw scores into percentages that sum to 1. Because the total is fixed, the options compete; if one gets a bigger share, the others shrink.

Jev has no loop. It reads the text once. The options you provide become the entire vocabulary of that question: five labels instead of a hundred thousand tokens, and just once. Confidence isn’t a separate feeling, it’s the shape of the distribution: piled up in one place means high, spread out means low.

Left, an LLM: one pass per token. Right, Jev: the state is read once, all questions in the same pass.

Why it’s cheap

There are no output tokens, so output is free. What you pay for is input. Asking 13 questions in one call came out about 12.2× cheaper than asking them one by one, because the text travels once instead of 13 times. On Vercel AI Gateway the price is 0.042 dollars per million input tokens.

3. Its relative from school: the classifier

If you studied computer science or AI, you know its relative: the classifier. Spam or not, a review positive or negative. The BERT models on HuggingFace work like this. First they learn general language from a huge amount of text (pre-training), then they are adjusted on labelled data for specific classes (fine-tuning).

BERT is an encoder: it reads the whole sentence at once. Self-attention in one line: each word decides for itself how much to look at every other word in the sentence when working out its meaning. At the end, each class gets a probability.

What BERT is for: sentence and word classification; pre-training first, then task-specific fine-tuning. (Animation labels in Turkish.)

Jev looks like the frontier version of that family. The difference: in a classic classifier the classes are fixed during training. With Jev you write the classes at call time, and you don’t train anything.

One confusion to close early: Jev is not an embedding model. An embedding gives you a list of numbers (a vector) and you compute similarity yourself. Jev gives you a decision.

The transformer and self-attention, the frontier version where you supply the labels, and the “Jev is an embedding model” mistake.

4. What gets rewarded: RLHF, RLVR, RLCD

A model first learns by reading the internet. In a second stage it is steered with reinforcement learning: the model does something, gets a reward if it is good, and shifts towards whatever earns rewards. Like training a dog. Whatever you reward, the model becomes. The only difference between the three methods below is what the treat is given for.

Reinforcement learning: whatever you reward, the model becomes. (Animation labels in Turkish.)

RLHF: human preference

Reinforcement Learning from Human Feedback. Three stages: the model learns by reading, people rank answers and a reward model is trained on those rankings, then the main model goes into a loop scored by the reward model. The human is no longer in the loop; the reward model scores in their place.

The result: polite, fluent, persuasive answers. It is also why chat models always sound confident. They were rewarded for being liked, and confident answers get liked more. Being liked is not being right; on “how sure am I, in percent” they are not honest.

RLHF in three stages: pre-training, reward model, reward loop.

RLVR: an answer key

Reinforcement Learning with Verifiable Rewards. When the answer can be checked exactly, the reward is easy: right or wrong. Maths, code. Reasoning models are trained this way; they are accurate but slow and expensive because they think at length. But “should this sentence be cut in the edit?” has no answer key.

RLVR: the reward goes to a verifiable result; maths and code.

RLCD: an honest percentage

The method TypeSafe uses for Jev is called Reinforcement Learning for Calibrated Decisions. Here the reward goes to whether the stated probability matches reality. If the things it calls 70% come true 7 times out of 10, reward; if it says 90% and gets half right, penalty. This is called calibration: what you call 20% should really happen 1 time in 5.

The practical value of a calibrated percentage is that you can set thresholds: above 85% auto-approve, below 20% auto-reject, everything in between goes to a human. If the percentage isn’t honest, those thresholds mean nothing.

A scale: 80 percent on one pan, eight correct out of ten on the other
Calibration: what you call 80% should come true 8 times in 10. Generated illustration.
RLCD: the reward goes to whether your percentage holds up.
Same model, different treat: human preference, answer key, honest percentage.

5. “It can’t hallucinate”?

That sentence is half true. The official claim is only about type errors: Jev can’t make a type error. The schema is defined up front and the answer is always one of the labels you gave it. Ask an LLM for JSON and sometimes it comes back broken; that doesn’t happen with Jev.

But it can give a wrong answer that is still valid. The label is valid, the decision is wrong. The structure doesn’t break, the judgement does. So the real question is: how sure is it when it’s wrong? I measured that below.

TypeSafe’s own “jaggedness” doc is open about what it can’t do:

  • it doesn’t count
  • it doesn’t do arithmetic
  • it doesn’t compare dates
  • it reads text literally
  • performance drops with a large state
  • it is open to adversarial text written to fool it

My rule: maths in code, judgement in Jev.

6. Kıyafet Bul: the catalogue in code, the decision in Jev

Kıyafet Bul project cover
Code: github.com/selmakcby/jev-kiyafet-bul

I left theory and built a product. I pulled 561 outerwear and knitwear products from Trendyol once (Apify, 63 cents); they sit in a file. Kıyafet Bul (“find an outfit”) isn’t a search engine. You write what you want as a normal sentence and the rest goes through this pipeline. Code: github.com/selmakcby/jev-kiyafet-bul

  1. 1youyou write“a brown coat for the office, up to 2500 lira”
  2. 2Jevextracts filterstype, colour, length, belted, hooded: all in one call.
  3. 3codenarrows downfrom 561 products to at most 60 candidates. Price and size come from code, because Jev doesn’t read numbers.
  4. 4Jevscores candidatesa “does it match the request” score and confidence per product; the reason when it doesn’t.
  5. 5coderankssorts by score × confidence and lays out the results.
Kıyafet Bul screen: a search for a brown coat, the filters Jev extracted, and the five-step panel on the right
Left: how Jev understood my sentence. Right: the five steps. The budget and size boxes say KODDAN, “from code”.
Kıyafet Bul screen: a search for a long beige trench coat and the Jev panel
Second search: “long beige trench, belted, size M, under 3000 TL”. Size and budget again come from code.

Type “mayo” (swimsuit) and Jev says “not in this catalogue”, and none of the remaining steps run. That is the cleanest example of a smart if: it can say “no” to something that isn’t in the option list.

Kıyafet Bul searching for a swimsuit: scope out of catalogue, steps 3 to 5 skipped
“mayo” → out of catalogue 99%. Narrowing, scoring and ranking skipped.
a few secper search; up to 5 seconds when 60 candidates are scored
38ktokens, in the largest search
< 1 centa fraction of a cent per search
561products, pulled once

7. Where would you use this?

Clothing was just one example. The logic is the same everywhere: you need a decision, and you want to know how sure that decision is.

Customer supportWhich team should this message go to, is it urgent? Jev routes; a person or an LLM writes the reply.choicenoul
InboxInvoice, meeting, ad, does it need a reply? A label and a percentage for every email.choicenoul
Thesis and survey codingFor students: sorting 2,000 open-ended survey answers into categories takes weeks by hand, minutes with Jev.choice
Search and recommendationHow well does a candidate product or document fit? Jev scores, code does the ranking.score
Agent routerIs the question easy or hard? Send it to Sonnet or Opus? No point paying an expensive model to make that call.scorechoice
Auto-approval thresholdApplications, refunds, content: above 90% confidence it passes automatically, below that it goes to a person.noul
Live signalListening to speech and saying “cut”, or run-or-fight in a game. Where speed matters.choicenoul
In my own editing: the sentences from a shoot go through Jev, a cut list comes out, Claude Code cuts in CapCut. My current method is having a big model read the whole transcript.
Which one when: decision and percentage Jev, sentence LLM, numbers, arithmetic, dates code. Best is all three together.

8. I measured the calibration

This is what I really wanted to know: when Jev says 80%, is it really 80%? I took the answer key from my own video. I pulled 598 sentences from the raw footage; I had actually cut 392 of them in the edit. For each sentence I asked Jev “should this sentence be cut in the edit?” and compared its probability with my real decision.

598sentences, from my own footage
392actually cut in the edit
0.110ECE: 11 points apart on average between stated and real
64%accuracy (at a 0.5 threshold)
Calibration chart: Jev's stated probability against the real cut rate, blue line above the diagonal
Blue: all 598 sentences. Dashed line: perfect calibration. Bars below: sentences per bucket. (Chart labels in Turkish.)
bucketsentencesJev saidactually cut
0.3–0.4340.360.62
0.4–0.51660.450.51
0.5–0.62160.540.67
0.6–0.71360.640.75
0.7–0.8420.730.83
0.8–0.940.841.00

The curve rises overall: the higher Jev’s number, the more often I had cut, so it knows the order. But it is timid: it never says 95%, and the top bucket has only 4 sentences. The observed rates across the buckets are 0.62 / 0.51 / 0.67 / 0.75 / 0.83; in most buckets the real cut rate is higher than what Jev said.

598 requests took 23 seconds, about 1.3 cents in total.

Calibration: 598 sentences, ECE 0.110.

9. The live judge

Live judge project cover
Code: github.com/selmakcby/jev-canli-yargic

Then I hooked Jev up to myself. The Mac’s microphone listens in 3-second chunks, mlx-whisper transcribes, each sentence goes to Jev, and a panel labels it: teaching, filler, repetition, self-criticism, transition; plus a “cut” percentage. These are my own editing rules. Code: github.com/selmakcby/jev-canli-yargic

  1. 1codemicrophone3-second chunks
  2. 2codewhispertranscribes, about 0.6 s
  3. 3JevJevlabel + cut percentage, median about 290 ms
  4. 4codepanelputs the label and percentage on screen
Microphone, whisper, Jev, panel. The slow part isn’t Jev, it’s the audio side.

The slow part isn’t Jev, it’s the audio side. In rehearsal it labelled six sentences out of six correctly: “I think I explained that badly” got self-criticism 1.0; saying the same sentence a second time got repetition 0.98.

Live judge panel: self-criticism label and cut 72 percent
A frame from the real run: “Ben çıkartalım” (“let me cut that”) self-criticism, cut 72%. For this sentence whisper took 665 ms, Jev 350 ms.

10. Claude Code + Jev: a job-listing matcher

The last experiment is a job-listing matcher. I told Claude Code I was looking for listings that fit my LinkedIn profile. Claude Code pulled current LinkedIn listings with Apify, asked Jev about each one (does it fit me, what percentage; which role; which level) and wrote the result to Excel. Claude Code doesn’t know Jev, so I installed the TypeSafe skill from their docs; the skill works as its user manual. The keys were already in .env.

Claude Code doesn’t know Jev; once the TypeSafe skill from their docs is installed, it knows how to call it.

Out of 45 listings: 1 “apply for sure”, 16 “could apply”, 28 “pass”. The Jev side cost 0.23 cents. Claude Code fetched the data, Jev made the decisions, Claude Code wrote the Excel.

My LinkedIn profile, 45 listings via Apify, Jev: “does this listing fit my profile?”, code applies thresholds, Excel. 45 listings in 1.84 seconds, 0.23 cents; best match AI Engineer, Rotterdam, 0.82.

11. Bonus: Minecraft

Minecraft bot project cover
Code: github.com/selmakcby/jev-minecraft-bot

Everyone has done this; I tried it for a minute. Planning and movement live in code, reflexes in Jev. Every 600 milliseconds the code writes the state: health, hunger, is it night, is the nearest enemy close, mid or far. Jev decides in one call: fight, flee, eat, gather wood, shelter, explore, wait; plus a danger probability. Code carries it out.

One decision: code writes the state, Jev answers three questions in one call (explore 0.69), code carries it out with a pathfinder. About 300 ms per decision.
Jev as a Minecraft character, in a forest
Jev in Minecraft. A generated illustration, not gameplay.

I give the options, Jev makes the decision, code does the doing. That has a consequence: if “gather iron” isn’t in the list, Jev won’t mine iron even when it sees it. The panel logged 2893 decisions, about 7 cents in total. Code: github.com/selmakcby/jev-minecraft-bot

Who gives the options, who decides, who executes: me, Jev, code. (Labels in Turkish.)
What the “shelter” decision means in code: dig 3 blocks down, cover the top with dirt.
Jev as a Minecraft character, by a lake
The same character, another frame. Also a generated illustration.
Minecraft figures: zombie, creeper, skeleton, apple, log, sword, torch
The figures from the board: zombie, creeper, skeleton, apple, log, sword, torch. Generated illustrations.

12. Sources

You can get a Jev key in two places: the TypeSafe console (there is a waitlist) or Vercel AI Gateway.

Where in your work is there a decision being made that you can’t write code for? That’s Jev’s spot. If you try it, tell me in the comments what came out.

Selma Kocabıyık — AI Engineer