Skip to content
Made with Jev

Jev confidence: when to act, ask or escalate

Jev always answers. It picks an option every time and reports how sure it is, and that number is what turns a model call into an if: act when it's high, ask when it's middling, hand off when it's low. This page covers how TypeSafe computes it, why a Noul doesn't have one, the thresholds TypeSafe's own cookbooks used, the lines builders drew, and how to set yours from the cost of a mistake. Source: TypeSafe confidence docs.

Updated 8 Oct 2026 · by Made with Jev

In short

  • Every Choice and Score answer carries confidence, from 0 to 1, computed from the answer's own probabilities: 1 when all of it is on one outcome, 0 when it is spread evenly. confidence
  • A Noul has no confidence field. Its probability already carries the uncertainty, and a value near 0.5 means unsure. Noul
  • Calibration is a property of many answers: of the answers given 0.8, about 80% should be right. It is not a guarantee about any one answer. System One
  • No threshold is a default. TypeSafe says to start conservative, test on your own data and set different lines for actions with different risks. Thresholds scale with risk
  • Run in shadow first: log what Jev would have done, label a sample, then choose the line.

What the number is

confidence is a statistic computed from probabilities, which every Choice and Score answer also returns, so you can always check it or replace it. TypeSafe publishes the formulas.

Choice. With n options and p_max the probability of the chosen one, confidence = (p_max − 1/n) / (1 − 1/n). It measures how far the top option sits above an even split. Only the top probability counts, so (0.6, 0.3, 0.1) and (0.6, 0.2, 0.2) both give 0.40. In TypeSafe's ticket example, returns gets 0.61 of three departments and billing 0.35, and TypeSafe prints a confidence of 0.42 (these rounded probabilities give 0.415).

Score. Levels are ordered, so probability on a neighbouring level costs less than probability on a distant one. With three levels, (0, 0.5, 0.5) gives 0.25, a model torn between two adjacent levels, while (0.5, 0, 0.5) gives 0.00, torn between opposite ends. TypeSafe's bug report scored (0, 0.57, 0.43), which gives 0.35.

TypeSafe calls this "one reasonable way to summarize a distribution, not the only one". Two alternatives it suggests trying: the top probability on its own, which reads as "how likely is this option" but means different things with 2 options and with 10, so set it per question; and the ratio of the top probability to the second, for decisions that come down to two candidates.

TypeSafe's formulas, in TypeScript
// Choice: how far the top option sits above an even split.
function choiceConfidence(p: number[]): number {
  const n = p.length;
  return (Math.max(...p) - 1 / n) / (1 - 1 / n);
}

// Score: spread around the most likely level, against an even spread.
function scoreConfidence(p: number[]): number {
  const n = p.length;
  const m = p.indexOf(Math.max(...p));
  const spread = p.reduce((s, pi, i) => s + pi * Math.abs(i - m), 0);
  const even = p.reduce((s, _, i) => s + Math.abs(i - (n - 1) / 2), 0) / n;
  return Math.max(0, 1 - spread / even);
}

// Noul: no field is returned; this puts it on the same scale.
const noulConfidence = (yes: number) => Math.abs(2 * yes - 1);

Sources: How confidence is calculated, Choice, Score.

Why a Noul has no confidence

A Noul answer is one probability over two outcomes, so the number is the answer and the certainty at once. If you want Nouls and Choices on the same scale, use the distance from 0.5: |2p − 1|. firehose-judge does exactly this so that one review rule covers all three types.

If you threshold a Noul directly, TypeSafe's advice is: use 0.5 when yes and no are equally easy to act on. Raise the line when acting on a false yes is expensive, such as paging someone or issuing a refund. Lower it when missing a true yes is expensive, such as a safety issue. Send the middle to a person. Its support-routing example uses YES = 0.8 and NO = 0.2, with everything between going to review. If reviewers see too many cases, narrow the gap. If too many wrong routes get through, widen it.

Sources: Reading a Noul, firehose-judge src/jev.ts.

Calibrated across many answers, not for one

TypeSafe trains Jev with what it calls reinforcement learning for calibrated decisions. The target: across many answers, outcomes given 0.2 happen about 20% of the time and outcomes given 0.8 about 80%. That is a statement about groups. A confidence of 1.0 on a Score "describes the model's answer, not a guarantee that the answer is correct", and a reworded description that raises confidence hasn't been shown to be better until you check it against labels.

Calibration is also measured on someone's data, not yours. Bouncer publishes a per-question table from its own fixtures. Its destructive question was right 92% of the time over 26 answers, but its 0.8–0.9 bin held 2 answers and got 1 right. Small bins swing. Bouncer's bar for turning a question on is 85% accuracy among answers at 0.8 confidence or higher, compared on the exact ratio, because 11 of 13 is 84.6% and would round to a passing-looking 85%.

Sources: AI primer, System One, Reading a Score, Bouncer README.

Bouncer's figures are the authors' run, as of Oct 2026; the README table updates with each calibration run.

What a 0.98 means for a judge is on Jev as a judge.

Act, confirm or route

TypeSafe's starting pattern splits confidence into three ranges. High: act automatically. Medium: ask the user to confirm, flag it for review, or gather more information first. Low: don't act; route it to a person, ask for clarification, or fall back to another system. A single system usually needs more than one line, because a wrong read costs more for some actions than for others.

Thresholds by risk, after TypeSafe's banking example
const action = response.answers.action;

if (action.confidence < 0.5) {
  // Genuinely unsure: don't guess.
  return routeToHuman(message);
}
if (action.choice === "check_balance") {
  // Low stakes: the wrong screen is recoverable.
  return showBalance(accountId);
}
if (action.choice === "approve_transfer") {
  // High stakes: even a confident read is confirmed.
  return action.confidence > 0.9
    ? confirmThenExecute(accountId)
    : askUserToConfirm(accountId);
}

In TypeSafe's version, the transfer is confirmed even above 0.9; the threshold only decides how it is confirmed. Its confidence-routing pattern uses the same idea with a 0.6 floor and 0.85 for a transfer. And not every answer needs a gate. If all you need is the best option, take the top one; TypeSafe's agent-skill page warns against "using confidence thresholds everywhere".

Sources: Thresholds scale with risk, Confidence-gated routing, agent skill: common issues.

The lines TypeSafe's cookbooks used

Each of these numbers was chosen for one dataset and one model version. They are examples, not defaults.

CookbookQuestionThe lineWhat it didRun
Classification using confidenceOne Choice over 75 SIC industry groups, 60 filingsConfidence 0.9Split them in half: the confident half 90% right, the rest 40%. Reporting the rest one level up raised them to 70%, with no second calljev-1.12, 2026-08-12
Self-consistency: choices8 Choice questions, 15 repeated runsTop probability 0.60; below → "uncertain", human reviewAgreement across repeated runs rose from 90.8% to 99.2%; 74.2% of answers automaticjev-latest (returned jev-1.13.0), sampled 2026-09-11
Self-consistency: nouls14 Noul questions about one insurance claim, 15 runs0.30–0.70 → "uncertain"Mean per-question standard deviation 0.0102; one answer (covered) ranged 0.43–0.53 across runs, crossing 0.5jev-latest (returned jev-1.13.0), sampled 2026-09-11
Classifying RAG passages4 Nouls per passageInjection > 0.70 drop; contradiction > 0.70 conflict; relevance < 0.45 drop; evidence > 0.55 includeRemoved the planted injection that cosine similarity ranked first; TypeSafe calls the numbers a starting point for its corpusjev-1.12, 2026-08-27
Date extractionOne Choice per date partA date's confidence is its weakest part's; under 0.60 → a personDates code can't assemble also go to a personjev-1.12

The lines builders drew

These are each author's own settings, from their repositories or posts. We haven't re-run them.

BuildThe lineWhat happensSource
firehose-judgeReview if confidence < 0.4 (Choice), < 0.1 (Score), < 0.16 (Noul, as |2p − 1|)The post goes to the "needs a human" lane. "Scores legitimately land between two adjacent levels, so they get a much lower bar."src/jev.ts
jev-antispam-botSPAM_THRESHOLD 0.81 (default)Deletes only matches above it, and leaves messages alone if Jev is unavailableREADME
winnowHide a block only if P(needed) < 0.1; 0.5 is the keep lineAnything in between is kept. 0.1 is "the bin that came back clean on hand-labeled replay"README
BouncerA 0.40–0.60 uncertainty band asks yougit stash clear was caught at 0.42 by the band, not by its own rule, which needed 0.70. The authors call the verdict right and "the reason was an accident"README
fraud-detection-jev-kimiUnder 95% confidence → Kimi K331 of 100 emails escalated; the pipeline got 96 right; Jev's share of the bill was $0.003Hassan, on X
semdecide0.700 in the README example0.860 → TRUE, and the exit code follows the comparison, so it composes with &&README
jev-as-judge0.2 / 0.8 verdict lines, 0.6 confidenceThe author calls all three "illustrative and uncalibrated"README
stanford-rag-jev-loopPassages above 0.6 reach the writerAs described by N01ennn; we haven't found the group's own write-upx.com/N01ennn

Set the line from the cost of a mistake

A threshold is a trade between two costs: acting on a wrong answer, and sending a case to someone who didn't need to see it. For a Noul you act on as a yes, the average cost of acting is (1 − p) × the cost of a wrong yes; the cost of review is whatever review costs. Acting is cheaper when p > 1 − review cost ÷ cost of a wrong yes. The same reasoning on the other side gives the line below which a "no" is safe: p < review cost ÷ cost of a missed yes. Between the two lines, a person decides.

With hypothetical costs, chosen only to show the arithmetic: if a wrong automatic refund costs as much as 20 reviews and a missed refund request as much as 4, act on yes above 0.95, treat it as no below 0.25, and review in between.

Our arithmetic; the costs are hypothetical.

This only works if p means what it says on your data, so check it against labels first. For a Choice or Score, confidence is not the probability of being right. Replace (1 − p) with the error rate you measured at that confidence in your labelled sample. jeval automates this: it reads your logged answers and human corrections, plots the expected cost for every candidate line, and writes out the one to deploy. Its example report recommends moving a line from 0.60 to 0.75, but that report is built from a synthetic demo log, not a real deployment.

Sources: Noul thresholds, jeval README.

Tune it in shadow

  1. Log, don't act. Run Jev beside the current path and record every answer, with nothing changed. winnow has WINNOW_MODE=shadow ("run it for a week without trusting it"), Bouncer ships watching and logging, and 0xCodila ran a Grok Bot router in shadow and read the logs before switching it on.
  2. Label a sample spread across the range. winnow's replay tool draws 100 blocks stratified by the judge's probability, so the middle isn't under-sampled.
  3. Plot confidence against accuracy, per question. That is TypeSafe's own instruction for testing thresholds. firehose-judge serves the review share and human agreement at 21 thresholds, so you can see what a stricter bar costs in human hours.
  4. Set one line per action, from the cost of a mistake (above).
  5. Pin the model version once the lines are set. The response's model field names the version that answered (as of Oct 2026, jev-1.13.0).
  6. Keep questions and thresholds in one file, where a reviewer can find them.
  7. Re-tune when something changes: the model, a question's wording, or the classifier behind the same API. jev-antispam-bot's README says to retune SPAM_THRESHOLD after switching.
  8. Watch both directions. TypeSafe's checklist: a line set too high causes false negatives, one set too low causes false positives, and sometimes the question needs to be more specific instead.

Sources: How to build: route on uncertainty, agent skill: common issues, winnow, firehose-judge, 0xCodila, models.

A lane for the unsure ones

Each of these acts on the confident answers and hands the rest to a person, a larger model or a do-nothing default. Hassan sent the emails under 95% confidence to Kimi K3, and firehose-judge puts posts Jev won't commit on in a lane marked "needs a human".

GitHubSecurity and abuse

firehose-judge

The live Bluesky firehose, judged post by post, with a lane for humans.

Leo Mata

0~$0.00003

GitHubSocial feeds

LinkedIn NoSlop

A Chrome extension that blurs low-value LinkedIn posts.

sushrutb17

0

GitHubSecurity and abuse

Jev Anti-Spam Bot

A Telegram bot that deletes only the spam Jev is sure about.

Nikita Kolmogorov

16

Hassan

@nutlope

Jev + Kimi K3 for fraud detection! TLDR: Jev classified 100 emails in 1.42 seconds, then I routed the uncertain cases to Kimi K3. The full pipeline got 96/100 correct for only ~$0.07. Video is not sped up, check out the live run! Here was my process: I gave Jev 100 emails to classify (a mix of 50 legit & 50 fraudelent emails). It classified all of them in 1.42 seconds. An underrated feature about Jev is it will give you the confidence score for a classification, so I routed any prediction under 95% confidence to Kimi K3 to be fully sure. 31 emails fell below that threshold. After routing those to Kimi K3, the combined pipeline reached 96% accuracy. The full run took 16 seconds & ~$0.07 in inference costs: - $0.068 from Kimi K3 on @togethercompute - $0.003 (1/3 of a cent) from Jev on @typesafeai. I think this is a really interesting pattern: use a fast specialized model like Jev for the narrow task, then route the uncertain cases to a larger LLM. I feel like this kind of approach could be a game changer for use cases like fraud or anything realtime. You can use the speed & low cost of Jev while having a larger LLM as a fallback to ensure high accuracy.

XSecurity and abuse

Pick

Fraud detection with Jev and Kimi K3

Emails
100 in 1.42 s
Correct
96/100
Cost
~$0.07 pipeline ($0.003 Jev, $0.068 Kimi K3)

GitHubSearch

SemDecide

grep for meaning and jq for judgment, with a Unix exit code.

Sharvil Saxena

76

GitHubSecurity and abuse

jevmod

Moderation with a probability per category and thresholds you set.

Omar Hernandez

1

Akshay 🚀

@akshay_pachaar

GitHubBenchmarks and evals

jev-as-judge

Traces
10 frozen
Questions
4 in one request

NO1ennn

@N01ennn

this is pure f*cking treasure A Stanford AI research group finally drew the perfect RAG system: retrieval, Jev and agents in one loop, and it fixes the 3 things that break every RAG app: > the LLM reads 20 passages when only 3 matter > it answers questions your docs can't answer > it trusts whatever text it retrieves here's how it runs: > a lead agent sends the query > hybrid search pulls the top 20 candidates, dense + keyword > ONE Jev request scores all 20 + 2 gates: answerable? injection? > only passages above 0.6 reach the writer agent > a second Jev call checks every claim against its source > grounded answer, with citations and when the docs don't have it: > answerable fails, the writer never runs > a researcher agent rewrites the query and retries once > still nothing? "not in the docs". zero tokens spent on a guess retrieval casts the net. Jev decides what's real. agents do the work save this before you build your next RAG

XSearch

A RAG loop where Jev decides what is real

Candidates scored per request
20
Passage threshold
0.6

Run it in shadow first

Judge everything, log it, change nothing, then read the logs. winnow's README calls its shadow mode a week of running it without trusting it, and Bouncer ships watching and logging rather than blocking.

GitHubContext and memory

winnow

A context sieve for Claude Code: Jev judges each tool result before it lands.

Ghaleb Dweikat

103

GitHubSecurity and abuse

Bouncer

A judgment layer that checks every Claude Code tool call against your policy.

Clownware

1~$0.04/day

codila

@0xCodila

Jev + GrokBot is the best AI agent system I’ve built in my life It just made my setup CHEAPER and FASTER than what 95% of people are running... setup takes literally 7 minutes: prompt → GrokBot → Jev decision → GrokBot execution → result step 1 → open @typesafeai , create API key (keep it off chat paste) step 2 → tell Grok Bot: store TYPESAFE_API_KEY in the secure field step 3 → prompt Grok Bot: install typesafe-sdk on Agent Computer + smoke system_one (Choice) step 4 → tell Grok Bot: build the usage lab (router, dry-run, config, logs) - or clone Github below step 5 → add skill jev-usage-router: before browser / research / retry / extra bot → call the router, honor action step 6 → stay shadow first, read logs, then active when you trust it - kill switch: bypass jev or enabled: false step 7 → flip active: GrokBot obeys route - Jev decides - GrokBot executes - humans control irreversible actions the result: Jev + GrokBot the best and fastest agent running directly on your computer rn, I’ve already tested it on routine tasks - and the results are genuinely incredible You can come up with endless ways to use Jev + GrokBot - but the most important thing is to install it as soon as possible Copy this 2028 setup, explore my repo below - then read the full Jev deep dive ↓

XRouting and model choice

A usage router for Grok Bot

Tools that find the line

Calibration checks and threshold pickers that work on your labelled answers. jeval turns what a mistake costs into the point where a person should take over.

GitHubBenchmarks and evals

jeval

Works out what a Jev confidence score is actually worth.

hope

22

SkillBenchmarks and evals

tenbin

Split a judgment into Choice, Score and Noul, lint it, then measure it.

Shingo Imota

4

GitHubBenchmarks and evals

jev-calibrate

Tune your Jev questions against your own labels, then confirm on held-out data.

smkrv

32

GitHubBenchmarks and evals

jev-align

Turns human labels into calibrated Jev classifiers with GEPA.

Sutro

305

GitHubBenchmarks and evals

jev-evaluation

Nine experiments and 28 predictions, all fixed before any data.

Will Kelly

0123,805

GitHubBenchmarks and evals

daf-jev

A Python toolkit: question builders, confidence gates and calibration.

Daniel Ari Friedman

6

Reading

Ricky Grannis-Vu

@RickyGrannisVu

Jev gives you a probability. You set two lines on it: above the top one it's a yes, below the bottom one it's a no, anything in the gap runs a third branch. It's just a normal branch, so you can handle uncertainty flexibly.

Guidex.com

Jev’s probability thresholds

Jev outputs a probability, two thresholds split the answer into yes, no and a third branch between them.

Akshay 🚀

@akshay_pachaar

Guidex.com

Build a Jev Judge

Akshay Pachaar’s worked example: evaluate a refund-support agent with Jev instead of a generative judge, and record the verdicts in Opik. Separates the judge from the evaluation system, and is explicit that a 0.98 is a probability about one proposition, not a percentage of the answer.

Guideblog.lepine.pro

Let’s look at Jev

Jean-François Lépine on Choice, Noul and Score, calibrated confidence and batching questions, with a complete Python project.

Akshay 🚀

@akshay_pachaar

LLMs vs. Jev, clearly explained! TL;DR The key difference is not that Jev generates faster. Jev does not generate text at all. A traditional LLM receives context and produces an answer one token at a time. Even when the output is a small JSON object, every token depends on those generated before it. Jev receives the same context but evaluates predefined decisions directly. When those decisions are independent, it can evaluate all of them in parallel. Consider an agent handling a failed deployment. It may need to determine: → Whether the incident is urgent → Which team should handle it → Whether the proposed command is risky → Whether the task is complete An LLM generates a response containing these answers sequentially. The application then parses and validates it. With Jev, you define the questions and expected answer types upfront. It evaluates them together and returns typed answers with probabilities. Jev supports three decision primitives: 1. **Choice** selects from known options, such as engineering, billing, or sales. 2. **Score** places the input on an ordered scale, such as low, medium, or high risk. 3. **Noul** evaluates a yes-or-no condition and returns the probability that it is true. The probabilities matter as much as the selected answers. If engineering receives 91% probability and billing receives 9%, automatic routing may be reasonable. If the probabilities are 52% and 48%, the system can escalate, gather more context, or call a stronger model. This keeps control inside ordinary software. Code owns the thresholds and consequences. Jev supplies the semantic judgment that a normal `if` statement cannot derive from unstructured text. It works best when the possible answers are known, the decision depends on meaning, and a careful person could judge the input quickly. It is not designed for writing, summarization, code generation, arithmetic, or decisions requiring several dependent reasoning steps. Independent questions can run in parallel, but decisions that depend on earlier results must remain sequential. Jev also cannot return an option outside the declared schema, but it can still select the wrong valid option. Type safety prevents malformed outputs, not incorrect judgments. The clean mental model is this: LLMs generate new language when the answer space is open. Jev evaluates known paths when the answer space is bounded. I wrote the full breakdown explaining Jev and where it fits. The article is quoted below.

Guidex.com

LLMs vs. Jev, clearly explained

Akshay Pachaar: Jev does not generate text at all. It answers Choice, Score and Noul questions in parallel, and your code owns the thresholds.

Guideyoutube.com/@GaryExplains

Jev: fast and cheap, but there is a caveat

Gary Explains on the catch: Jev understands natural language but it is not an LLM. It answers with structured values and a confidence level.

Where to go next

Confidence is only as good as the question behind it. If answers come back unsure across the board, check the wording first: Jev question types. The same gates, applied to labels, are on Jev for classification. Routing each turn to a cheaper or larger model on Jev's read is on the Jev router page. Your first call, and where the confidence appears in it, are on how to use Jev.

Common questions

What is confidence in Jev?
A number from 0 to 1 on every Choice and Score answer, computed from the answer's own probabilities: 1 when all the probability is on one option or level, 0 when it is spread evenly. TypeSafe publishes the formula.
Why doesn't a Noul answer have a confidence?
A Noul has two outcomes, so its probability of yes already says how sure it is: near 0 or 1 is sure, near 0.5 is not. For a comparable number, use |2p − 1|.
Is confidence the probability that the answer is right?
No. It summarises how the probabilities are spread. Jev is trained for calibration, which holds across many answers, not for any single one. Measure accuracy at each confidence level on your own labelled data.
What confidence threshold should I use?
There is no default. TypeSafe's examples use a 0.5 or 0.6 floor, with higher lines of 0.85 to 0.9 for risky actions, and its cookbooks used 0.9 and 0.60 on their own datasets. Set yours from the cost of a mistake and test it in shadow first.
Should I threshold confidence or the top probability?
Either can work. TypeSafe suggests also trying the top probability, set per question because its meaning depends on the number of options, and the ratio of the top two probabilities. If you only need the best option, take it without a threshold.
What is shadow mode?
Running Jev beside your current system and logging what it would have done, without acting on it. You label a sample of those answers, then pick the threshold from what you measured.

Made with Jev is independent and not affiliated with TypeSafe AI. Every figure on this page is the one its author published, linked to where it can be checked.