Jev question types: Choice, Score and Noul
Every Jev call is a set of questions, and each one is a Choice, a Score or a Noul. The type decides what comes back: one option from your list, a position on levels you wrote, or the probability that the answer is yes. This is the reference for all three, using TypeSafe's own examples, plus the habits that make an answer mean what you meant by the question. Source: TypeSafe primitives.
Updated 8 Oct 2026 · by Made with Jev
In short
- Choice picks one option from up to 255 and returns the choice, a probability for every option and a confidence. Choice
- Score places the text on 2 to 10 ordered levels you describe. The score can fall between two levels, and it comes with a legend, probabilities and a confidence. Score
- Noul answers one yes/no question with the probability of yes, and nothing else: there is no confidence field. Noul
- The model never sees your question ID, so write the whole question in instructions. Option names and their descriptions do reach the model. Choice
- Put every question you might need in one request. They are answered independently and in parallel, and an unneeded answer costs only its tokens. Speculative fan-out
Which type to use
Pick the type whose answer your code can act on directly. A Choice between refund, rebook and information maps onto three code paths, a Score of frustration onto a threshold, and a Noul onto an if.
| Type | Asks | criteria | Returns | Limit |
|---|---|---|---|---|
| Choice | Which of these options? | A map of option → description (a description can be null) | choice, probabilities, confidence | Up to 255 options |
| Score | Which level? | An ordered array of level descriptions, low to high | score, legend, probabilities, confidence | At least 2 levels, up to 10 |
| Noul | Is this true? | Optional {true, false} descriptions | noul (0 to 1) | One condition per question |
Use a Noul for a yes/no judgment and a Score for a position on a spectrum. TypeSafe's docs ask "Is the candidate strong in Python?" of four CVs. As a Noul the answers run 0.03, 0.14, 0.81, 0.92. As a Score with four described levels they land at 0.0, 1.0, 2.05 and 2.89, each on or near a level someone wrote. A Noul of 0.5 means yes and no are equally likely. It does not mean "medium".
Source: Noul, Choose a question type.
Choice
A Choice has a type, instructions (the question) and criteria, which is a map from option name to description. The model gets both the names and the descriptions, so write descriptions that tell the options apart. Use null when a name says it all. Give the full list rather than a shortlist; each extra option costs a few tokens. Add an other or none of the above option when the list might not cover every input.
{
"state": "I was charged twice for my subscription this month.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or plan questions"
}
}
}
}{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 0.88, "technical": 0.12, "sales": 0.0 },
"confidence": 0.81
}
}
}The 0.81 is TypeSafe's figure. The published formula gives 0.82 on these rounded probabilities.
Past 255 options, or for a deep taxonomy, chain Choices level by level and walk the tree in code. TypeSafe's hierarchical classification cookbook keeps the best few paths at each level rather than committing to one. For a flat list that is too long, score it in windows and then choose within the winner.
Sources: Choice, hierarchical classification cookbook.
Score
A Score's criteria is an ordered array of level descriptions. A level's number is its position in the array, starting at 0. The model sees only the descriptions and judges each level against the state separately, so it never sees the numbers. The score that comes back is each level number times its probability, summed. It can land between two levels. legend maps each number back to its description, and the Python SDK keys legend and probabilities by integer rather than by string.
{
"frustration": {
"type": "score",
"score": 1.05,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.0, "1": 0.95, "2": 0.05 },
"confidence": 0.92
}
}Here the score is 0 × 0.0 + 1 × 0.95 + 2 × 0.05 = 1.05. The same score can hide different answers: 1.0 can mean all the weight is on level 1, or half on level 0 and half on level 2. Read probabilities and confidence alongside it. A score is a position, not a measurement. Threshold it, sort by it, or round it to a level, but don't interpolate an exact quantity between two levels; TypeSafe says jev-1.13's levels are weak in numerical calibration. To combine Scores of different lengths, first divide each by its top level number (the number of levels minus one).
Sources: Score, jaggedness: maths using score.
Noul
A Noul has instructions and optional criteria with true and false descriptions of what a yes and a no mean. The answer is one number, the probability of yes. There is no separate confidence because, with only two outcomes, that number already says everything: near 1 is a clear yes, near 0 a clear no, and near 0.5 means the model is unsure.
{
"is_urgent": { "type": "noul", "noul": 0.95 }
}| Message | noul |
|---|---|
| Thanks, that fixed it! | 0.02 |
| I need this sorted today, whatever it takes. | 0.26 |
| Are you a bot? | 0.40 |
| I have asked three times now. Can I please just talk to a real person? | 0.99 |
The middle two are where a threshold in your code decides. How to set it is on Jev confidence.
Source: Noul.
Instructions and criteria
instructions is the question. criteria is the set of possible answers: options for a Choice, levels for a Score, and an optional definition of yes and no for a Noul. Treat the criteria as an extension of the instruction. When they pull in different directions, for example a Noul whose true describes a no, jev-1.13 does worse.
Both fields take a string, an object or an array. Start with strings. Switch to an object when one string is doing several jobs:
- Two options keep getting confused. Give each one fields for what it covers, what belongs to a neighbouring option instead, and a few examples. Use the same field names on every option. The names are yours; none are reserved, and the model reads them, so keep them short.
- Part of the question comes from your code. Put the database row in its own field and refer to it by name in backticks, rather than splicing it into a sentence. TypeSafe's Noul example compares one CV against several candidate records this way, with one question per record and all of them in a single request.
- The state has several parts. Point the question at one with a backticked path such as
support.tickets[0].message.
Examples steer the model, and higher confidence doesn't prove a description is better. On TypeSafe's Safari bug report, adding a matching example to each Score level moved the answer from 1.43 at confidence 0.35 to 1.03 at 0.96. An example unrelated to browsers left it at 1.43 and 0.35. Choose examples whose correct level you already know, then test on different inputs. One builder measured this: winnow's author scored three wordings of the same question against 97 blind hand labels. The structured version (what / not for / examples) and the one-line version were both clean below 0.1. A stricter one-line wording was overconfident: 23% of its lowest bin turned out to be needed.
Sources: Advanced: structure, Score: structured levels, jaggedness: contradictory, winnow README.
Writing Score levels
- Describe situations, not degrees. "Broken or degraded feature, but workaround exists" gives the model something to match against; "moderately severe" doesn't.
- Numbers don't help. Each level is judged on its own, so "worse than the previous level" means nothing, and neither do numbers in the descriptions. In TypeSafe's test, number-only levels split a cosmetic bug between 0 and 1. With descriptive levels it scored 0.0 at confidence 1.0.
- One dimension per Score. "Punctual and smart and experienced" measures three things. Split it into three Scores and weight them in code.
- Give a rare extreme its own level when you would act on it differently, such as "abusive or threatening" above "very angry".
- Use as many levels as you can describe distinctly, up to 10. Three is fine.
- If there is no in-between, use a Choice or several Nouls.
firehose-judge's questions.json shows both styles side by side. Its bait Score describes situations, from "Just talking, no hook" to "Full ragebait or outrage farming". Its sentiment Score uses degree words, from "Very negative" to "Very positive". By TypeSafe's advice, the second is the one to rewrite first if its answers drift.
Sources: Score: writing good levels, firehose-judge questions.json.
How to write a good question
- One snap judgment per question. TypeSafe's test is whether a knowledgeable person could answer it in a second. "Does this message convey urgency?" passes. "Analyze this message and determine the best course of action" doesn't.
- Split broad questions. TypeSafe's docs replace "Is
messagespam?" with six Nouls: does it ask for a password, does it claim an unexpected reward, does it pressure the reader to act quickly, does the sender's name conflict with the email domain, does the link's domain conflict with it, and does the link text hide where it goes. Each answer can then be inspected, weighted and tuned. - Write the whole question in instructions. The ID
safe_to_publishtells the model nothing, because the model never sees it. - For a Noul, one condition and a clear boundary. "Is the customer angry and asking for a refund?" is two Nouls. Phrase it so that a high value means yes: "Does the message contain personal data?", not "Is the message free of personal data?". Words like "any" remove the middle ground. A statement ("The customer is requesting a refund") works as well as a question. Try both, and try with and without criteria.
- Say exactly what you mean. jev-1.13 reads scoping words, negations and implied conditions at face value. When you catch yourself explaining what you meant after a wrong answer, that explanation is the missing half of the instruction.
- Send only what the question needs. Unrelated text in the state costs accuracy.
- Keep questions and thresholds in one file. These are what a reviewer needs to read. TypeSafe's agent-skill page adds: "Agents aren't great at writing questions, so expect to edit collaboratively with them."
Sources: primitives, How to build with TypeSafe, Writing a Noul question, literal reading, agent skill.
A first call, line by line, is on how to use Jev.
Ask the questions you might need
Questions in one request share the state and are evaluated independently and in parallel. One answer is never context for another, and adding a question costs its tokens and barely changes the response time. So ask everything your code might need, including questions that only matter for some inputs, and ignore the answers you don't use. TypeSafe calls this speculative fan-out. In its parallel-questions cookbook, 13 questions in one call were 12.2× cheaper and 10.0× faster than 13 separate calls, with the same answers.
Make a second request only when your code cannot build it until it has the first answer. That happens when the answer decides what to fetch, what the state is made of, or which options the next question offers.
Sources: Speculative fan-out, parallel questions cookbook, When one question depends on another.
Common mistakes
| Mistake | What happens | Instead |
|---|---|---|
| Relying on the question ID | The model never sees it | Put the whole question in instructions |
| Two conditions in one Noul | The value means less | One Noul each; combine in code |
| A Noul for a degree ("strong in Python") | 0.5 reads as "medium" but means "unsure" | A Score with described levels |
| Inverted wording, or true that describes a no | Worse answers; code reads it backwards | High value = yes; criteria extend the instruction. See contradictory instructions |
| Levels written as degrees or numbers | Probability splits between levels | Describe situations |
| Several things in one Score | Confidence drops; the score means less | One Score per dimension |
| Reading a Score as a quantity | Interpolation between levels is unreliable | Threshold, sort or round it. See maths and numbers |
| Implied conditions | Answered literally | State the exact condition. See literal reading |
| A property of a property | Accuracy drops with each hop | Point at the field by path. See indirection |
| The whole document as state | Distractors cost accuracy | Filter first. See large state |
| Trusting option order | jev-1.13 can lean to the first option | Reorder and check. See option order |
| One question per call | Up to 12.2× the cost and 10.0× the time (TypeSafe's cookbook) | One request per state |
Sources: Jev 1.13 jaggedness, primitives, Score, Noul.
Question sets you can read
Each of these keeps its questions in one file you can open: firehose-judge's questions.json mixes two Choices, two Scores and four Nouls about every Bluesky post, and Bouncer's policy is a YAML file of plain-English questions and thresholds.

GitHubSecurity and abuse
firehose-judge
The live Bluesky firehose, judged post by post, with a lane for humans.
Leo Mata
0~$0.00003Akshay 🚀
@akshay_pachaar



GitHubSecurity and abuse
Jev Anti-Spam Bot
A Telegram bot that deletes only the spam Jev is sure about.
Nikita Kolmogorov
16
GitHubSecurity and abuse
Bouncer
A judgment layer that checks every Claude Code tool call against your policy.
Clownware
1~$0.04/day
GitHubContext and memory
winnow
A context sieve for Claude Code: Jev judges each tool result before it lands.
Ghaleb Dweikat
103Many questions in one request
Questions in a request are answered independently and in parallel, so a long rubric costs little more than one question. Rob Hallam asks 61 about a draft post for $0.0004, and Movez asked 14 yes-or-no questions of each of 100,000 posts for $0.67.
Rob Hallam
@robj3d3
Jev + SuperX = virality solved ✅ Every post gets 61 questions in ~1s for $0.0004 🤯 > fitted on 9,481 real posts from 207 creators > picks the viral post 2 in 3 times > never rewards reply bait So: write, score, rewrite, stop when it peaks. Free, no signup. try it below ↓
Movez
@0xMovez
I just built a Jev X Viral Post Analyser. 100,000 viral X posts. 20.4 seconds. $0.67. Claude Opus 5, same corpus, same clock, got through 214 posts and spent $0.98. per post that is ~680x cheaper the full Opus pass would have run $458. viral analysis is the perfect Jev job. • it is not writing, it is 14 yes/no calls per post: > does the hook open a loop, > is there a number in the first line, > is the proof real or claimed. classification, not prose. • what it found: 1,220 posts broke into the top 1%. baseline 1.22%. > superlative claim - 2.34% viral. 1.92x baseline > contrarian take - 1.59%. 1.31x > launch / tool drop - 1.46%. 1.19x and numbered lists, the thing everyone writes: 0.55%. below baseline. the most used hook is the least viral one. full stop. • what you are watching: left is the post under analysis, right is Jev answering 14 typed questions about it, each with a confidence score. the run stops at 20.4s because that is when Jev finished all 100k. pulled the corpus through a few X APIs, one parallel pass into Jev. should I drop it to public? Read my latest article on Jev Engineering below and turn your ideas into reality.
Ian Nuttall
@iannuttall
I gave Jev 3,282 of my X posts across 100M views and asked it to find what actually works for growth. 4,252,330 tokens $0.1282 for the full 8m 34s run! Each post got 8 questions about the topic, hook, tone, whether it teaches something, etc. How-to posts got 150 median likes vs the average median of 44. AI and coding was a 1.9x multiplier topic compared and SEO, despite recent posts, was right at base median 1.0x - surprisingly. The recommended topic + angle + voice formula was: AI coding + teach something + provocative
Dan Shipper
@danshipper
we almost never test new foundation models but we've been testing this for ~a week @every and it's pretty wild. the kind of things that will be obviously indispensible in 6-12 months it doesn't produce words as output, it produces probabilities. so it can efficiently act as a judge in cases where you'd need a Fable-level model—but in our testing was 25x faster and 600x lower priced excellent vibe check by @hammer_mt on @every: https://every.to/also-true-for-humans/mini-vibe-check-typesafe-s-jev-judged-everything-i-ve-written-in-0-7-seconds?utm_cta_source=home_main_a_3
ArticleBenchmarks and evals
PickEvery’s editorial vibe check
- Judgments
- 1,709
- Total cost
- <$0.01
- Median per passage
- 0.35 s
Steve Krouse
@stevekrouse
typesafe's jev is fun! live demo you can play with: typesafe-demo.val.run
Ackerman
@Yarilo7brigada
98% of my feed is junk. Now I don’t even see it I built a filter that reads my feed for me. It runs on Jev - a model that doesn’t generate text; it just makes binary decisions, in milliseconds and for pennies. For every post, it evaluates four questions: Is it relevant to my niche? Does it provide real value? What format is it (breakdown, news, ad, meme)? And does it make loud claims with zero proof? Out of 1,000 posts, only 20 remained. 97 milliseconds per post. 1.2 cents for the entire morning. The most frustrating takeaway, nearly half of my feed was ads and memes, not the creators I originally followed for substance. An hour of mindless scrolling turned into five lines over breakfast.
Tools that write, lint and test questions
Agent skills that know the three types, and tools that check a question set before and after it ships: jevkit lints one offline, jev-calibrate tunes criteria against labels you already have, and Sutro's jev-align optimises the question itself.

GitHubSDKs and integrations
jevkit
A Rust CLI for Jev, with a linter that checks questions before you pay.
Ariel Frischer
4
SkillBenchmarks and evals
tenbin
Split a judgment into Choice, Score and Noul, lint it, then measure it.
Shingo Imota
4
GitHubBenchmarks and evals
jev-calibrate
Tune your Jev questions against your own labels, then confirm on held-out data.
smkrv
32
GitHubBenchmarks and evals
jev-align
Turns human labels into calibrated Jev classifiers with GEPA.
Sutro
305
SkillSDKs and integrations
Building with Jev
A skill for writing and fixing programs that call Jev.
Drew Breunig
153
SkillAgents and browsers
Augustus
An agent skill for designing systems around Jev’s judgments.
Basit Mustafa
11
GitHubSDKs and integrations
mcp_jev
A local MCP server that runs ready-made Jev question packs.
Pedro Knigge
1Reading

Guidedocs.typesafe.ai
Quickstart: state, questions, typed answers
The three question types, Choice, Score and Noul, in one support-ticket example.

Guideblog.lepine.pro
Let’s look at Jev
Jean-François Lépine on Choice, Noul and Score, calibrated confidence and batching questions, with a complete Python project.
Guideyoutube.com/@daveebbelaar
Jev Explained for Python Developers
Dave Ebbelaar in Python: a support-ticket classification first, then Choice, Score and Noul, several questions in one call, and latency and price next to Claude Haiku, Opus 5 and Fable 5.1.
Akshay 🚀
@akshay_pachaar
LLMs vs. Jev, clearly explained! TL;DR The key difference is not that Jev generates faster. Jev does not generate text at all. A traditional LLM receives context and produces an answer one token at a time. Even when the output is a small JSON object, every token depends on those generated before it. Jev receives the same context but evaluates predefined decisions directly. When those decisions are independent, it can evaluate all of them in parallel. Consider an agent handling a failed deployment. It may need to determine: → Whether the incident is urgent → Which team should handle it → Whether the proposed command is risky → Whether the task is complete An LLM generates a response containing these answers sequentially. The application then parses and validates it. With Jev, you define the questions and expected answer types upfront. It evaluates them together and returns typed answers with probabilities. Jev supports three decision primitives: 1. **Choice** selects from known options, such as engineering, billing, or sales. 2. **Score** places the input on an ordered scale, such as low, medium, or high risk. 3. **Noul** evaluates a yes-or-no condition and returns the probability that it is true. The probabilities matter as much as the selected answers. If engineering receives 91% probability and billing receives 9%, automatic routing may be reasonable. If the probabilities are 52% and 48%, the system can escalate, gather more context, or call a stronger model. This keeps control inside ordinary software. Code owns the thresholds and consequences. Jev supplies the semantic judgment that a normal `if` statement cannot derive from unstructured text. It works best when the possible answers are known, the decision depends on meaning, and a careful person could judge the input quickly. It is not designed for writing, summarization, code generation, arithmetic, or decisions requiring several dependent reasoning steps. Independent questions can run in parallel, but decisions that depend on earlier results must remain sequential. Jev also cannot return an option outside the declared schema, but it can still select the wrong valid option. Type safety prevents malformed outputs, not incorrect judgments. The clean mental model is this: LLMs generate new language when the answer space is open. Jev evaluates known paths when the answer space is bounded. I wrote the full breakdown explaining Jev and where it fits. The article is quoted below.
Guidex.com
LLMs vs. Jev, clearly explained
Akshay Pachaar: Jev does not generate text at all. It answers Choice, Score and Noul questions in parallel, and your code owns the thresholds.
Ricky Grannis-Vu
@RickyGrannisVu
We rewrote the oldest primitive in programming. The if statement takes a sentence now, and Jev decides. That nested boolean mess eleven rules deep, four AND/ORs, one comment nobody understands anymore? Replace all of it with "if this expense is allowed under the expense policy." Judgment Mode, live now.
Guidex.com
Rewrite if-statements with Jev
Ricky Grannis-Vu on replacing deeply nested boolean expressions with one natural-language sentence for Jev to decide.
Erik Spock Gafni (hiring!)
@EGafni
my personal faves are: AutoResearch for Feature Extraction Use LLMs to generate Jev questions's who's probabilities are fed into a classic ML algorithm like logistic regression to predict a label. Hierarchical Classification Traverse a classification taxonomy using Jev and beam sesarch. Jev is great for graph traversal.
Guidex.com
Two favourite Jev patterns: AutoResearch features and hierarchical classification
Erik Gafni of TypeSafe names his favourites from the cookbooks. AutoResearch for feature extraction: an LLM writes Jev questions, and their probabilities feed a classic algorithm such as logistic regression to predict a label. Hierarchical classification: traverse a taxonomy with Jev.
Where to go next
Once the questions are right, the next decision is what to do with an unsure answer: Jev confidence. The failure modes behind the mistakes above, with the builds that work around them, are on Jev limitations. For whole jobs built from these questions, see Jev for classification and Jev as a judge. To set up a coding agent that writes Jev calls for you, see Jev prompts.
Common questions
- What are Jev's question types?
- Three: Choice picks one option from a list you give it, Score places the text on ordered levels you describe, and Noul returns the probability that a yes/no statement is true. One request can mix all three.
- What is a Noul in Jev?
- A yes/no question. The answer is a single number from 0 to 1, the probability of yes. It has no separate confidence field: a value near 0.5 already means the model is unsure.
- How many options can a Choice question have?
- Up to 255. For more, or for a deep taxonomy, chain Choice questions level by level, or score the list in windows and then choose within the best one.
- How many levels can a Score have?
- At least two, and the API accepts up to ten. Use only as many as you can describe distinctly; three is fine.
- What is the difference between instructions and criteria?
- instructions is the question. criteria is the set of answers: the options for a Choice, the levels for a Score, or optional definitions of yes and no for a Noul. Both can be a string, an object or an array.
- Does Jev see the question ID?
- No. The ID is only the key your answer comes back under, so write the full question in instructions. Choice option names and descriptions are sent to the model.
- Can one question use another question's answer?
- Not in the same request: every question is answered independently. If your code needs the first answer to build the next question or state, make a second request.
Made with Jev is independent and not affiliated with TypeSafe AI. Every figure on this page is the one its author published, linked to where it can be checked.