Jev for classification
Classification is the job Jev was built for: a fixed set of labels, one answer, and a probability for each. Builders who published both a volume and a bill paid between $0.0067 and $0.2000 per thousand items. This page covers the question type to pick, how to get several labels in one call, and where to hand off the unsure ones.
Updated 8 Oct 2026 · by Made with Jev
In short
- Choice picks one label from up to 255 options. Score places an item on 2 to 10 ordered levels. Noul is a yes/no probability. docs.typesafe.ai/primitives
- For labels that can co-occur, ask one Noul per label in the same request. A Choice always returns exactly one option. primitives/noul
- Put every question in one call. TypeSafe measured 13 questions batched at 12.2x cheaper and 10.0x faster than 13 separate calls, with the same answers. parallel_questions cookbook
- You set the thresholds. Calibration is measured across groups of predictions; it doesn't guarantee that any single answer is correct. concepts/system-one
What builders paid per item
| Build | What | Items | Their bill | Per 1,000 items |
|---|---|---|---|---|
| 100,000 viral posts in 20.4 seconds | posts scored | 100,000 | $0.67 | $0.0067 |
| X timeline labeler | posts labelled | 1,000 | $0.03 | $0.0300 |
| 3,282 posts, eight questions each | posts analysed | 3,282 | $0.1282 | $0.0391 |
| 1,891 ads in 19 seconds | ads classified | 1,891 | $0.12 | $0.0635 |
| 1,315 posts across eight dimensions | posts sorted | 1,315 | $0.086 | $0.0654 |
| 500 emails for 3.5 cents | emails triaged | 500 | $0.035 | $0.0700 |
| 1kpapers | papers classified | 1,018 | $0.08 | $0.0786 |
| A 26-sheet plan set in 2.9 seconds | drawing sheets read | 26 | $0.0052 | $0.2000 |
Each author's volume and bill; the division is ours. Jev bills input tokens only, so the figure follows how much text each item carries and how many questions it gets. Movez asked 14 per post. Median: $0.0644 per 1,000 items. Jev statistics and Jev pricing.
Choice, Score or Noul for a label
| Type | Use it for | Returns | Limit |
|---|---|---|---|
| Choice | one label from a known set with no order | choice, probabilities, confidence | up to 255 options; add an "other" or "none of the above" option |
| Score | a position on levels you describe (severity, frustration) | score (can fall between levels), probabilities, legend, confidence | 2 to 10 levels |
| Noul | a clean yes/no | noul (0–1), no confidence field | — |
- A Noul of 0.5 means the model is unsure, not "medium". If the label is a degree, use a Score. docs.typesafe.ai/primitives
- In some cases jev-1.13 leans toward the first option, so reorder the labels and check the answer holds. See option order.
- For a taxonomy too big or too deep for one question, chain Choices level by level and keep the best K paths (beam search). TypeSafe runs this over patents, Shopify products, MeSH and a codebase. hierarchical_classification cookbook.
Several labels on every item, in one call
Each question in a request sees the same state and is evaluated independently, in parallel, so adding labels barely changes latency and costs only the question tokens. primitives/noul ("For a checklist of conditions, ask many Noul questions in one request").
from typesafe_sdk import Noul, TypeSafeClient
TAGS = {
"billing": "Is `ticket` about a charge, invoice or refund?",
"bug": "Does `ticket` report something that is broken?",
"feature_request": "Does `ticket` ask for something the product does not do yet?",
}
with TypeSafeClient() as client:
response = client.system_one(
state={"ticket": text},
questions={tag: Noul(instructions=q) for tag, q in TAGS.items()},
)
labels = [tag for tag in TAGS if response.answers[tag].noul > 0.7] # threshold: illustrative, set yours on labelled dataThresholds, abstaining and human review
- Choice and Score answers carry
confidence, from 0 to 1, computed from how the probabilities are spread. A Noul doesn't; a value near 0.5 already says "unsure". The docs suggest three paths: act, confirm, or send to a person or another system. "Start with conservative thresholds, test with your own data." docs.typesafe.ai/confidence - Three examples from TypeSafe, all labelled with their model and date:
- 60 SEC filings into 75 industry groups (jev-1.12, 2026-08-12). A 0.9 confidence cutoff splits them in half: the confident half is right 90% of the time and the rest 40%. Reporting the rest one level up, as a division, raises them to 70%, with no second call. classification_using_confidence cookbook
- Moderation with 8 Choice questions × 15 runs (jev-latest, 2026-09-11). Requiring a top probability of at least 0.60, and calling everything else "uncertain", raised agreement from 90.8% to 99.2%, with 74.2% of answers automatic. consistency_choice_cookbook
- An insurance claim with 14 Noul questions × 15 runs (jev-latest, sampled 2026-09-11). Values from 0.30 to 0.70 go to human review. One answer moved between 0.43 and 0.53 across the runs, crossing a 0.5 line, which is why the middle band exists. consistency_noul_cookbook
- Builder example: Hassan routed every prediction under 95% confidence to Kimi K3. That was 31 of 100 emails, and the result was 96 of 100 correct in 16 s for about $0.07, of which $0.003 was Jev. x.com/nutlope
- Builder example: firehose-judge sends posts Jev won't commit on to a "needs a human" lane. github.com/ragelink/firehose-judge
- Builder example: linkedin-noslop keeps anything uncertain visible. github.com/sushrutb17/linkedin-noslop-extension
- Builder example: jeval sets the hand-off line from the cost of a mistake, measured on labelled data. github.com/rlaope/jeval
Labelling in bulk
- The API reference lists no batch endpoint (as of Oct 2026). Throughput comes from many questions per request plus concurrent requests. docs.typesafe.ai/api
- Limits are 100K tokens/s and 80 requests/s, "adjusting dynamically", and higher on enterprise plans. The SDKs retry 429s with backoff. docs.typesafe.ai/models
- Computed caveat: Movez reports 100,000 posts in 20.4 s. One request per item would take about 21 minutes on one key at 80 requests/s. Treat his figure as his own and plan on your own limits.
- Builds, from the bulk group: DuckDB at about 10 s per 1,000 rows (Hamilton Ulmer); 129 rows in about 1 s from Postgres (Zachi); jev-curate's 24.0 rows/s, which comes from a mock bench, and whose 1,500 rows/s figure is a target, not a result; jevkit's offline lint of a question set before any call.
Compare it against an LLM before you switch
- system-one-adapter-python, MIT, from TypeSafe: "A drop-in replacement for typesafe_sdk's system_one evaluation API, backed by LLM APIs instead of TypeSafe. Useful for comparing TypeSafe against an LLM on cost/speed/intelligence." Install
system-one-adapter[openai],[anthropic]or[gemini].llm_answer_modeis "probabilities" or "discrete".response.usageadds total tokens, retries and latency. github.com/typesafe-ai/system-one-adapter-python - Run the same labelled set through TypeSafeClient and through the adapter, then compare accuracy, cost and latency.
What the TypeSafe comparisons showed, credited to TypeSafe:
- Noul answers: Jev's mean per-question standard deviation was 0.0102, lower than every LLM probability condition.
- Choice labels: Jev repeated its plurality label 90.8% of the time, against 87.5% to 100% for the LLM settings, and flipped on 2 of 8 questions.
- Sources: the choice consistency cookbook and the noul consistency cookbook.
Builder results, both directions:
- construction-plan-classifier matched GPT-4.1 and GPT-6 Astra on 100% of sheet-level classifications, 17 to 21 times cheaper (Trinay Hari's claim).
- Mike Taylor's Every test found 6 of 7 planted defects; only the larger model found the seventh.
- browser-agent-2s-to-02s: −6 points of accuracy.
- a-local-classifier-instead: a local classifier was as fast and as good for that author.
One label from a fixed list
A Choice question with every category as an option. Hassan sorted 1,018 papers into 24 topics for $0.08 in total, at a median of 256 ms a paper.
Hassan
@nutlope
I used Jev to classify 1,018 AI research papers. The result: $0.08 total cost and 256ms median end-to-end latency per paper. The pipeline was: 1. Summarize each paper with DeepSeek V4 Flash 2. Send the title + summary + 24 possible topics to Jev 3. Use Jev to classify each paper 4. Visualize everything on http://1kpapers.com The summaries cost $3.99 on @togethercompute. The classifications cost $0.08 on @typesafeai. So for just over $4 of inference, I ended up with a pretty useful way to explore the top AI research papers from the past year. I think this is where things are heading: different models for different parts of the workflow, instead of using one model for everything. I’m running evals on the Jev classifications before replacing the current ones, but the site is already live: http://1kpapers.com
Peter Wang
@the_cyw
I made a chrome extension to label all the X posts on my timeline. It tells me if each post is clean, engagement bait, promo, secondhand or filler. $0.03 for 1000 posts. Open sourced if you want to try it out.
ares. 🎧
@aresotik
ESTA HERRAMIENTA ACABA DE ROMPER TODO EL MERCADO DEL AD SPY Maxfusion ha cogido JEV, el modelo nuevo de TypeSafe, y le ha metido la ad library entera de una marca → 1.891 anuncios clasificados → 19 segundos → 0,12 $ Y no es un resumen: cada anuncio etiquetado por etapa del funnel y estilo creativo, más la radiografía completa de la cuenta Llega pronto al MCP de maxfusion
Riley Brown
@rileybrown
Yeah Jev by @typesafeai is very cool. It classified 500 emails in seconds. And it costed 3.5 cents.
Trinay Hari
@hari_trinay
Built a construction plan-set classifier with Jev. Proq turns civil and building plan sets into bills of materials using an LLM pipeline we built on GPT-4.1. Jev classified an entire 26-sheet plan set in 2.9 seconds for $0.0052. It matched GPT-4.1 and GPT-6 Astra on 100% of sheet-level classifications while running 17–21x cheaper and 5x faster than our production pipeline.
Sabrina
@sabrinaesaquino
Jev is now live on the Venice API. Watch it classify 24,000 Hacker News posts into 12 categories in about 2 minutes
Kenny Chen|AI 实战
@KennyChinaTech
Jev 这个案例很适合小团队:2,300 篇 AI 论文,约 83 秒,成本 $0.14。 但这个数字不能直接当成 Jev 的单模型成本。旧标签先由 DeepSeek V4 Flash 跑过,真正该测的是整条分类链路:预处理、Jev 决策和人工抽检加起来还剩多少。 x.com/omarsar0/statu…
Fayaz Ahmed
@fayazara
Made myself a little image classifier with OCR + Jev It was able to categorise ~900 images in 40 seconds Pretty cool
Several labels on every item, in one call
One request carries every question, so a post can get a topic, a hook, a format and a bait check at once. Movez asked 14 yes-or-no questions of each of 100,000 posts and paid $0.67.
Movez
@0xMovez
I just built a Jev X Viral Post Analyser. 100,000 viral X posts. 20.4 seconds. $0.67. Claude Opus 5, same corpus, same clock, got through 214 posts and spent $0.98. per post that is ~680x cheaper the full Opus pass would have run $458. viral analysis is the perfect Jev job. • it is not writing, it is 14 yes/no calls per post: > does the hook open a loop, > is there a number in the first line, > is the proof real or claimed. classification, not prose. • what it found: 1,220 posts broke into the top 1%. baseline 1.22%. > superlative claim - 2.34% viral. 1.92x baseline > contrarian take - 1.59%. 1.31x > launch / tool drop - 1.46%. 1.19x and numbered lists, the thing everyone writes: 0.55%. below baseline. the most used hook is the least viral one. full stop. • what you are watching: left is the post under analysis, right is Jev answering 14 typed questions about it, each with a confidence score. the run stops at 20.4s because that is when Jev finished all 100k. pulled the corpus through a few X APIs, one parallel pass into Jev. should I drop it to public? Read my latest article on Jev Engineering below and turn your ideas into reality.
Ian Nuttall
@iannuttall
I gave Jev 3,282 of my X posts across 100M views and asked it to find what actually works for growth. 4,252,330 tokens $0.1282 for the full 8m 34s run! Each post got 8 questions about the topic, hook, tone, whether it teaches something, etc. How-to posts got 150 median likes vs the average median of 44. AI and coding was a 1.9x multiplier topic compared and SEO, despite recent posts, was right at base median 1.0x - surprisingly. The recommended topic + angle + voice formula was: AI coding + teach something + provocative
Yum⋆₊˚
@yuhasbeentaken
Jev classified 1,315 X posts for about $0.086 in estimated model cost 😂 seeing everyone's Jev demos made me want to build something for my own content research. i'd collected a lot of posts, but figuring out what they had in common still meant opening them one by one and taking notes. so i built a dashboard around Jev. it labels each post across 8 dimensions, including topic, hook and writing style. now i can filter by topic and hook, compare engagement, and open the original posts to see the examples behind each pattern. my archive is a lot easier to learn from now.
Ackerman
@Yarilo7brigada
98% of my feed is junk. Now I don’t even see it I built a filter that reads my feed for me. It runs on Jev - a model that doesn’t generate text; it just makes binary decisions, in milliseconds and for pennies. For every post, it evaluates four questions: Is it relevant to my niche? Does it provide real value? What format is it (breakdown, news, ad, meme)? And does it make loud claims with zero proof? Out of 1,000 posts, only 20 remained. 97 milliseconds per post. 1.2 cents for the entire morning. The most frustrating takeaway, nearly half of my feed was ads and memes, not the creators I originally followed for substance. An hour of mindless scrolling turned into five lines over breakfast.
Matthew Berman
@TheMattBerman
jev is INSANE. in 40 seconds it broke down 724 live ads from 37 brands. every hook. every format. offer. cta. awareness stage. landing page mismatch. used 9 cents of tokens. (will be avail in @stealads + mcp)

GitHubCoding and code review
Jev PR Labeler
GitHub PR labels chosen by Jev, with size measured as scope, not lines.
Jeremy Huang
4Dan Shipper
@danshipper
we almost never test new foundation models but we've been testing this for ~a week @every and it's pretty wild. the kind of things that will be obviously indispensible in 6-12 months it doesn't produce words as output, it produces probabilities. so it can efficiently act as a judge in cases where you'd need a Fable-level model—but in our testing was 25x faster and 600x lower priced excellent vibe check by @hammer_mt on @every: https://every.to/also-true-for-humans/mini-vibe-check-typesafe-s-jev-judged-everything-i-ve-written-in-0-7-seconds?utm_cta_source=home_main_a_3
ArticleBenchmarks and evals
PickEvery’s editorial vibe check
- Judgments
- 1,709
- Total cost
- <$0.01
- Median per passage
- 0.35 s
Hand the unsure ones to a person or a larger model
Every label comes with a probability, and code decides what is sure enough to act on. Hassan sent the 31 of 100 emails under 95% confidence to Kimi K3. The pipeline got 96 right, and Jev's share of the bill was $0.003.
Hassan
@nutlope
Jev + Kimi K3 for fraud detection! TLDR: Jev classified 100 emails in 1.42 seconds, then I routed the uncertain cases to Kimi K3. The full pipeline got 96/100 correct for only ~$0.07. Video is not sped up, check out the live run! Here was my process: I gave Jev 100 emails to classify (a mix of 50 legit & 50 fraudelent emails). It classified all of them in 1.42 seconds. An underrated feature about Jev is it will give you the confidence score for a classification, so I routed any prediction under 95% confidence to Kimi K3 to be fully sure. 31 emails fell below that threshold. After routing those to Kimi K3, the combined pipeline reached 96% accuracy. The full run took 16 seconds & ~$0.07 in inference costs: - $0.068 from Kimi K3 on @togethercompute - $0.003 (1/3 of a cent) from Jev on @typesafeai. I think this is a really interesting pattern: use a fast specialized model like Jev for the narrow task, then route the uncertain cases to a larger LLM. I feel like this kind of approach could be a game changer for use cases like fraud or anything realtime. You can use the speed & low cost of Jev while having a larger LLM as a fallback to ensure high accuracy.
XSecurity and abuse
PickFraud detection with Jev and Kimi K3
- Emails
- 100 in 1.42 s
- Correct
- 96/100
- Cost
- ~$0.07 pipeline ($0.003 Jev, $0.068 Kimi K3)

GitHubSecurity and abuse
firehose-judge
The live Bluesky firehose, judged post by post, with a lane for humans.
Leo Mata
0~$0.00003


SkillBenchmarks and evals
tenbin
Split a judgment into Choice, Score and Noul, lint it, then measure it.
Shingo Imota
4
GitHubSecurity and abuse
jevmod
Moderation with a probability per category and thresholds you set.
Omar Hernandez
1
Label a table or a dataset where it lives
Jev as a SQL function or a streaming filter, so the labels land next to the rows. Hamilton Ulmer reports about ten seconds per thousand rows from DuckDB.
Hamilton Ulmer
@hamiltonulmer
I made a DuckDB extension where you can use @typesafeai 's Jev to do quick classification of rows in any csv/parquet file or duckdb table about 10sec for 1k rows ~ better than using an LLM, way more ergonomic than a classifier game-changing for data analysis!

GitHubSDKs and integrations
duckdb-jev
Ask a question about every row in SQL, and get a real SQL type back.
Colliber
28Zachi
@iam_zachi
I think I just cooked something 🔥 jev(): a PostgreSQL extension that searches your whole database in natural language. No index, no embeddings, just one function. WHERE jev(people, 'could work from home') or WHERE jev(people, 'name sounds european') 129 rows judged in ~1s for $0.0009. Second run: 6ms from cache.

GitHubDocuments and OCR
jev-curate
A Rust dataset sifter that keeps clean rows verbatim and drops the rest.
Akash Priyadarshi
106
GitHubSDKs and integrations
jevkit
A Rust CLI for Jev, with a linter that checks questions before you pay.
Ariel Frischer
4Check it against what you use now
Results from people who measured before switching, including the ones where Jev lost. xunaoo made a browser-agent step ten times faster with Jev and lost six points of accuracy.
Vincent.seo
@xunaoo
Jev 가 우리 로직을 대체할 수 있을까요? 아닙니다. 대체가 아니라 로직 안으로 들어앉습니다. 제 브라우저 에이전트에 붙여 봤습니다. 판단 한 단계가 2초에서 0.2초, 모델 호출이 열 번에서 두 번. 그런데 정확도는 6%p 낮습니다.
Petru - Tech Driven
@techdrivenpetru
Yes but not with Jev. I used another classifier, locally, just as fast, just as good. Example: setup a "server" in python that loads the classifier model. Add a hook in Claude that fires on "pre-tool-use" and next time you ask Claude a random question like "how do I lint check a project?" and it tries to freelance and read your entire repo only to intoxicate itself and pollute its context, the classifier will slap its hand, say "no sir, you answer from knowledge" and deny the tool call. I tested this yesterday with success, but need to refine it as it misfires. Basically I was able to identify general queries, instances where I would ask something and Claude would rush ahead and run pip install without me asking or just write code instead of answering. A classifier is hypercheap compared to a regular LLM and would catch all of these. The model I used was DeBERT large. It's still stupid fast, I tested it on an Apple with M1 (regular) and you don't feel it running.
XOpen source
A local classifier instead of Jev

GitHubBenchmarks and evals
jev-acento
A pre-registered audit of how Jev handles Spanish.
Marcos Martinez
0VertrAI
@vertr_ai
154ms vs 860ms is a useful routing result, scoped to these 200 synthetic classification cases—not a blanket ChatGPT comparison. We credited your latency and cost charts in our Jev video, alongside the game/browser demos: x.com/vertr_ai/statu…
XBenchmarks and evals
154 ms against 860 ms, and the scope of it
- Routing
- 154 ms vs 860 ms
- Cases
- 200 synthetic

Live siteOpen source
Julia 1
A 144.3M-parameter decision model that edges Jev on typed decisions.
Supersonic Labs
144.3MWhere it doesn't fit
- Images: Jev is text only, so put OCR or a vision model first. See Jev images.
- Non-English labels score lower. See Jev limitations.
- Counting labels: do it in code. See maths and numbers.
- You can't train it on your labels. See Jev fine-tune.
- Keys and signup status as of Oct 2026. See Jev access.
Reading
Guideyoutube.com/@samwitteveenai
Jev: the ultimate classification model?
Sam Witteveen on System 1 thinking, then demos of Choice, Score and Noul, a practical example, and chained actions.
saturn
@SatOnchain
this paper is f*cking gold. someone made handbook on "how jev can be used as classifiers for evals on scale" anyone who solve this has chance to make millions. Bookmark and read full handbook.
Guidex.com
A handbook on Jev as a classifier
SatOnchain shares a handbook on using Jev as a classifier for large-scale evals.
cocktail peanut
@cocktailpeanut
Jev is cool not because it re-invented classification, but because it makes ARBITRARY classification into a type-safe programmable primitive. A general purpose zero shot decision model whose native interface is RUNTIME-DEFINED typed decisions, optimized for that exact interface
Guidex.com
Arbitrary classification as a primitive
cocktail peanut: Jev did not reinvent classification. It makes any classification a typed decision you define at runtime.
Erik Spock Gafni (hiring!)
@EGafni
my personal faves are: AutoResearch for Feature Extraction Use LLMs to generate Jev questions's who's probabilities are fed into a classic ML algorithm like logistic regression to predict a label. Hierarchical Classification Traverse a classification taxonomy using Jev and beam sesarch. Jev is great for graph traversal.
Guidex.com
Two favourite Jev patterns: AutoResearch features and hierarchical classification
Erik Gafni of TypeSafe names his favourites from the cookbooks. AutoResearch for feature extraction: an LLM writes Jev questions, and their probabilities feed a classic algorithm such as logistic regression to predict a label. Hierarchical classification: traverse a taxonomy with Jev.
Ricky Grannis-Vu
@RickyGrannisVu
Jev gives you a probability. You set two lines on it: above the top one it's a yes, below the bottom one it's a no, anything in the gap runs a third branch. It's just a normal branch, so you can handle uncertainty flexibly.
Guidex.com
Jev’s probability thresholds
Jev outputs a probability, two thresholds split the answer into yes, no and a third branch between them.
Guideyoutube.com/@vogeldev
Meet Jev: tested on 1,000 emails
Ryan Vogel runs Jev on 100 and then 1,000 of his emails: category, priority, spam and reply predictions.
Guideyoutube.com/@daveebbelaar
Jev Explained for Python Developers
Dave Ebbelaar in Python: a support-ticket classification first, then Choice, Score and Noul, several questions in one call, and latency and price next to Claude Haiku, Opus 5 and Fable 5.1.
Guidepypi.org
langchain-typesafe
The LangChain package with TypeSafeClassifier, to use Jev inside a LangChain app.
Common questions
- What does it cost to classify with Jev?
- Builders who published a volume and a bill paid between $0.0067 and $0.2000 per 1,000 items. Jev bills input tokens only, $0.042 per million as of Oct 2026.
- Should I use Choice, Score or Noul for labels?
- Choice for one label from a set, Score for a degree on levels you describe, and Noul for yes/no.
- Can Jev assign more than one label to an item?
- Yes. Ask one Noul per label in the same request and keep every label above your threshold. A single Choice returns one label.
- How many categories can a Choice question have?
- Up to 255. For deeper taxonomies, chain Choice questions level by level.
- What confidence threshold should I use?
- There is no universal one. TypeSafe's examples use 0.9 for industry codes and a 0.60 top probability for moderation. Set yours from the cost of a mistake and test it on labelled data.
- How do I compare Jev with an LLM on my own data?
- Run the same questions through TypeSafe's system-one-adapter-python, which answers the system_one API with OpenAI, Anthropic or Gemini models, and compare accuracy, cost and latency.
- Can Jev classify images?
- Not directly. It takes text, so convert the image with OCR or a vision model first.
Made with Jev is independent and not affiliated with TypeSafe AI. Every figure on this page is the one its author published, linked to where it can be checked.