Can Jev do maths, dates and counting?
No, and TypeSafe says so itself. Its Jev 1.13 jaggedness page, last reviewed 2 Oct 2026, lists 9 failure modes. Every one has the same fix: code computes, Jev judges. Source: TypeSafe Jev 1.13 jaggedness.
Updated 8 Oct 2026 · by Made with Jev
In short
- Arithmetic, counting and date comparison belong in code.
- Jev picks among options; it doesn't write. Extract candidates first.
- More state is not better: irrelevant detail costs accuracy.
- Text in the state can steer the answer, so screen it and test.
- Confidence is calibrated across many answers; it isn't a guarantee for any single one. Source: TypeSafe System One.
The 9 limits, at a glance
Most of these start as a question worded the wrong way. How to write one that avoids them: Jev question types.
| Limit | What goes wrong | Do this instead |
|---|---|---|
| Literal reading | Boundary cases and implicit conditions are unreliable | State the exact condition, put boundary cases in the criteria, split interpretation into two literal questions |
| Maths and numbers | Arithmetic, counting, hex and RGB comparisons struggle | Compute in code; ask one yes/no question per item; convert numbers to named buckets |
| Dates | Date comparison and extraction are unreliable | Ask Choice questions for date parts, build the date in code, do every comparison in code |
| Indirection | “Judge the field the user asked about” fails | Give direct instructions and name the state field the question is about |
| Large state and context | Irrelevant detail lowers accuracy | Filter before sending, or run a Noul relevance pass first |
| Adversarial content and prompt injection | Text in the state can move the answer | Write explicit criteria, test edge cases, screen input and output in a separate request |
| Contradictory instructions and criteria | When criteria invert the instruction, jev-1.13 might get confused | Make the criteria extend the instruction, never invert true/false |
| Option order | jev-1.13 can lean toward the first option in some cases | Reorder the options and check the answer stays the same |
| Writing (generation) | Jev cannot generate text, code or structured data | Extract candidates with regex or an LLM, then let Jev choose |
1. Literal reading
Jev reads the state and the question literally. Boundary cases and implicit conditions are unreliable.
Fix: State the exact condition, put boundary cases in the criteria, and split interpretation into two literal questions that code combines.
Source: TypeSafe: Literal reading.
2. Maths and numbers
Arithmetic, counting and number comparisons are unreliable; do the maths in code.
Counting: Loop over the items in code, ask one Noul per item, and sum in code.
let count = 0;
for (const item of items) {
const { noul } = await jev(item, {
matches: { type: "noul", instructions: "This item matches the criteria" }
});
if (noul > threshold) count++;
}
return count;Numbers as text: Hex and RGB comparisons are unreliable, so convert in code to a value or a named bucket. See Jev images for colour names.
Score interpolation: Don't try to recover exact magnitudes from a Score; only threshold it.
Source: TypeSafe: Maths and numbers.
3. Dates
Date comparison and extraction are unreliable.
Fix: Ask Choice questions for the date's parts, including a “not stated” option; code builds the date and does every comparison.
Source: TypeSafe date extraction cookbook.
No build in this directory does this yet.
4. Indirection
Instructions like “Judge the field the user asked about” fail.
Fix: Give direct instructions and name the state field the question is about.
Source: TypeSafe: Indirection.
5. Large state and context
Irrelevant detail lowers accuracy.
Fix: Filter before sending, or run a Noul relevance pass first.
Limits: 64k per request, of which 32k is state plus the longest question. Gateways list 32k.
Sources: TypeSafe: Large state and context, classifying RAG passages cookbook, TypeSafe models.
6. Adversarial content and prompt injection
The doc says state “is data, and jev-1.13 does not treat it as hostile by default”, so injected text can move the answer.
Fix: Write explicit criteria, test edge cases before rollout, and screen input and output in a separate request.
Note the tension: A screen built on Jev can itself be steered, so test it too.
Sources: TypeSafe: Adversarial content, LLM guardrails cookbook, classifying RAG passages cookbook.
The jev-as-judge build's author cautions that it “is not a proven defence against injection”.
7. Contradictory instructions and criteria
When criteria invert the instruction, jev-1.13 might get confused.
Fix: Make the criteria extend the instruction, and never invert true/false between them.
Source: TypeSafe: Contradictory instructions.
8. Option order
In some cases jev-1.13 leans toward the first option.
Fix: Reorder the options and check the answer stays the same.
Source: TypeSafe: Option order.
9. Writing (generation)
Jev cannot generate text, code or structured data.
Fix: Extract candidates with regex or an LLM, then let Jev choose.
Source: TypeSafe pre-parsed value extraction cookbook.
The haiku-correction build shows that the claim “Jev produced the hiragana” doesn't hold.
Two more limits from the docs and testers
Non-English text: Accepted, but English is the primary training language and other languages, CJK included, score lower.
Source: TypeSafe models: Language support.
Alternative: Laya publishes a multilingual checkpoint (the author's claim).
Knowledge you didn't supply: jev-capability-atlas reports that Jev is accurate when the answer is in the state, and often confidently wrong when it isn't (an independent finding).
Images: Jev is text only, so put OCR or a vision model first. See Jev images.
Fine-tuning: You cannot fine-tune Jev itself. What people train instead: Jev fine-tune.
Literal reading
Boundary cases and implicit conditions are unreliable. State the exact condition, put boundary cases in the criteria, and split interpretation into two literal questions that code combines.

SkillBenchmarks and evals
tenbin
Split a judgment into Choice, Score and Noul, lint it, then measure it.
Shingo Imota
4
GitHubBenchmarks and evals
jev-evaluation
Nine experiments and 28 predictions, all fixed before any data.
Will Kelly
0123,805Maths and numbers
Arithmetic, counting and number comparisons are wrong. Loop over items in code, ask one Noul per item, and sum in code. Convert hex and RGB to named buckets.
ラプター | 高専 ロボコン ビジコン
@Raptor_zip
型付き意思決定AIモデル「Jev」で双腕ロボを動かした結果まとめてみた🦾 3層制御のうち意思決定(2層)をJevに任せ、IKや物理計算はコード側に分離する構造を作りました。500msの応答と0.5円/試行で早くて激安。
XRobotics and devices
A dual-arm robot with Jev in the middle layer
- Response
- ~500 ms
- Cost per attempt
- ~¥0.5
Justin Schroeder
@jpschroeder
I rebuilt Tesla Full Self Driving with Jev in less than an hour. This model is a total unlock.
Saeed
@stringsaeed
built a highlighter on jev paste any language → my code tokenizes → jev names the lang, colours every word, then says which of 9 lint rules fire and where nine rules in code. jev just answers. near instant lab.saeed.sh/highlight
Ackerman
@Yarilo7brigada
98% of my feed is junk. Now I don’t even see it I built a filter that reads my feed for me. It runs on Jev - a model that doesn’t generate text; it just makes binary decisions, in milliseconds and for pennies. For every post, it evaluates four questions: Is it relevant to my niche? Does it provide real value? What format is it (breakdown, news, ad, meme)? And does it make loud claims with zero proof? Out of 1,000 posts, only 20 remained. 97 milliseconds per post. 1.2 cents for the entire morning. The most frustrating takeaway, nearly half of my feed was ads and memes, not the creators I originally followed for substance. An hour of mindless scrolling turned into five lines over breakfast.

Indirection
Instructions like 'Judge the field the user asked about' fail. Give direct instructions and name the state field the question is about.

GitHubCoding and code review
Jev Review
A staged code-review workflow with a local dashboard.
Dev Agrawal
675Large state and context
Irrelevant detail lowers accuracy. Filter before sending, or run a Noul relevance pass first. The hard limit is 64k per request.
NO1ennn
@N01ennn
this is pure f*cking treasure A Stanford AI research group finally drew the perfect RAG system: retrieval, Jev and agents in one loop, and it fixes the 3 things that break every RAG app: > the LLM reads 20 passages when only 3 matter > it answers questions your docs can't answer > it trusts whatever text it retrieves here's how it runs: > a lead agent sends the query > hybrid search pulls the top 20 candidates, dense + keyword > ONE Jev request scores all 20 + 2 gates: answerable? injection? > only passages above 0.6 reach the writer agent > a second Jev call checks every claim against its source > grounded answer, with citations and when the docs don't have it: > answerable fails, the writer never runs > a researcher agent rewrites the query and retries once > still nothing? "not in the docs". zero tokens spent on a guess retrieval casts the net. Jev decides what's real. agents do the work save this before you build your next RAG
Coach Shweta Bajaj
@shwetabjaj
Jev Skill Suggestion for Claude Code is a smart idea. Instead of loading every skill into context, Jev decides which one is actually relevant and injects only that. My takeaway: better context hygiene, less clutter, and potentially more efficient agent workflows. @typesafeai @vercel
XContext and memory
Skill suggestion injects one skill

GitHubContext and memory
Jev Sift
Classify first, read selectively: an agent plugin and MCP tool.
Kush Bhuwalka
46Adversarial content and prompt injection
Text in the state can steer the answer. Screen input and output, write explicit criteria, and test before rollout. Note the tension: a screen built on Jev can itself be steered.


GitHubSDKs and integrations
jev-mcp
An MCP server with claim checks, screening and ranking built on Jev.
Joey Kudish
506NO1ennn
@N01ennn
this is pure f*cking treasure A Stanford AI research group finally drew the perfect RAG system: retrieval, Jev and agents in one loop, and it fixes the 3 things that break every RAG app: > the LLM reads 20 passages when only 3 matter > it answers questions your docs can't answer > it trusts whatever text it retrieves here's how it runs: > a lead agent sends the query > hybrid search pulls the top 20 candidates, dense + keyword > ONE Jev request scores all 20 + 2 gates: answerable? injection? > only passages above 0.6 reach the writer agent > a second Jev call checks every claim against its source > grounded answer, with citations and when the docs don't have it: > answerable fails, the writer never runs > a researcher agent rewrites the query and retries once > still nothing? "not in the docs". zero tokens spent on a guess retrieval casts the net. Jev decides what's real. agents do the work save this before you build your next RAG

GitHubBenchmarks and evals
awesome-jev-robustness
Tests, calibration audits and failure-mode studies of Jev.
Yifan-Lan
5Contradictory instructions and criteria
When criteria invert the instruction, the answer is wrong. Make the criteria extend the instruction, and never invert true/false between them.

SkillBenchmarks and evals
tenbin
Split a judgment into Choice, Score and Noul, lint it, then measure it.
Shingo Imota
4Option order
In some cases jev-1.13 leans toward the first option. Reorder the options and check the answer stays the same.

GitHubBenchmarks and evals
awesome-jev-robustness
Tests, calibration audits and failure-mode studies of Jev.
Yifan-Lan
5Writing (generation)
Jev cannot generate text, code or structured data. Extract candidates with regex or an LLM, then let Jev choose.
たく|ガチのCopilot達人
@taku_ai_case
これはめちゃくちゃ勉強になりました。 Jevに回答させるのではなく、CodexやClaude Codeへ渡す記憶候補を選ばせる。文章を書けないAIの特性を、ここまで実用的な仕組みに落とし込む発想がすごいです。
XContext and memory
Let Jev pick the memories, not the answer
Justin Schroeder
@jpschroeder
I rebuilt Tesla Full Self Driving with Jev in less than an hour. This model is a total unlock.

GitHubSDKs and integrations
jev-mcp
An MCP server with claim checks, screening and ranking built on Jev.
Joey Kudish
506yuri
@yurinakanishi33
昨日のJevハッカソンで発表した「Jevを使った俳句生成」について訂正があります。 発表では「JevのChoiceで、ひらがなを1文字ずつ生成している」と説明しましたが、間違いでした。 実際には、季語・名詞・助詞・動詞・副詞・結びの定型句などの単語をコード側であらかじめ用意し、それらを組み合わせて5音・7音の完成行候補まで先に作っていました。 Choiceに渡していたのは、その時点の文字列から、いずれかの完成行に到達できる「次の1文字」だけです。 さらに、候補が1つならコードで確定。複数ならChoiceが返した確率の上位3候補を再正規化し、コード側でもう一度抽選していました。そのため、Choiceが選んだ文字を必ず使っていたわけでもありません。 つまり、単語や文章の道筋をコード側で先に用意したうえで、Choiceが返した確率を使ってコード側で分岐する実装でした。申し訳ありません。 添付動画は、コード側で用意していた語彙と単語のつながりに関する制約を外した版です。残した制約は、5・7・5の区切りと、使用できる文字を1字1音として扱える73字に限定する文字種制限だけです。 毎回、ここまでの句、季節・状況・気分などの文脈、残りの音数、73字すべてをChoiceに渡しています。Choiceが返した全73候補の確率をコード側で再正規化し、その確率を重みに次の1字を抽選。これを17回繰り返しています。 結果は......全くできていませんねw #aimeetup
XBenchmarks and evals
A correction about generating haiku with Jev
Where these come from
Guideyoutube.com/@GaryExplains
Jev: fast and cheap, but there is a caveat
Gary Explains on the catch: Jev understands natural language but it is not an LLM. It answers with structured values and a confidence level.
Jason Zhu
@GoSailGlobal
拿 Jev 做搜索重排,我先泼一盆冷水:单独用,它没打赢向量检索 TypeSafe 的 Jev 这阵子很火,一堆项目拿它做重排。我们在 Agent Skills Hub 的 33,047 条目录上认真测了一次,164 条中英文真实查询,9,831 对分级标注,整套只花了 2.6 美元 三个结论 01|单独重排,约等于没赢 Jev 重排 bge-m3 的前 30 条,NDCG@10 只多了 0.012,置信区间跨过零。MRR 和前三命中率倒是涨得明显,它很会把最强的那一条顶到第一,后面几条基本是重新洗牌 02|裁判偏差,被我们量出来了 Jev 自己也参与了打标注,这就是循环 只用 Jev 当裁判,它领先 0.053 两个裁判合并,领先 0.012 只用 Haiku 当裁判,反而落后 0.028 同一组比较,换个裁判结论直接翻面。所有涉及 Jev 的结论,我们只认 Haiku 那一列 03|真正稳赢的是融合 把 Jev 和 bge-m3 的排序做 RRF 融合,NDCG@10 到 0.864,比纯向量高 0.064 到 0.116,三种裁判下都成立。代价是每次查询多一次 API 调用,约 0.0002 美元 顺便暴露了我们自己的问题:Hub 线上 CLI 用的关键词排序只有 0.609,短板是召回。相关结果有一半压根没进候选,后面怎么重排都救不回来 数据、标注、每条查询的得分全部开源,不用 API key 就能复现打分
Guidex.com
Jev as a reranker: an honest negative
Jason Zhu tested Jev reranking on 164 real queries. Alone it did not clearly beat vector search; fused with it, it did. In Chinese.
Isaac Flath
@isaac_flath
I've been using Jev by @typesafeai Here's the six things i've tried and am confident I'll still use Jev for 60 days from now. There's many more experiments, ideas, and things I think I will use it for. It's a big deal (more on why in next post). But I am only sharing things that I am 99% sure will lead to stuff I will still be using Jev for in 60 days. That means I started with small, boring, but useful, stuff. - Fact-checking my scripts - Ranking my news feed - Finding the right text in PDFs - Checking citations - Grouping my review notes - Figuring out why agents fail (eval over traces) https://isaacflath.com/writing/six-things-i-tried-with-jev
Guidex.com
Six things I will still use Jev for
Isaac Flath’s shortlist: fact-checking scripts, ranking a news feed, finding text in PDFs, checking citations, grouping notes, and evals over agent traces.
Ricky Grannis-Vu
@RickyGrannisVu
Jev gives you a probability. You set two lines on it: above the top one it's a yes, below the bottom one it's a no, anything in the gap runs a third branch. It's just a normal branch, so you can handle uncertainty flexibly.
Guidex.com
Jev’s probability thresholds
Jev outputs a probability, two thresholds split the answer into yes, no and a third branch between them.
Jiayuan (JY) Zhang
@jiayuan_jy
一些关于 Jev (@typesafeai) 的 notes 研究了一天 Jev,带来的新鲜感迅速回落,这好像就是一个更快的通用分类器/决策器,LLM 完全可以做到。 而且因为不了解模型背后的参数规模(应该不会很大),世界知识可能不一定有常规 LLM 那么全,在这种情况下,它的复杂场景的决策结果是不是真的准还是要打一个问号的。 在一些有限解空间 + 低延迟要求的问题上,Jev 可能是一个很好的方向,加上形式化输出从程序上保证了正确性。 一个最常见的场景就是 Computer Use,太适合 Jev 了,网页的 Dom 元素是一个有限集,完全可以让 Jev 来做操作,这里面的 loop 相当于是 dom list -> jev action -> new dom list 这样的循环,但是这里 Jev 的推理能力和长上下文情况下的 computer use 效果如何还不清楚,目前还没有看到 benchmark(直观判断肯定是不如 GPT 6 Astra 的,但是速度太快了)。 昨天有尝试把 Pi Agent 中间的一些模块用 Jev 来重写一下,发现可以优化的地方不是特别多,tool using 部分的选择还是不能用 Jev 来代替,因为这不是一个有限集(每一步 tool using 实际上带了很多参数,比如 edit tool,会有具体的 lines 等参数,这部分是需要模型推理出来的,没有办法提前加到 Jev 的决策列表中),但是有一些地方是可以的,比如说 Compaction,可以让 Jev 快速做分类器(LLM 也能做,这里面差异化不太大)。 另外一类场景是依赖决策树逻辑的,例如: - 游戏 AI - 机器人 - 自动驾驶 而且这些场景对实时性要求比较高,传统的 LLM 来做这些事情的一个很大问题就是太慢了(plus 很大一部分场景是缺乏数据来训练的)。 目前正在做的两个 demo: 1. Poker AI,Poker 非常适合这个场景,而且决策树非常长 + 复杂,正在用 Jev 和其他模型做对抗式 battle。(btw 传统的 GTO Wizard 用来做训练实在是太难用了) 2. Pokemon VGC AI,这是严肃的宝可梦双打对战,有世界锦标赛的那种,每个赛季都有对应的规则,因为数据很全,所以非常适合用来研究 AI 的决策,plus 之前竟然没有一个很好的用来个人训练的 AI(这个做完了打算用这个 AI 实际在 Pokemon Champion 排位赛里用一下)。 这两个 demo 场景都是偏回合制的,其实用 LLM 也能做。
Guidex.com
Day-one notes from a skeptic, in Chinese
Jiayuan Zhang: a faster classifier that an LLM can also be, but a good fit for computer use, games and robotics where the options are known.
Common questions
- Can Jev do maths?
- No. TypeSafe recommends doing arithmetic in code and asking Jev only for the judgement.
- Can Jev count?
- Not reliably. Ask one yes/no question per item and sum the answers in code.
- Can Jev compare dates?
- No. Have it extract the date parts as Choices, then build and compare the dates in code.
- Can Jev write text?
- No. Generate or extract candidates elsewhere and let Jev pick one.
- Is Jev vulnerable to prompt injection?
- Yes. Text in the state can steer the answer. Screen input and output, write explicit criteria, and test before rollout.
- Does more context help Jev?
- No. Irrelevant state lowers accuracy, so filter first. The hard limit is 64k per request.
- Does Jev work in other languages?
- Yes, but with lower accuracy than English. Test on your own content.
Made with Jev is independent and not affiliated with TypeSafe AI. Every figure on this page is the one its author published, linked to where it can be checked.