Skip to content
Made with Jev

Can Jev do maths, dates and counting?

No, and TypeSafe says so itself. Its Jev 1.13 jaggedness page, last reviewed 2 Oct 2026, lists 9 failure modes. Every one has the same fix: code computes, Jev judges. Source: TypeSafe Jev 1.13 jaggedness.

Updated 8 Oct 2026 · by Made with Jev

In short

  • Arithmetic, counting and date comparison belong in code.
  • Jev picks among options; it doesn't write. Extract candidates first.
  • More state is not better: irrelevant detail costs accuracy.
  • Text in the state can steer the answer, so screen it and test.
  • Confidence is calibrated across many answers; it isn't a guarantee for any single one. Source: TypeSafe System One.

The 9 limits, at a glance

Most of these start as a question worded the wrong way. How to write one that avoids them: Jev question types.

LimitWhat goes wrongDo this instead
Literal readingBoundary cases and implicit conditions are unreliableState the exact condition, put boundary cases in the criteria, split interpretation into two literal questions
Maths and numbersArithmetic, counting, hex and RGB comparisons struggleCompute in code; ask one yes/no question per item; convert numbers to named buckets
DatesDate comparison and extraction are unreliableAsk Choice questions for date parts, build the date in code, do every comparison in code
Indirection“Judge the field the user asked about” failsGive direct instructions and name the state field the question is about
Large state and contextIrrelevant detail lowers accuracyFilter before sending, or run a Noul relevance pass first
Adversarial content and prompt injectionText in the state can move the answerWrite explicit criteria, test edge cases, screen input and output in a separate request
Contradictory instructions and criteriaWhen criteria invert the instruction, jev-1.13 might get confusedMake the criteria extend the instruction, never invert true/false
Option orderjev-1.13 can lean toward the first option in some casesReorder the options and check the answer stays the same
Writing (generation)Jev cannot generate text, code or structured dataExtract candidates with regex or an LLM, then let Jev choose

1. Literal reading

Jev reads the state and the question literally. Boundary cases and implicit conditions are unreliable.

Fix: State the exact condition, put boundary cases in the criteria, and split interpretation into two literal questions that code combines.

Source: TypeSafe: Literal reading.

2. Maths and numbers

Arithmetic, counting and number comparisons are unreliable; do the maths in code.

Counting: Loop over the items in code, ask one Noul per item, and sum in code.

Counting pattern
let count = 0;
for (const item of items) {
  const { noul } = await jev(item, {
    matches: { type: "noul", instructions: "This item matches the criteria" }
  });
  if (noul > threshold) count++;
}
return count;

Numbers as text: Hex and RGB comparisons are unreliable, so convert in code to a value or a named bucket. See Jev images for colour names.

Score interpolation: Don't try to recover exact magnitudes from a Score; only threshold it.

Source: TypeSafe: Maths and numbers.

3. Dates

Date comparison and extraction are unreliable.

Fix: Ask Choice questions for the date's parts, including a “not stated” option; code builds the date and does every comparison.

Source: TypeSafe date extraction cookbook.

No build in this directory does this yet.

4. Indirection

Instructions like “Judge the field the user asked about” fail.

Fix: Give direct instructions and name the state field the question is about.

Source: TypeSafe: Indirection.

5. Large state and context

Irrelevant detail lowers accuracy.

Fix: Filter before sending, or run a Noul relevance pass first.

Limits: 64k per request, of which 32k is state plus the longest question. Gateways list 32k.

Sources: TypeSafe: Large state and context, classifying RAG passages cookbook, TypeSafe models.

6. Adversarial content and prompt injection

The doc says state “is data, and jev-1.13 does not treat it as hostile by default”, so injected text can move the answer.

Fix: Write explicit criteria, test edge cases before rollout, and screen input and output in a separate request.

Note the tension: A screen built on Jev can itself be steered, so test it too.

Sources: TypeSafe: Adversarial content, LLM guardrails cookbook, classifying RAG passages cookbook.

The jev-as-judge build's author cautions that it “is not a proven defence against injection”.

7. Contradictory instructions and criteria

When criteria invert the instruction, jev-1.13 might get confused.

Fix: Make the criteria extend the instruction, and never invert true/false between them.

Source: TypeSafe: Contradictory instructions.

8. Option order

In some cases jev-1.13 leans toward the first option.

Fix: Reorder the options and check the answer stays the same.

Source: TypeSafe: Option order.

9. Writing (generation)

Jev cannot generate text, code or structured data.

Fix: Extract candidates with regex or an LLM, then let Jev choose.

Source: TypeSafe pre-parsed value extraction cookbook.

The haiku-correction build shows that the claim “Jev produced the hiragana” doesn't hold.

Two more limits from the docs and testers

Non-English text: Accepted, but English is the primary training language and other languages, CJK included, score lower.

Source: TypeSafe models: Language support.

Alternative: Laya publishes a multilingual checkpoint (the author's claim).

Knowledge you didn't supply: jev-capability-atlas reports that Jev is accurate when the answer is in the state, and often confidently wrong when it isn't (an independent finding).

Images: Jev is text only, so put OCR or a vision model first. See Jev images.

Fine-tuning: You cannot fine-tune Jev itself. What people train instead: Jev fine-tune.

Literal reading

Boundary cases and implicit conditions are unreliable. State the exact condition, put boundary cases in the criteria, and split interpretation into two literal questions that code combines.

SkillBenchmarks and evals

tenbin

Split a judgment into Choice, Score and Noul, lint it, then measure it.

Shingo Imota

4

GitHubBenchmarks and evals

jev-evaluation

Nine experiments and 28 predictions, all fixed before any data.

Will Kelly

0123,805

Maths and numbers

Arithmetic, counting and number comparisons are wrong. Loop over items in code, ask one Noul per item, and sum in code. Convert hex and RGB to named buckets.

ラプター | 高専 ロボコン ビジコン

@Raptor_zip

型付き意思決定AIモデル「Jev」で双腕ロボを動かした結果まとめてみた🦾 3層制御のうち意思決定(2層)をJevに任せ、IKや物理計算はコード側に分離する構造を作りました。500msの応答と0.5円/試行で早くて激安。

XRobotics and devices

A dual-arm robot with Jev in the middle layer

Response
~500 ms
Cost per attempt
~¥0.5

Justin Schroeder

@jpschroeder

I rebuilt Tesla Full Self Driving with Jev in less than an hour. This model is a total unlock.

GitHubGames and real time

Pick

JevPilot

Peak rate
4/s
Built in
<1 hour

Saeed

@stringsaeed

built a highlighter on jev paste any language → my code tokenizes → jev names the lang, colours every word, then says which of 9 lint rules fire and where nine rules in code. jev just answers. near instant lab.saeed.sh/highlight

Live siteCoding and code review

A syntax highlighter on Jev

Lint rules
9

Ackerman

@Yarilo7brigada

98% of my feed is junk. Now I don’t even see it I built a filter that reads my feed for me. It runs on Jev - a model that doesn’t generate text; it just makes binary decisions, in milliseconds and for pennies. For every post, it evaluates four questions: Is it relevant to my niche? Does it provide real value? What format is it (breakdown, news, ad, meme)? And does it make loud claims with zero proof? Out of 1,000 posts, only 20 remained. 97 milliseconds per post. 1.2 cents for the entire morning. The most frustrating takeaway, nearly half of my feed was ads and memes, not the creators I originally followed for substance. An hour of mindless scrolling turned into five lines over breakfast.

XSocial feeds

A feed filter that kept 20 of 1,000 posts

Per post
97 ms
Morning cost
$0.012
Kept
20 / 1,000

GitHubSocial feeds

LinkedIn NoSlop

A Chrome extension that blurs low-value LinkedIn posts.

sushrutb17

0

Indirection

Instructions like 'Judge the field the user asked about' fail. Give direct instructions and name the state field the question is about.

GitHubCoding and code review

Jev Review

A staged code-review workflow with a local dashboard.

Dev Agrawal

675

Large state and context

Irrelevant detail lowers accuracy. Filter before sending, or run a Noul relevance pass first. The hard limit is 64k per request.

NO1ennn

@N01ennn

this is pure f*cking treasure A Stanford AI research group finally drew the perfect RAG system: retrieval, Jev and agents in one loop, and it fixes the 3 things that break every RAG app: > the LLM reads 20 passages when only 3 matter > it answers questions your docs can't answer > it trusts whatever text it retrieves here's how it runs: > a lead agent sends the query > hybrid search pulls the top 20 candidates, dense + keyword > ONE Jev request scores all 20 + 2 gates: answerable? injection? > only passages above 0.6 reach the writer agent > a second Jev call checks every claim against its source > grounded answer, with citations and when the docs don't have it: > answerable fails, the writer never runs > a researcher agent rewrites the query and retries once > still nothing? "not in the docs". zero tokens spent on a guess retrieval casts the net. Jev decides what's real. agents do the work save this before you build your next RAG

XSearch

A RAG loop where Jev decides what is real

Candidates scored per request
20
Passage threshold
0.6

Coach Shweta Bajaj

@shwetabjaj

Jev Skill Suggestion for Claude Code is a smart idea. Instead of loading every skill into context, Jev decides which one is actually relevant and injects only that. My takeaway: better context hygiene, less clutter, and potentially more efficient agent workflows. @typesafeai @vercel

GitHubContext and memory

Jev Sift

Classify first, read selectively: an agent plugin and MCP tool.

Kush Bhuwalka

46

Adversarial content and prompt injection

Text in the state can steer the answer. Screen input and output, write explicit criteria, and test before rollout. Note the tension: a screen built on Jev can itself be steered.

GitHubAgents and browsers

dsh-jev-tools

Judgment, not generation, inside DeepSeek Harness.

HorusEyes

12

GitHubSDKs and integrations

jev-mcp

An MCP server with claim checks, screening and ranking built on Jev.

Joey Kudish

506

NO1ennn

@N01ennn

this is pure f*cking treasure A Stanford AI research group finally drew the perfect RAG system: retrieval, Jev and agents in one loop, and it fixes the 3 things that break every RAG app: > the LLM reads 20 passages when only 3 matter > it answers questions your docs can't answer > it trusts whatever text it retrieves here's how it runs: > a lead agent sends the query > hybrid search pulls the top 20 candidates, dense + keyword > ONE Jev request scores all 20 + 2 gates: answerable? injection? > only passages above 0.6 reach the writer agent > a second Jev call checks every claim against its source > grounded answer, with citations and when the docs don't have it: > answerable fails, the writer never runs > a researcher agent rewrites the query and retries once > still nothing? "not in the docs". zero tokens spent on a guess retrieval casts the net. Jev decides what's real. agents do the work save this before you build your next RAG

XSearch

A RAG loop where Jev decides what is real

Candidates scored per request
20
Passage threshold
0.6

GitHubBenchmarks and evals

awesome-jev-robustness

Tests, calibration audits and failure-mode studies of Jev.

Yifan-Lan

5

Contradictory instructions and criteria

When criteria invert the instruction, the answer is wrong. Make the criteria extend the instruction, and never invert true/false between them.

SkillBenchmarks and evals

tenbin

Split a judgment into Choice, Score and Noul, lint it, then measure it.

Shingo Imota

4

Option order

In some cases jev-1.13 leans toward the first option. Reorder the options and check the answer stays the same.

GitHubBenchmarks and evals

awesome-jev-robustness

Tests, calibration audits and failure-mode studies of Jev.

Yifan-Lan

5

Writing (generation)

Jev cannot generate text, code or structured data. Extract candidates with regex or an LLM, then let Jev choose.

たく|ガチのCopilot達人

@taku_ai_case

これはめちゃくちゃ勉強になりました。 Jevに回答させるのではなく、CodexやClaude Codeへ渡す記憶候補を選ばせる。文章を書けないAIの特性を、ここまで実用的な仕組みに落とし込む発想がすごいです。

Justin Schroeder

@jpschroeder

I rebuilt Tesla Full Self Driving with Jev in less than an hour. This model is a total unlock.

GitHubGames and real time

Pick

JevPilot

Peak rate
4/s
Built in
<1 hour

GitHubSDKs and integrations

jev-mcp

An MCP server with claim checks, screening and ranking built on Jev.

Joey Kudish

506

yuri

@yurinakanishi33

昨日のJevハッカソンで発表した「Jevを使った俳句生成」について訂正があります。 発表では「JevのChoiceで、ひらがなを1文字ずつ生成している」と説明しましたが、間違いでした。 実際には、季語・名詞・助詞・動詞・副詞・結びの定型句などの単語をコード側であらかじめ用意し、それらを組み合わせて5音・7音の完成行候補まで先に作っていました。 Choiceに渡していたのは、その時点の文字列から、いずれかの完成行に到達できる「次の1文字」だけです。 さらに、候補が1つならコードで確定。複数ならChoiceが返した確率の上位3候補を再正規化し、コード側でもう一度抽選していました。そのため、Choiceが選んだ文字を必ず使っていたわけでもありません。 つまり、単語や文章の道筋をコード側で先に用意したうえで、Choiceが返した確率を使ってコード側で分岐する実装でした。申し訳ありません。 添付動画は、コード側で用意していた語彙と単語のつながりに関する制約を外した版です。残した制約は、5・7・5の区切りと、使用できる文字を1字1音として扱える73字に限定する文字種制限だけです。 毎回、ここまでの句、季節・状況・気分などの文脈、残りの音数、73字すべてをChoiceに渡しています。Choiceが返した全73候補の確率をコード側で再正規化し、その確率を重みに次の1字を抽選。これを17回繰り返しています。 結果は......全くできていませんねw #aimeetup

Where these come from

Guideyoutube.com/@GaryExplains

Jev: fast and cheap, but there is a caveat

Gary Explains on the catch: Jev understands natural language but it is not an LLM. It answers with structured values and a confidence level.

Jason Zhu

@GoSailGlobal

拿 Jev 做搜索重排,我先泼一盆冷水:单独用,它没打赢向量检索 TypeSafe 的 Jev 这阵子很火,一堆项目拿它做重排。我们在 Agent Skills Hub 的 33,047 条目录上认真测了一次,164 条中英文真实查询,9,831 对分级标注,整套只花了 2.6 美元 三个结论 01|单独重排,约等于没赢 Jev 重排 bge-m3 的前 30 条,NDCG@10 只多了 0.012,置信区间跨过零。MRR 和前三命中率倒是涨得明显,它很会把最强的那一条顶到第一,后面几条基本是重新洗牌 02|裁判偏差,被我们量出来了 Jev 自己也参与了打标注,这就是循环 只用 Jev 当裁判,它领先 0.053 两个裁判合并,领先 0.012 只用 Haiku 当裁判,反而落后 0.028 同一组比较,换个裁判结论直接翻面。所有涉及 Jev 的结论,我们只认 Haiku 那一列 03|真正稳赢的是融合 把 Jev 和 bge-m3 的排序做 RRF 融合,NDCG@10 到 0.864,比纯向量高 0.064 到 0.116,三种裁判下都成立。代价是每次查询多一次 API 调用,约 0.0002 美元 顺便暴露了我们自己的问题:Hub 线上 CLI 用的关键词排序只有 0.609,短板是召回。相关结果有一半压根没进候选,后面怎么重排都救不回来 数据、标注、每条查询的得分全部开源,不用 API key 就能复现打分

Guidex.com

Jev as a reranker: an honest negative

Jason Zhu tested Jev reranking on 164 real queries. Alone it did not clearly beat vector search; fused with it, it did. In Chinese.

Isaac Flath

@isaac_flath

I've been using Jev by @typesafeai Here's the six things i've tried and am confident I'll still use Jev for 60 days from now. There's many more experiments, ideas, and things I think I will use it for. It's a big deal (more on why in next post). But I am only sharing things that I am 99% sure will lead to stuff I will still be using Jev for in 60 days. That means I started with small, boring, but useful, stuff. - Fact-checking my scripts - Ranking my news feed - Finding the right text in PDFs - Checking citations - Grouping my review notes - Figuring out why agents fail (eval over traces) https://isaacflath.com/writing/six-things-i-tried-with-jev

Guidex.com

Six things I will still use Jev for

Isaac Flath’s shortlist: fact-checking scripts, ranking a news feed, finding text in PDFs, checking citations, grouping notes, and evals over agent traces.

Ricky Grannis-Vu

@RickyGrannisVu

Jev gives you a probability. You set two lines on it: above the top one it's a yes, below the bottom one it's a no, anything in the gap runs a third branch. It's just a normal branch, so you can handle uncertainty flexibly.

Guidex.com

Jev’s probability thresholds

Jev outputs a probability, two thresholds split the answer into yes, no and a third branch between them.

Jiayuan (JY) Zhang

@jiayuan_jy

一些关于 Jev (@typesafeai) 的 notes 研究了一天 Jev,带来的新鲜感迅速回落,这好像就是一个更快的通用分类器/决策器,LLM 完全可以做到。 而且因为不了解模型背后的参数规模(应该不会很大),世界知识可能不一定有常规 LLM 那么全,在这种情况下,它的复杂场景的决策结果是不是真的准还是要打一个问号的。 在一些有限解空间 + 低延迟要求的问题上,Jev 可能是一个很好的方向,加上形式化输出从程序上保证了正确性。 一个最常见的场景就是 Computer Use,太适合 Jev 了,网页的 Dom 元素是一个有限集,完全可以让 Jev 来做操作,这里面的 loop 相当于是 dom list -> jev action -> new dom list 这样的循环,但是这里 Jev 的推理能力和长上下文情况下的 computer use 效果如何还不清楚,目前还没有看到 benchmark(直观判断肯定是不如 GPT 6 Astra 的,但是速度太快了)。 昨天有尝试把 Pi Agent 中间的一些模块用 Jev 来重写一下,发现可以优化的地方不是特别多,tool using 部分的选择还是不能用 Jev 来代替,因为这不是一个有限集(每一步 tool using 实际上带了很多参数,比如 edit tool,会有具体的 lines 等参数,这部分是需要模型推理出来的,没有办法提前加到 Jev 的决策列表中),但是有一些地方是可以的,比如说 Compaction,可以让 Jev 快速做分类器(LLM 也能做,这里面差异化不太大)。 另外一类场景是依赖决策树逻辑的,例如: - 游戏 AI - 机器人 - 自动驾驶 而且这些场景对实时性要求比较高,传统的 LLM 来做这些事情的一个很大问题就是太慢了(plus 很大一部分场景是缺乏数据来训练的)。 目前正在做的两个 demo: 1. Poker AI,Poker 非常适合这个场景,而且决策树非常长 + 复杂,正在用 Jev 和其他模型做对抗式 battle。(btw 传统的 GTO Wizard 用来做训练实在是太难用了) 2. Pokemon VGC AI,这是严肃的宝可梦双打对战,有世界锦标赛的那种,每个赛季都有对应的规则,因为数据很全,所以非常适合用来研究 AI 的决策,plus 之前竟然没有一个很好的用来个人训练的 AI(这个做完了打算用这个 AI 实际在 Pokemon Champion 排位赛里用一下)。 这两个 demo 场景都是偏回合制的,其实用 LLM 也能做。

Guidex.com

Day-one notes from a skeptic, in Chinese

Jiayuan Zhang: a faster classifier that an LLM can also be, but a good fit for computer use, games and robotics where the options are known.

Common questions

Can Jev do maths?
No. TypeSafe recommends doing arithmetic in code and asking Jev only for the judgement.
Can Jev count?
Not reliably. Ask one yes/no question per item and sum the answers in code.
Can Jev compare dates?
No. Have it extract the date parts as Choices, then build and compare the dates in code.
Can Jev write text?
No. Generate or extract candidates elsewhere and let Jev pick one.
Is Jev vulnerable to prompt injection?
Yes. Text in the state can steer the answer. Screen input and output, write explicit criteria, and test before rollout.
Does more context help Jev?
No. Irrelevant state lowers accuracy, so filter first. The hard limit is 64k per request.
Does Jev work in other languages?
Yes, but with lower accuracy than English. Test on your own content.

Made with Jev is independent and not affiliated with TypeSafe AI. Every figure on this page is the one its author published, linked to where it can be checked.