Projects, posts and guides about Jev, the System One model from TypeSafe AI. Each entry links to its source and shows the cost and speed its author reported.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
the natural endpoint of coding agents is adhd maxxing
while my other agents are working i’ve been building a nintendo ds game with gpt-6 astra, fable 5.1 and jev
today i added perry the platypus so he can push you off a ledge
side quests are getting out of hand
Can #Jev really control embodied robots? @Anno_YanzheChen showed ShowHarness in #VLA. @Vanzllll told me: as DoF scales toward humanoids, discrete actions may struggle, while continuous policies scale better with #diffusion. Without AR, how do we model cause and effect? #Robotics
للمبرمجين 👨💻✨
Jev من TypeSafe AI يقدم فكرة مختلفة عن نماذج الذكاء الاصطناعي المعتادة.
بدل التركيز على المحادثات وتوليد النصوص، تم تصميمه لمساعدة التطبيقات على اتخاذ قرارات واضحة وسريعة داخل النظام ⚙️🤖
مناسب للمهام المتكررة مثل التصنيف، التوجيه، والأتمتة.
I think Jev and Laya were the missing pieces to algorithmic trading. Sure you can hardcode gates and adjust variables, but being able to add a real time decision maker within your algorith is a game changer. I can smell and taste an early retirement.
Stop settling for mediocre AI outputs. 🤖
Introducing slop-grader: The new Jev-AI CLI tool designed to audit text against your specific rulesets. It’s time to take control of the "slop." 🛠️
Here is the breakdown 🧵
#AITools #ProductivityAI
The useful engineering lesson is the async split: planning can continue while Jev selects the next action. That makes sense without assuming it proves AGI. We explain the roles using this demo in our 105-second video: x.com/vertr_ai/statu…
The small-LLM fallback for typing is an important detail: Jev chooses actions, while text generation stays with an LLM. We featured your 1x demo, with credit, in our short explainer of that division of labor: x.com/vertr_ai/statu…
Akshay Pachaar’s X article on Jev as a millisecond decision layer, and where it sits next to an LLM.
VertrAI
@vertr_ai
154ms vs 860ms is a useful routing result, scoped to these 200 synthetic classification cases—not a blanket ChatGPT comparison. We credited your latency and cost charts in our Jev video, alongside the game/browser demos: x.com/vertr_ai/statu…
ChatGPT thinks through the plan. Jev helps pick the next move. ⚡
Minecraft, browser automation, speed and cost—Jev explained in under 2 minutes. 👇
Featuring demos by @rronak_ and @gregpr07, benchmarks from @OpenRouter, and Jev by @typesafeai.
#Jev #AI
$10k/month just to decide which agent should answer.
not to answer. to decide.
group chat + @mentions: fast, cheap, terrible UX.
an LLM orchestrator: nice UX, 4-7s and $0.00046 a message.
@typesafeai jev: 0.29s median, $0.00002.
same decision, $20/month.
just shipped jev-studio v0.2.0
@typesafeai Jev in one pip install
- MCP tools for Choice / Noul / Score
- `jev` CLI now with dry-run provenance
- ready-made prompt libraries + slash commands for every cookbook
- Claude Code + Codex plugin manifests
pip install jev-studio
Most AI models are built to generate.
But Jev is trying something different:
Don’t generate. Decide.
Give it some context and a specific question, and it returns a structured decision with a probability.
I found this interesting for things like model routing, agent workflows, tool safety and RAG.
The idea is simple --Instead of calling a big LLM for every small decision, let your code handle the workflow and use a smaller model where human-like judgment is actually needed.
Honestly, this feels like an interesting direction for AI systems.
🚨 A LOCAL 421M MODEL JUST ATE CLOUD JEV ON SPEED
Laya is an open-source System 1 decision model that runs on your machine.
Laya: 86.5 decisions/sec, P50 ~9ms, score 46
Jev: 3.2 decisions/sec, 317ms API round-trip, score 1
Ships with:
— 421M params
— ~1GB inference memory
— Millisecond local calls, no network
— Typed decisions in one forward pass
— Apache-2.0 and self-hosted
While the cloud waits, local decides.
Agent memory is usually just an append-only Markdown file that grows forever, and most frameworks load the whole thing into context on every single run. That has been bothering me for a while, but never quite enough to fix it for our agents. Jev feels like it might be a low effort patch to this, by simply scoring each memory against the prompt first, then load only the relevant bits for that specific request.
Yes but not with Jev. I used another classifier, locally, just as fast, just as good.
Example: setup a "server" in python that loads the classifier model. Add a hook in Claude that fires on "pre-tool-use" and next time you ask Claude a random question like "how do I lint check a project?" and it tries to freelance and read your entire repo only to intoxicate itself and pollute its context, the classifier will slap its hand, say "no sir, you answer from knowledge" and deny the tool call.
I tested this yesterday with success, but need to refine it as it misfires.
Basically I was able to identify general queries, instances where I would ask something and Claude would rush ahead and run pip install without me asking or just write code instead of answering.
A classifier is hypercheap compared to a regular LLM and would catch all of these.
The model I used was DeBERT large. It's still stupid fast, I tested it on an Apple with M1 (regular) and you don't feel it running.
Created a slop detector extension with Jev
It scans all the post on the screen in the real time and classifies it on categories like scam, slop, clean, etc. Shows a minimal badge on the post with the confidence score.
Comment bellow if you want to try the extension.
Jev is live in New API now!
Choice. Score. Noul.
Structured judgments that drop straight into code.
Same TypeSafe SDK. Point it at your New API gateway.
One plugin. No rewrite.
newapi.pro/zh/plugins
Creators saying their new tool “killed” another is understandable. It’s marketing and rage bait.
But people retweeting it without even trying the tool? That’s the worrying part.
I’ve seen at least 10–20 “Jev killers”(@typesafeai) on my timeline already. Tried a few myself. Most weren’t even close.
We’re amplifying opinions before forming our own.
I keep seeing the same fix across agent stacks this week: teams are ripping the expensive model out of the middle of their decision loops.
TypeSafe shipped Jev in early access on September 15. Not a chat model. They call it a System One model: hand it a state and typed options, it returns calibrated probabilities, no text generation at all.
The mechanism is simple. Most agent loops burn a full LLM call on decisions that never needed generation: route to worker A or B, is this result relevant, approve or block this action. TypeSafe claims up to 200x faster inference and 400x lower cost on that class of decision. Ricker's own tests below land at 193x and 444x.
What happened next is the real story. Within four days, Cognition's Jared Palmer shipped Kev, a LoRA adapter on Qwen2.5-0.5B trained in 1 hour 45 minutes on a MacBook Pro, with an API close enough to point TypeSafe's own SDK at it. Kev now scales up to a 9B version that trails Jev by about 4.5 points on held out evaluation. Laya-MLX arrived the same week targeting millisecond decisions on Apple Silicon. A community leaderboard, JevBench, already ranks a dozen of these models.
The adoption signal convinces me this is not a toy. TanStack AI shipped a native decide() API for typed choices, scores and booleans. Beacon, an open source memory layer, uses Jev to score which coding sessions are worth turning into reusable lessons. Three teams, one primitive, inside a week.
RouteLLM out of Berkeley showed back in 2024 that routing simple queries to a cheap model cuts cost over 85% while holding 95% of GPT4 quality, and production semantic routers report 40 to 90% savings today. What changed is that the router stopped being a side project and became a shipped, benchmarked model category with a name.
The bottleneck this solves is real: every agent framework has a model sitting in a loop answering questions that never needed a sentence back. My read is that the decision layer becomes as standard a piece of the agent stack as the vector database became for retrieval, and whoever owns the default there owns a lot of the unit economics conversation for the next year of agent infrastructure.
https://x.com/0xRicker/status/2101705843200721203
1,000 AI papers sorted into 24 topics for $0.0585. Then Opus 5 graded the labels.
@nutlope's Jev paper map went viral, but the pipeline never shipped and the eval was "still running". So I rebuilt both and opened them.
The first judge run came back empty: Opus spent its whole budget thinking and answered nothing. Reasoning off, second run: it agreed with Jev on 85 of 100 papers, at 153x the cost and 1.9s against 57ms per paper.
The 15 misses are not random. One number Jev already returns tells you which labels to recheck.
Cheap models sort. Expensive models audit only what the cheap one flags.
Repost if you classify anything at scale, because the eval rows are public and anyone can rerun them with their own judge in one command.
Code in the reply.
What people posted on X while they built with Jev. Every card plays the original video or shows its images here, so you see the demo before you open the thread.