822 projects, posts and guides about Jev, the System One model from TypeSafe AI. Each entry links to its source and shows the cost and speed its author reported.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
Found this pretty interesting repo: herdr + jev = agent router
You give it a coding task, and it routes it to the right agent + model + reasoning effort based on the task and your available quota.
Supports Cursor, Claude Code, Codex, and OpenCode.
Basically a load balancer for your coding agents.
https://github.com/nidhi-singh02/agent-router
Thanks @nidhisinghattri
We at @Parall_HQ are open-sourcing a universal agentic monitor built around Jev (@typesafeai). It’s a fun little toy, but a lot handier than expected.
➡️ github.com/parall-hq/jeva…
Catching AI agents that secretly collude usually means reading their minds.
I just read their chat instead, using Jev by TypeSafe AI
It named the colluders 85% of the time.
Built for the AI Swarm Dynamics Hackathon by @grove_research & @aidigest_. Thanks for letting people join online, I really enjoyed it!
Jev and Laya have been getting a lot of buzz lately. Can we use them in ComfyUI too?
Yes. I built a node pack that branches a workflow on a plain-English question, like "Too graphic for kids?"
Runs locally with Laya, or flip one dropdown to Jev.
github.com/dante01yoon/Co…
Agent: "done" ✅
UI: looks fine ✅
Database: the free plan just published a 2nd project and took the 1st one offline ❌
I stopped trusting my AI agent's "done". Now the database is the judge.
jev-check, an open-source skill for Claude Code and Codex 👇
Built an open-source 𝕏 filter powered by uprising decision models!
Instead of slow, pricey LLMs, it leverages @typesafeai Jev & @CloudflareDev Clef to detect politics, soccer & crypto noise in <100ms.
Zero tracking & 100% local. Links below 👇
Built a tool to help fill out your LA County ballot for Nov 3: take a quiz, enter your address, get a match for every contest, from governor to water board.
Claude did the research, Jev the scoring. Open source: check the logic or fork it.
la-ballot-match.vercel.app
Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta.
The Decisions API makes decisions up to 10x faster than GPT-6 Luna through the Responses API.
You can now train your own Decision model like Jev locally!
We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM.
Turn any LLM like Qwen3.8, Gemma 4 into decision models with our open-source Unsloth repo.
We fine-tuned with a Clef head using Unsloth and LoRA (r=64) for one epoch, increasing downstream accuracy from 30–37% to 78%.
GitHub: https://github.com/unslothai/unsloth
Guide and Notebooks: https://unsloth.ai/docs/basics/train-your-own-decision-model-with-unsloth
Announcing d1 with vision. 👁️👁️ Our first decision model now supports images, text or both as inputs. We tested d1 against GPT-6.1 Sol and Claude Opus 5.5 on six real applications, from filtering support tickets to inspecting circuit boards. d1 matches or beats GPT-6.1 Sol on four of them. It costs 19x to 200x less than both models and answers significantly faster on every task.
> probabilities for yes/no, choice, or score questions
> one forward pass, without generating tokens
> text decisions in 200 to 300 ms
> Liquid API: https://console.liquid.ai
🧵
d1-3B ranks first among models under 10B on the Decision Index v0.2.1, a benchmark for structured decision-making.
Built from LFM2.5-VL-3B, it makes decisions from text and images in one pass.
Use it for reranking, agent guardrails, and visual inspection.
2/
.@Cloudflare's decision models are available on Ollama!
Decision models allow you to classify an image, label a bug report or route a support ticket to the right team.
Clef (27B):
ollama pull clef
Clef Flash (9B):
ollama pull clef-flash
pplx-decider-v1.1-27b, our updated open weights multimodal decision model, is now available.
It scores the highest on the new @huggingface Decision Index 0.3 benchmark. pplx-decider-v1.1-27b costs half as much as v1, at $0.02 per million input tokens.
huggingface.co/spaces/multimo…
New: Decision Model Rankings!
See the spend share and token share of different decision models and task types, including Relevance, Correctness, Instruction Following and more.
@typesafeai is leading all categories today.
Pokémon FireRed’s Elite Four + Champion, cleared in one shot - with decision-making powered by Perplexity’s Decisions API.
The actual run stats:
592 ms median API response
987 ms p95 - 96.4% of responses under one second
$0.028 estimated inference cost, 137 live API calls
the jev Decision Index v0.3 is here 🎯
featuring a way stronger set of benchmarks, now with a strong private half 🔒 so the ranking tracks skills beyond benchmarks
and, as promised, vision 👁️
huggingface.co/spaces/multimo…
Found an interesting open-source project: Bud Decision Studio.
Think of it as #LMStudio for decision models. Run #Jev -like System One models locally, experiment with them, compare models, evaluate decisions, and expose them through an API.
Repo : github.com/BudEcosystem/B…
Launching Decision models on @useRouterPlus. You can now use various system 1 decision models such as Jev, @perplexity_ai Decider, @bespokelabsai Nimble, Mercury Decide, Clef, Kev 4B and Decider 2B with 3 models free !!
app.routerplus.com/playground
I benchmarked all the top Decision Models against Jev and there is a CLEAR WINNER.
Spoiler alert 🚨: GLiDE wins by a big margin over Jev, Cloudflare, and Perplexity's Decision Models.
You also have to understand that the decision index is the M.O.A.B for decisions.
M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam. It is the source of truth and the other benchmarks are just noise.
GLiDE was never trained on a single row of Decision Index so it's not benchmaxxed at all.
Also their " clef-flash runs ~13x faster than Jev" claim is dishonest because theyre comparing their local inference to the Jev API latency.
Sage by @LevantoLabs is the first decision model, launched in July, and currently the best on the market.
It wins across all the evals proposed by @perplexity_ai's CEO (nice job @AravSrinivas!), and it's first on JevBench, beating Jev and Google's models on every dimension.
BUT frankly, benchmarks are not what excites us.
I'm BORED by the dozens of new decision models launching every day with a tiny +0.02 on some eval. They all look the same.
The magic is somewhere else.
There are thousands of new things that could be invented in the space of decision models and models for machines. The invention space is huge!
After launching Sage, we also pioneered:
- The first guardrails and model routers built on decision models
- Automated context engineering with "grounding" (automatic web search when confidence is low). Decision models are only as smart as their context, especially because they can't make tool calls.
- Reasoning. In many use cases you don't know whether your question is System 1 or System 2. With reasoning, the model is as fast as a System 1 model when the question is System 1, and takes more time when it's needed. This enables hard, multi-step decisions. Our clients just want the correct answer, and in some use cases they're happy to wait 0.5s more. This will soon become the market standard (GLiDE/@george_onx and @TypeLLM already followed).
And *so* much is yet to come. Models for machines: it's just the beginning!
We're cooking. Follow @levantolabs for new product primitives that will surprise you.
Open innovation is HEALTHY.
Competition is HEALTHY.
Benchmark-race is BORING.
Product inventions are FUN.
Go go go!!!
I wanted to understand @typesafeai's Jev and where open source scores against it
The popular OpenJev's like Von were good but the tiny GLiNER models (some only 74M params) performed phenomenally — even beating Jev at Doom
Full breakdown...
You'd think an AI would turn when the option says THE SNAKE DIES.
Three new decision AIs. One game of Snake.
Amazon Strands Decider 2B.
Cloudflare Clef-flash.
Perplexity pplx-decider.
One of them keeps crashing into the thing it was just warned about.
Every move is a real call to the model.
Every option tells it what's there.
Game, snakes and video. All code, written by Claude Opus 5.5.
Jev now has an open competitor from an actual lab
Maincode just released Matilda Jev: a 26B decision model trained on their own infra in Australia.
it doesn’t generate text. give it a decision and you get probabilities for every option in one pass.
they claim it beats the original Jev on 27/43 benchmarks and the whole thing is open source under Apache 2.0.
Open decision models just keep leveling up.
Pour one out for @CompleteSkeptic. OpenAI, Cloudflare and Perplexity all have decision models now, and ruby_decision_model 0.2 talks to all of them, Jev included, through one Ruby client. Same support ticket, five models, results inside. x.com/i/article/2107…
I put OpenAI’s new Decisions API through the same 1,565-email benchmark I’ve been running on Jev, Perplexity and Clef.
6 models from 4 companies now.
Perplexity v1.1 came out best overall 🚀
OpenAI was the fastest⏱️
But the confidence behavior is still the most interesting part to me.
Jev was the only model that said “≥99% sure” hundreds of times without getting one of those answers wrong.
Its the cocky guy in the class who is also right most of the time🥲
Clef has massive imposter syndrome! it's much better than its confidence suggests.
And Perplexity changed a lot in just 4 days.
v1 gave just 3 answers at ≥99% confidence. Four days later, v1.1 gave 323.
It also got slightly more accurate and cut the price in half.
kudos @AravSrinivas and @perplexity_ai for making such massive strides and shipping so quickly!
There's no doubt now that @typesafeai has opened a pandora's box and as @CompleteSkeptic said, forked a new branch in the AI timeline with these decision models!
Meet Kev - open source version of Jev
Everyone is talking about Jev right now
But it is a proprietary model and you dont' know what goes behind the scences
Kev built on Qwen3.5 based on Jev's architecture. You can use pretrained weights or train your own
Open source and 100% local.
Shanghai AI Lab's Intern-Decision comes in 0.8B, 2B and 4B sizes and returns choices, scores and yes/no answers with probabilities; MetaX shipped Day-0 support. x.com/i/article/2107…