Timothy Kassis
@TimothyKassis
We built a really nice free tool on top of @typesafeai Jev to see how rigorous an academic paper is.
XBenchmarks and evals
Updated Oct 3, 2026
Projects, posts and guides about Jev, the System One model from TypeSafe AI. Each entry links to its source and shows the cost and speed its author reported.

Diogo Almeida
@CompleteSkeptic
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
Timothy Kassis
@TimothyKassis
We built a really nice free tool on top of @typesafeai Jev to see how rigorous an academic paper is.
XBenchmarks and evals
Yannis
@yannnis
Enough is enough! Is it that LinkedIn sucks or what? Let's settle this. I go viral on LinkedIn for slop (attached) Then I spend 2 days to build a JEV product to scan Reddit for leads. And get tumbleweed!?! How??
XSales and leads
Ackerman
@Yarilo7brigada
98% of my feed is junk. Now I don’t even see it I built a filter that reads my feed for me. It runs on Jev - a model that doesn’t generate text; it just makes binary decisions, in milliseconds and for pennies. For every post, it evaluates four questions: Is it relevant to my niche? Does it provide real value? What format is it (breakdown, news, ad, meme)? And does it make loud claims with zero proof? Out of 1,000 posts, only 20 remained. 97 milliseconds per post. 1.2 cents for the entire morning. The most frustrating takeaway, nearly half of my feed was ads and memes, not the creators I originally followed for substance. An hour of mindless scrolling turned into five lines over breakfast.
AI Insider
@TheAIInsiderN
JEV + OPUS 5.5 IS INSANE. I built an AI system that analyzes an X post before you publish it. Paste any X link. Press Start. Get a virality score in minutes. The entire production version was built in 9 minutes. Here’s what happens: Jev analyzes 800 viral posts as a live baseline Groups them by hook type: demo, launch, proof, contrarian take Runs 12 structured checks: hook, numbers, media, CTA and more Finds the 5 most similar viral posts Opus 5.5 explains why those posts spread Then your post receives: Virality score out of 100 Estimated likes and views Percentage of the viral baseline it beats Direct comparison with similar winners It works on drafts too. So instead of posting and hoping, you can identify weak hooks, missing proof and unclear CTAs before anyone sees them. Hover over any result to inspect the original post, its media, every check and the final score. JEV DECIDES FAST. OPUS EXPLAINS WHY. #JEV
Avid
@Av1dlive
Jev + GPT-6 Astra just built the most TERRIFYING AI trading setup on the internet... [this article covers 90% of what is required to build quant-level systems] /1 GPT-6 Astra reads the order book, the tape and 3 correlated futures /2 Jev turns the signal into a trade and checks the risk limit /3 computer use clicks the order screen, no broker API needed /4 the full loop runs in 6 ms, signal to fill steal this setup in the article below👇
Proliquid
@proliquid_xyz
Did you know that we have @typesafeai Jev AI giving you the sentiment on the news? Don't miss the next ethereum:0xfaba6f8e4a5e8ab82f62fe7c39859fa577269be3 like move. One click long on bullish news only on app.proliquid.xyz/terminal 🦈
XTrading and markets
Jetwani Avinash
@jetwaniavinash
First thing I built on Jev: a memory gate for Claude Code. Every message gets one question: is this worth remembering? 0.3s later it's either skipped or saved as one line in JEVMEM.md in the repo. Open source, link in my pinned post.

Prasad Pilla
@prasad_pilla
AGREED. team built an internal tool using Jev to qualify candidates, and Jev itself is genuinely impressive at what it’s designed todo. but recruiting is hard. there are too many variables that don’t fit neatly into a rubric: context, trajectory, ownership, quality of work, potential, and the signals you only catch by actually digging into someone’s work tbh https://x.com/idleshubh/status/2102429846660079971?s=20
XSales and leads
morph
@morpphhhaw
TypeSafe founder Diogo Almeida introduced Jev, and the field I would inspect first in any agent built around it is not model. It is state An agent can have room for a long conversation and still miss the one line that matters. A refund request needs the current order, the charge record and the policy that applies today. It does not need twelve earlier attempts to sound helpful Here is the state packet I would put in front of Jev: step 1 → name the decision the code needs to make now, not the entire job the agent was given step 2 → pull facts from the system of record: order status, amounts, timestamps and permissions step 3 → include the customer's actual words as evidence, without asking Jev to reconstruct the whole conversation step 4 → state the constraints that can change the route, such as a refund window or an account hold step 5 → build the available actions from live code, so a closed account never appears as a valid destination step 6 → put the question in the question field. State is evidence, not a second prompt hiding instructions step 7 → leave the sums and date arithmetic to code, then send Jev the result it needs to judge step 8 → keep the exact packet with the answer, because you cannot debug a decision from its label alone The point is not to compress everything until it looks clever. The point is to make it obvious which piece of evidence could change the answer The document below shows that packet on one page. The article goes further into shaping state without turning it back into a transcript
XContext and memory
OrcDev
@orcdev
TanStack AI just shipped subagents. Jev picks which models / agents should run for the task, in parallel, in sequence, or mixed. Best Jev use case I've seen so far. TanStack team cooked. 🔥
XRouting and model choice
rewind
@rewind02
built a full production app with opus 5.5 and jev... fed it a url, it screenshots the site, scores it against a list of known patterns, and ships a shareable report here's the actual pipeline: - gemini flash takes the screenshot, that's the only vision step needed - jev scores it against a fixed list of criteria in a fraction of a second, no text generation, just structured ratings - opus 5.5 writes the entire app, front to back, off a single planning prompt - product os turns the plan into a spec, then a roadmap, then working code, task by task - code review, security audits, and testing happen automatically as it builds, not after - a deploy checklist gets generated too, walks you through auth setup, hosting, env variables, step by step - cost to run per scan lands in fractions of a cent, both models chosen specifically because they're cheap at their job what actually makes this work: opus 5.5 handles everything generative, jev handles every decision that doesn't need generation, and neither model gets used for a job it's not built for that split is why the whole thing runs on pennies instead of frontier-model pricing pair that with a structured build system instead of one long freeform prompt, and a non-coder ships a live, secured, production app in an afternoon, not a sprint
aiasssistantstore
@aiassistantstor
CLM-8B Just Dropped: The New Open AI Model That Claims to Be Up to 9× Faster Than Jev for Agents
Abdullah
@Abdullah_Ops1
التجربة الثانية مع JEV 👇 سويت له Vibe Coder Launch Inspector تحط رابط مشروعك يفحصه ويطلع لك هل جاهز للإطلاق ولا يحتاج شغل يعطيك Score + المشاكل + الأدلة + وش اللي يحتاج تعديل والأحلى كل مشكلة معها Fix with Codex يطلع لك برومبت جاهز ترسله لكودكس #AI #JEV #Codex #VibeCoding
XCoding and code review
apolinario (poli)
@multimodalart
congrats on the release! i've included xor on the Decision Index 0.2 benchmark of jev-like models we don't have a vision benchmark yet, but the text-only version got 8th place, strong in arts & human taste x.com/multimodalart/…
PeterY
@yudeewxy
用JEV + Codex + TG Bot,开发了一个小工具,用于实时判断9个coin的价格走势和概率。交易信号来自前几个月开发的trading信号平台。同时给这个Bot授权了1万刀,看看他围绕这些信号来实时交易的表现会怎样。 开发过程: - JEV开通,获取API,配置到本地 - 配置TG Bot - 连接Binance和OKX的API - 信号聚合平台:http://pytrading.vercel.app - 或者找到trading view或者自己的一套交易因子 - Bot配置钱包,授权10K usdt,在OKX上进行交易 JEV的优势是:快速高效的反馈价格及进行交易行动、相比LLM大量节省Token
apolinario (poli)
@multimodalart
congrats on the release, i've added it to the jev decision index 0.2, it's the #2 open weight jev-like model 🥳 and the strongest in chess and for its size x.com/multimodalart/…
ImRobot
@th3nolo
Been building an agent-agnostic permission gate: Jev (TypeSafe) + hooks. Many harnesses copy Claude Code's hook format. Some barely have hooks at all. This makes me happy: agy in YOLO mode. chmod -R gets blocked, the agent asks me, I say yes, it runs once. Again? Blocked again.
XSecurity and abuse
Cartwright
@CartwrightApp
Built Cortex, a local MCP server: Claude plans, a fast layer does the clicking. Same Mac-app task: computer use 23 s → Cortex 0.8 s One action: ~5 s Claude round → 0.25 s, no model 24 tools: browser, Mac apps, GitHub search, safety gate Measured on my Mac. @typesafeai #JEV
TactiX Trading Panel
@TactiXPanel
The Jev-powered AutoScalper passed all Tests during the Testnet Live Test with an average win rate of 82% trading on the 1m chart, using a 120 candle lookback window. So 2h structures get scalped top-to-bottom and bottom-to-top Now its time to aim for mainnet deployment!
Awais.
@abbas_kazmi066
plenty of use-cases for JEV. here's one: Hover Explanations. This simple demo testing cost me around: $0.006. Definitely very cheap and fast.
Cuth
@ItsCuthulhu
Now that I've collected so much data on System One models on my DGX Spark. Lets see if I can beat Jev on size, speed, and accuracy. It will be MIT for anybody to build on. Everything I do on here is open source.
XOpen source
Browser Use
@browser_use
Luna (planner) -> Jev (actor) playing poker and winning
XGames and real time
Chris Brownridge
@chrisbrownridge
another @treg_ai /Jev demo to analyze data SUPER fast used treg to pull meta ads and then Jev to teardown the landing pages so you can understand where brands are sending traffic to categorizes the type of page, offers they're using and language used to sell also shows ad creative launch velocity, creative mix takes a few seconds to do it all end to end.
XAds and marketing
OpenMed
@OpenMed_AI
Four typed questions across four authored fictional notes, labels written before the calls. Jev 16/16, Laya 13/16. On the medication-conflict question alone: 4/4 vs 2/4. A useful diagnostic, not a clinical accuracy estimate.
Ricker
@0xRicker
Jev Engineering turns a static agent workflow into a graph that can rewrite itself while running. the video is basically the problem at scale: hundreds of routes → thousands of crossings → different agents → different tools → different confidence levels Jev Engineering doesn’t control every step. it controls the crossings. when two paths compete: → score both → kill the weak route → reroute the task so instead of one fixed chain, you get a live braid: state → decision → parallel paths → crossings → Jev → next state that’s the point of Jev Engineering: more parallel execution without letting the system lose the objective. the agents create the paths. Jev decides which paths survive. full breakdown below ↓
XAgents and browsers
Dain
@Dain0x
A beautiful pattern can still be noise. Here’s a JEV × Opus 5.5 workflow I’d test: → A data pipeline identifies candidate signals. → JEV selects which candidates to investigate using defined criteria. → Opus helps explore those candidates and propose explanations. → Statistical checks and human review test whether those explanations hold up. The critical question: what evidence would change our mind? That question belongs inside the workflow, before a promising pattern becomes a confident conclusion. This animation illustrates the concept. The particles and activity are simulated.
XBenchmarks and evals
frevana
@frevana_ai
Jev made competitor ad research 30x faster and ~90% cheaper. Typed “game” as the category. 20 seconds later: 294 top TikTok ads across 68 brands analyzed. Jev: 1/ pulled the top ads from the past 180 days 2/ scored every ad on stop-scroll, hook, format + trust 3/ clustered them into creative patterns 4/ classified winners, watches + dogs 5/ built a leaderboard of what’s winning 6/ recommended what to make next 20 sec · ~$0.02 For gaming: Winning hooks: offer/promo, challenge accepted Winning format: gameplay + UGC This is what we’re building at @frevana_ai with Jev: decode your category’s ads → know what to make next, in seconds. Try it free. Link in the first comment 👇
XAds and marketing

nicolay
@nicolaycz
Jev is the semantic if in your agent loop. LLM writes. Jev (TypeSafe System One) answers typed questions on the same state (Noul, Choice, Score) with calibrated probs. Your code branches. Pull the cheap decisions out of the LLM. Early access: console.typesafe.ai
XAgents and browsers
leiro
@leiroops
the lead engineer at a research company used Jev + GPT-6 Astra to build a system that analyzes 12,000,000 source updates every 6 hours the system processes massive streams of information, a process like this used to cost $1,000,000+, now Jev engineering does it for $200 the first version relied on GPT-6 Astra to interpret every result it worked, but processing millions of routine decisions through a large model created unnecessary latency and huge inference costs then the engineer rebuilt the decision layer around Jev instead of explaining every source, Jev reads the saved task state and decides what should happen next: Research, Write or Review Astra handles the open-ended work: comparing evidence, resolving contradictions and preparing the client brief code controls the queue, permissions, budget, retries and files at the core of this system is the same architecture I break down in the article: Jev routes the work, GPT-6 Astra handles the open-ended tasks, and code keeps the process running continuously the full step-by-step build of a client research process that runs 24/7 is below ↓ would you let AI decide which findings deserve further research?
gabidev
@GabiDev98
Here is another use case for Jev in DeFi. Vault risk classification. I pulled the signals that matter for Morpho vault risk (allocations, LLTV, utilization, idle, oracle, curator, APY sanity, liquidity and more) across 50 @Morpho vaults on Base, then asked Jev to score each one LOW / MEDIUM / HIGH / EXTREME. 50 judgments in 22.6s, ~452ms avg, ~60% median confidence, around $0.006 in costs Result mix: 30 MEDIUM - 12 HIGH - 8 EXTREME - 0 LOW. Jev also ranked curators and vaults by trust, who looks safest, who looks riskiest. Claude, Codex and other models can do this as well but it takes much more time and money.
XTrading and markets
keno
@kenonews
JEV can’t see images. So how do you give it eyes? What’s possible? What can run locally? How far can you get for free? The answers (and the catches) are in the full breakdown below:

XDocuments and OCR
Hassan
@nutlope
Just trained Tev1 0.8B, a tiny Jev-like classifier. Here it is running completely locally on my mac with @ollama & classifying some tasks. It's extremely fast: only ~50ms E2E latency. Video is not sped up! Releasing weights & benchmarks very soon so you can try it yourself :)
Viv
@Vtrivedy10
Jev for RAG in almost all cases you trust the semantic matching capability of Jev more than dot product similarity very useful as the direct similarity metric in small data cases and a great reranker with big data
XSearch
Jarek
@jarekceborski
I rebuilt CAPTCHA with Jev The browser measures how you fill in the form. Server turns that into a few plain sentences, and Jev decides: person or bot. Blog post with source code in the comments:
XSecurity and abuse
OpenRouter
@OpenRouter
Introducing typesafe/jev-router: a cache-aware model router powered by Jev and @typesafeai The Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. Here's how it works 👇🏻
XRouting and model choice
NO1ennn
@N01ennn
this is pure f*cking treasure A Stanford AI research group finally drew the perfect RAG system: retrieval, Jev and agents in one loop, and it fixes the 3 things that break every RAG app: > the LLM reads 20 passages when only 3 matter > it answers questions your docs can't answer > it trusts whatever text it retrieves here's how it runs: > a lead agent sends the query > hybrid search pulls the top 20 candidates, dense + keyword > ONE Jev request scores all 20 + 2 gates: answerable? injection? > only passages above 0.6 reach the writer agent > a second Jev call checks every claim against its source > grounded answer, with citations and when the docs don't have it: > answerable fails, the writer never runs > a researcher agent rewrites the query and retries once > still nothing? "not in the docs". zero tokens spent on a guess retrieval casts the net. Jev decides what's real. agents do the work save this before you build your next RAG
Gipp 🦅
@gippp69
Jev builders: the model bill is decided before the model is called /routing 20,000 events/day become: → 14,000 die as noise → 4,700 become logs → 800 earn a draft → 500 go to human review only 4% reaches writing. Jev costs $0.84; the full loop lands at $29.04. optimize the 4% before the writer.
XRouting and model choice
laxman
@llmluthor
Releasing Your Own Jev Post-train a 4B/8B/27B judge on your agent's traces. It beats Jev. 79.7% agreement with human labels vs Jev's 66.3% Beats DeepSeek-V4.1-Flash (763B) by 14 points 0.13s per step on a single GPU Completely open source: recipe, data, training, evals
XOpen source
What people posted on X while they built with Jev. Every card plays the original video or shows its images here, so you see the demo before you open the thread.
385 builds and 58 guides