Live site · by florianstandhar
JevBench
A reproducible benchmark for typed decision models.

florianstandhar’s JevBench is a reproducible benchmark for typed decision models, published on Benchmark Heaven.
Open the sourceLive site · by florianstandhar
A reproducible benchmark for typed decision models.

florianstandhar’s JevBench is a reproducible benchmark for typed decision models, published on Benchmark Heaven.
Open the sourceWayne Sutton
@waynesutton
Ask Jev anything. Give it a try at askjev.ai It won't answer. It will judge. Let's see if we can get to 1 million questions. @typesafeai 🤝 @convex work great together. @hmartenjoyer @CompleteSkeptic @justKDeng @mikeysee
Live siteBenchmarks and evals
OpenRouter
@OpenRouter
1/ Jev, a decision model by @typesafeai, sparked a burst of projects and discussion. We tested it using Ori Eval against popular LLMs on OpenRouter at judging. Jev was >5x faster than the next fastest model, and even its slowest requests beat every other model's median.
Dan Shipper
@danshipper
we almost never test new foundation models but we've been testing this for ~a week @every and it's pretty wild. the kind of things that will be obviously indispensible in 6-12 months it doesn't produce words as output, it produces probabilities. so it can efficiently act as a judge in cases where you'd need a Fable-level model—but in our testing was 25x faster and 600x lower priced excellent vibe check by @hammer_mt on @every: https://every.to/also-true-for-humans/mini-vibe-check-typesafe-s-jev-judged-everything-i-ve-written-in-0-7-seconds?utm_cta_source=home_main_a_3
ArticleBenchmarks and evals
Pick