Skip to content
Made with Jev

Article · by Vals AI

An independent evaluation of Jev, by Vals AI

Ties frontier models on claim verification, last of twelve on LegalBench.

Vals AI built two benchmarks that fit a model with no text output: 400 claim-verification items made from SEC filings, and a preregistered 12-subtask slice of LegalBench. It ran twelve systems on both. On claim verification Jev scores 0.975 at $0.02 per 1,000 cases and ties with GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. On LegalBench Jev finishes last of twelve on raw accuracy. Jev answers in about a tenth of a second for one judgment or thirty-two; at thirty-two its median latency is 194x lower than GPT-6 Astra's.

Open the source

More like this