Article · by Vals AI
An independent evaluation of Jev, by Vals AI
Ties frontier models on claim verification, last of twelve on LegalBench.
Vals AI built two benchmarks that fit a model with no text output: 400 claim-verification items made from SEC filings, and a preregistered 12-subtask slice of LegalBench. It ran twelve systems on both. On claim verification Jev scores 0.975 at $0.02 per 1,000 cases and ties with GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. On LegalBench Jev finishes last of twelve on raw accuracy. Jev answers in about a tenth of a second for one judgment or thirty-two; at thirty-two its median latency is 194x lower than GPT-6 Astra's.
Open the source