Skip to content
Made with Jev

X · by Stas Sorokin

A thousand papers sorted, then graded

Jev does the 24-topic sort, Opus 5 grades the labels.

Stanislav Sorokin

@stas_sorokin_

1,000 AI papers sorted into 24 topics for $0.0585. Then Opus 5 graded the labels. @nutlope's Jev paper map went viral, but the pipeline never shipped and the eval was "still running". So I rebuilt both and opened them. The first judge run came back empty: Opus spent its whole budget thinking and answered nothing. Reasoning off, second run: it agreed with Jev on 85 of 100 papers, at 153x the cost and 1.9s against 57ms per paper. The 15 misses are not random. One number Jev already returns tells you which labels to recheck. Cheap models sort. Expensive models audit only what the cheap one flags. Repost if you classify anything at scale, because the eval rows are public and anyone can rerun them with their own judge in one command. Code in the reply.

Sep 21, 2026 · 5 likesOpen on X

Stas Sorokin rebuilt Nutlope's Jev paper map pipeline: Jev sorts 1,000 AI papers into 24 topics, and Opus 5 grades the labels afterwards so the accuracy is checked rather than assumed.

Open the source

More like this