Every’s editorial vibe check
Eleven experiments, 1,709 judgments, under a cent in total.
Mike Taylor pointed Jev at everything he had written for Every: 37 documents, 21 questions each, answered in under 0.7 seconds. Against Fable 5.1 the median was 0.35 seconds per passage versus 8.83. It caught six of seven defects he had planted on purpose; the seventh only the larger model found.
Why it matters
As an early-warning pass it wins on the realistic comparison, which is not a slower model — it is nobody reviewing the draft at all.