Skip to content
Made with Jev

X · by Ranjan Kumar

Trust the ordering, not the confidence number

Where to place the decision model inside an agent harness.

Ranjan Kumar

@ranjankumar

𝐖𝐡𝐞𝐫𝐞 𝐉𝐞𝐯 𝐁𝐞𝐥𝐨𝐧𝐠𝐬 𝐢𝐧 𝐚𝐧 𝐀𝐠𝐞𝐧𝐭 𝐇𝐚𝐫𝐧𝐞𝐬𝐬 (𝐍𝐨𝐭 𝐚 𝐌𝐨𝐝𝐞𝐥 𝐒𝐰𝐚𝐩) Your agent harness needs a decision model. Where you place it depends on one brutal fact: Jev's ordering is trustworthy. Its confidence numbers are not. Most placement guides skip this distinction entirely. They tell you to pick a threshold and execute above it. That works if you actually have a probability. You might not. 𝐓𝐡𝐞 𝐜𝐨𝐫𝐞 𝐩𝐫𝐨𝐛𝐥𝐞𝐦: A model can rank beautifully and lie about magnitude simultaneously. When you gate execution on if confidence >= 0.85: approve_transfer, you are trusting a number that was never calibrated to your prevalence, your cost ratio, or your queue. Move the threshold up and you trade recall you never measured for precision you cannot state. You have no idea how far along an unmarked axis you moved. 𝐓𝐡𝐫𝐞𝐞 𝐝𝐞𝐜𝐢𝐬𝐢𝐨𝐧𝐬, 𝐭𝐡𝐫𝐞𝐞 𝐝𝐢𝐟𝐟𝐞𝐫𝐞𝐧𝐭 𝐧𝐞𝐞𝐝𝐬: Routing and ranking consume ordering only - which answer is best matters, the number beside it does not. Gated execution consumes magnitude - the threshold is your boundary and it must land on a calibrated scale. Relative logic like if top1 - top2 < 0.1: escalate consumes differences - and rescaling the score axis will flatten margins unevenly, breaking your code. The sepsis alert systems of 2020 learned this hard way. Michigan switched off alerts when COVID shifted patient prevalence beneath a fixed threshold. No weights changed. The denominator moved, the promise broke silently, and nurses drowned in false alarms. 𝐓𝐡𝐞 𝐟𝐢𝐱 𝐢𝐬 𝐬𝐞𝐪𝐮𝐞𝐧𝐜𝐞, 𝐧𝐨𝐭 𝐭𝐮𝐧𝐢𝐧𝐠: First question: is this number admissible as a probability at all? Run a calibration study on a few hundred labelled cases from your own queue. Samuel Sacco's measurements from 18 September and Adil Muhammad Pervez's 8,000 judgments both reach the same conclusion: fit your own map. The weights are shared across every account by design - there is no per-customer adaptation - so a deployer-side calibration map is your only mechanism. 𝐒𝐞𝐜𝐨𝐧𝐝 𝐪𝐮𝐞𝐬𝐭𝐢𝐨𝐧: given that the number is calibrated, where does your cost ratio put the line? That part of the familiar advice still works. Run them backwards and you are tuning a dial with no markings. Read the full analysis on calibration, harness placement, and where Jev actually fits: https://ranjankumar.in/jev-system-one-model-agent-harness-placement Follow for more practitioner insights on agentic AI systems and production AI engineering. #AgentiveAI #AIEngineering #SystemOne #Calibration #MLOps #DecisionModels #HarnessEngineering #Jev

Sep 21, 2026 · 0 likesOpen on X

A discussion of where Jev belongs in an agent harness. The author's rule: the ordering Jev produces can be trusted, the confidence numbers cannot, and the harness should be built around that.

Open the source

More like this