Benchmarked Laya locally on an ancient laptop:
💻 i5-5200U (2 cores) | 12GB RAM | 5400 RPM HDD
Kept it resident in RAM and got sub-second (~630ms) typed decisions on pure CPU! 🔥
Great open-weight release by @Nandakishorm1. No GPU or cloud LLM needed for fast routing.
#laya #jev
Farhan60291312 benchmarked Laya on an old laptop, an i5-5200U with 12GB of RAM and an HDD. It reached about 630 ms decisions on CPU alone, with no GPU.
you can make any open source model behave like jev with just a bit of inference engineering.
it's shockingly easy. to prove it, we built a new endpoint we're calling deepseek-v4.1-flash-jev. see the demo below. here's how it's done:
sglang (an inference engine) offers a scoring endpoint in addition to the normal generation one. in scoring mode, given an input & set of possible answers, it forces the model to produce probabilities for each one. example:
> input: what is most common letter in abcccde?
> possible answers: a, b, c
> output: (c, 0.9), (b, 0.0.5), (a, 0.05)
getting the above behavior instead of streamed output is as simple as using sglang's /v1/score endpoint instead of /generate. there's just one other trick required.
for deepseek, you have to add a closing think tag before the response. this forces a direct answer instead of a reasoning trace. if you want reasoning, you can do that too, but imo that makes things too slow to be worth it.
dsv4.1 flash is not as good as jev, but if we had enough spare compute to experiment with this same approach for a larger model then i think the decision quality would be at least as good, if not better.
also, somewhat unrelated, i think decision-making models kill all prospecting & sourcing work. i would have absolutely killed to have jev or similar when i was recruiting @mintlify. absolutely incredible.