Found something worth sharing?
Drop the link and Jev reads it on the spot. What belongs lands in the community feed with your name on it.
More from the community
GitHubBenchmarks and evals
Code regression detection benchmark comparing probes and CI tests
Reproducible benchmark measuring whether decision probes detect code regressions better than direct model judgment or execution tests
Jev · 456ms
@BeldureilGitHubBenchmarks and evals
Pixel-art chess benchmark matches Laya against Jev
Two AI decision models play pixel-art chess with identical typed questions, with live viewer and replays in 10 languages
Jev · 249ms
@JHONSU777GitHubRouting and model choice
Cost-aware three-tier LLM agent orchestration with verification
Coordinates AI agents to complete tasks, verifies results against the original request, and exposes every step with per-call cost tracking
Jev · 444ms
@mgtf