Jev Scores Agent Code Diffs
Jev evaluates diffs and returns a score without generating text. The author describes using it for 31 checks on a 340-line auth refactor.
vvtentt@Vvtentt101𝕏
the slow part of your agent is not the hard problem.
the problems start after it does the same work several times.
JEV DOESN'T WRITE. IT DECIDES.
one auth refactor. 340 lines. 31 checks: keep this hunk, reject that file, stop or go again. claude code spent a full generation on each one.
jev takes the diff and returns a score. 70-500ms. $0.042 per million in. no output text. output is free because there isn't any.
20-200x faster. 40-400x cheaper. the ceiling is only for loops that decide a lot and write a little.
The checks decide whether to keep a hunk, reject a file, or stop or continue. The author contrasts Jev's scoring with Claude Code spending a full generation on each check.
Also filed under Triage & routing
- Python Library for Jev Judgments
gut is an open-source Python library that turns Jev’s typed questions into a readable line of code.
- Real-Time Feedback Alert Classification
The author reports using Jev in a feedback platform, where real-time alert classification is already live.
- Yueli DEX Decision Exchange
Yueli DEX routes bounded business decisions to TypeSafe Jev or compatible providers and maps validated results to non-authorizing action intents.
- Medical Exam Option Scoring and Gating
Open Medical Jev is a frozen open-model medical decision stack that scores each exam option, then uses confidence and conformal gates to auto-release or escalate answers.