devwithjev
โ€” reading nowโ€” views
Submit a build

Jev Grades Code-Taste Tasks

Yuhan Luo ran a toy pipeline comparing Jev and Codex Luna as graders on FrontierCode-style code-taste tasks.

View on X cost<5% costtime<3% latency of luna
Yuhan Luo@_yuhanluo๐•
ran a toy pipeline comparing jev vs codex luna as graders on FrontierCode style "code taste" tasks: decomposed text-based rubrics into structured answers (e.g. are all changes necessary? do all changes consistently use the abstractions required by the task?) -> run both models in parallel to compare accuracy, latency, cost jev got similar accuracy at <5% cost and <3% latency of luna (tiny sample.
Sep 21, 2026X postsView on X
The pipeline converted text-based rubrics into structured questions about whether changes were necessary and consistently used the required abstractions, then ran both models in parallel. In a tiny sample, Jev had similar accuracy to Luna.

Also filed under Research & data