devwithjev
β€” reading nowβ€” views
Submit a build

Programmatic Evaluation with Jev and Claude

Pranav Ramesh tested Jev paired with Claude for programmatic evaluation and described the results as impressive.

Pranav Ramesh@PranavRamesh123𝕏
Tested Jev (@typesafe_ai) paired with Claude for the first time, and the results for programmatic evaluation are impressive. πŸ§΅πŸ‘‡ Takeaway: Don't force general LLMs to handle deterministic scoring. Pairing Claude’s/Codex reasoning What use cases would you test this on?
Sep 23, 2026X postsView on X
He says general LLMs should not be forced to handle deterministic scoring. He asks what use cases others would test this on.

Also filed under Tools & apps

  • Jev: a "System One" model that only makes decisions, $0.042 per million tokens with output free. 40-second explainer of why it matters for automation

    Jev: a "System One" model that only makes decisions, $0.042 per million tokens with output free. 40-second explainer of why it matters for automation Jev launched on Sep 15 from TypeSafe AI (founder Diogo Almeida, ex-OpenAI, credited on ChatGPT/InstructGPT). It does not generate text; it returns typed decisions (Choice / Score / true-false) with confidence. Price is $0.042 per million input tokens, output free. In its first week it became the fastest-adopted model on Vercel's AI Gateway, TypeSafe paused signups on Sep 22, Browser Use released jev-ultrafast (a web agent where Jev picks each action; Google Flights search in 7.1 s), and jaredpalmer/kev is an Apache-2.0 clone built on Qwen3.5 you can run yourself. My take in the video: the expensive part of automation was never the writing, i

  • Untitled

    There's a new AI model that refuses to talk. While everyone debates what a "decision model" is even for, builders already shipped the answer. 7 things people built with Jev in 9 days. Number 5 broke my brain πŸ‘‡

  • Memory System Built with Jev

    Opium says they made a memory system using Jev and that results are very promising so far.

  • Deterministic LLM Judge Aid

    Fadhli used Jev as a deterministic tool to aid an LLM as a judge (gpt-6) in a synthetic, simple-data use case.