devwithjev
— reading now— views
Submit a build

Jev End-to-End Testing

Jev runs end-to-end tests. Jason Lu reports a completed-run median of 47s and $0.0067 on the same eBay test flow.

View on X cost$0.0067time47s
Jason Lu@jasonlu_ai𝕏
JEV is changes the world of E2E testing! Same eBay test flow, completed-run medians: Jev: 47s / $0.0067 GPT-5.6 Luna: 62s / $0.0277 Claude Sonnet 5: 79s / $0.4062 Try jev-e2e. https://t.co/fyOdrNi3jF https://t.co/T2zaVnGvYO
Sep 18, 2026X postsView on X
On that flow, GPT-5.6 Luna had a 62s / $0.0277 median, and Claude Sonnet 5 had a 79s / $0.4062 median.

Also filed under Agents & browsers

  • Evaluates Tool-Using Agents with Jev in CI

    The project evaluates a real tool-using AI agent using code checks, a calibrated Jev judge, and an LLM judge. It is wired into CI.

  • Watchdog for AI Coding Agents

    Jev Watch is a watchdog for AI coding agents that stops Codex or OpenCode when they loop, stall, or drift off task, then resumes the same session with a correction.

  • Fraud Investigation Agent

    The author is building a fraud investigation agent using Jev and TigerGraph. Its flow goes from an alert through graph evidence, a Bayesian update, a pattern, Jev’s extra lookup, policy, a verdict, and human approval.

  • Coding-Agent Runtime Using Jev Decisions

    Distill is a coding-agent runtime that uses Jev decisions for model routing, context retention, and tool selection.