devwithjev
— reading now— views
Submit a build

Jev Guides Pi Agent Tool Decisions

Jev runs in Pi Agent, where it uses the current state to decide what to do next while Pi executes tools. In tests involving Git changes and failures, it dynamically changed the call path and completed all three scenarios.

View on X time平均延迟约 683ms
小墨同学@xiaomovps𝕏
Jev 我已经在 Pi Agent 里真正跑起来了🔥 我这次测试的重点不是测试分类,而是看它能不能参与真实的多工具决策:Pi 负责执行工具,Jev 根据当前 State 决定下一步该做什么 我正常跑了Git 修改、测试失败 3 个场景,Jev 都能根据状态动态改变调用路径,失败时还会主动增加 inspect_failure,而不是机械走固定流程。 13 次调用,3/3 跑通,平均延迟约 683ms。 中间第三次调用出现了失败,然后又做了进一步的调整,最后还是完成了目标 从这个结果角度来看,我希望把 skill 和 MCP 判断的能力交给其他的低端模型去完成,主模型只要去完成业务相关的内容,像工具使用或者结果类型判断,交给一些专门模型处理
Sep 18, 2026X postsView on X
The author reports 13 calls and says Jev added an `inspect_failure` step after a failure rather than following a fixed workflow. The author hopes specialized models can handle skill and MCP decisions while the main model focuses on business tasks.

Also filed under Agents & browsers

  • Evaluates Tool-Using Agents with Jev in CI

    The project evaluates a real tool-using AI agent using code checks, a calibrated Jev judge, and an LLM judge. It is wired into CI.

  • Watchdog for AI Coding Agents

    Jev Watch is a watchdog for AI coding agents that stops Codex or OpenCode when they loop, stall, or drift off task, then resumes the same session with a correction.

  • Custom Agent Harnesses with Jev

    Elvis uses Jev to build custom harnesses and reports success with guardrails, routing, and verifiers.

  • Fraud Investigation Agent

    The author is building a fraud investigation agent using Jev and TigerGraph. Its flow goes from an alert through graph evidence, a Bayesian update, a pattern, Jev’s extra lookup, policy, a verdict, and human approval.