Benchmarks Jev Against LLMs
The repository provides independent benchmark data comparing TypeSafe's Jev (System One model) with LLMs on accuracy, calibration, and cost.
Jevals@Jevals
Independent benchmark data for TypeSafe's Jev (System One model) vs LLMs: accuracy, calibration, cost. Boards + per-decision logs, CC-BY-4.0
It includes boards and per-decision logs. The data is licensed under CC-BY-4.0.
Also filed under Research & data
- Visualizes French Wikipedia Elites
The project is a data visualization of French elites on Wikipedia. Its crawl is ongoing, with new biographies arriving.
- GraphRAG with Swappable Laya and Jev Models
This project is an agentic GraphRAG pipeline with swappable local Laya or cloud Jev decision models.
- Jev Probability Calibration Experiments
The repository presents reproducible experiments on Jev probability calibration, uncertainty, and forecast preservation.
- Benchmarks Jev for Malicious Skill Detection
jev-skillbench benchmarks Jev as a malicious agent-skill detector on MalSkillBench with verify-and-escalate evaluation.