Japan 2026 Common Test AI Benchmark
A static dashboard benchmarks Jev against luna-none and luna-low on Japan's 2026 Common Test.
shibadogcap@shibadogcap
AI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard).
Also filed under Research & data
- Visualizes French Wikipedia Elites
The project is a data visualization of French elites on Wikipedia. Its crawl is ongoing, with new biographies arriving.
- GraphRAG with Swappable Laya and Jev Models
This project is an agentic GraphRAG pipeline with swappable local Laya or cloud Jev decision models.
- Jev Probability Calibration Experiments
The repository presents reproducible experiments on Jev probability calibration, uncertainty, and forecast preservation.
- Benchmarks Jev for Malicious Skill Detection
jev-skillbench benchmarks Jev as a malicious agent-skill detector on MalSkillBench with verify-and-escalate evaluation.