devwithjev
— reading now— views
Submit a build

Browse builds

32 builds · page 1 of 1

MohibShaikh@MohibShaikh
Benchmark of Jev as a malicious agent-skill detector on MalSkillBench with verify-and-escalate evaluation.
1206GitHub·Research & dataOriginal source ↗
KantaHayashiAI@KantaHayashiAI
Reproducible experiments on Jev probability calibration, uncertainty, and forecast preservation.
1126GitHub·Research & dataOriginal source ↗
bodepudimuneendra-netizen@bodepudimuneendra-netizen
Agentic GraphRAG pipeline with swappable local Laya or cloud Jev decision models.
0312GitHub·Research & dataOriginal source ↗
zekezeke@zeke
Research notes and an interactive Cloudflare Worker demo for Jev, TypeSafe AI's structured decision model
0054GitHub·Research & dataOriginal source ↗
erendikmenn@erendikmenn
Reproducible benchmark for measuring Jev reranking quality, latency, and cost in RAG.
1099GitHub·Research & dataOriginal source ↗
patryckalves@patryckalves
Reproducible evaluation of Jev Choice decisions on 182 valid ENEM 2025 questions, with raw results, accuracy, calibration, and latency analysis.
1207GitHub·Research & dataOriginal source ↗
steven-shoemaker@steven-shoemaker
Python library that maps Jev Choice, Score, and Noul questions onto classification, ranking, extraction, verification, and DataFrame workflows.
0360GitHub·Research & dataOriginal source ↗
yodablocks@yodablocks
Does ORDER BY over a Jev probability put rows in a defensible order? Independent ranking, calibration and invariant measurements of TypeSafe AI's Jev: passes six pre-registered gates on 360 labeled rows, fails four of six on graded product relevance.
1209GitHub·Research & dataOriginal source ↗
willkelly@willkelly
Adversarial Jev evaluation suite with preregistered predictions, request logs, and reproducible experiment artifacts.
1208GitHub·Research & dataOriginal source ↗
onlyoneaman@onlyoneaman
TypeSafe's Jev vs gpt-5.4-mini and gpt-5.6-luna on four public classification sets: cases, per-item answers, scoring, charts
1205GitHub·Research & dataOriginal source ↗
Jevals@Jevals
Independent benchmark data for TypeSafe's Jev (System One model) vs LLMs: accuracy, calibration, cost. Boards + per-decision logs, CC-BY-4.0
1204GitHub·Research & dataOriginal source ↗
kiarina@kiarina
Research workspace containing reproducible TypeSafe Jev evaluation and safety-judgment experiments.
1200GitHub·Research & dataOriginal source ↗
scienthoon@scienthoon
Reproducible evaluation of Jev calibration on synthetic support tickets and public classification benchmarks.
1118GitHub·Research & dataOriginal source ↗
fstandhartinger@fstandhartinger
JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.
1085GitHub·Research & dataOriginal source ↗
ktaletsk@ktaletsk
Semantic AI for pandas and Polars: classify text, analyze sentiment, and score DataFrame rows with natural-language questions and full probabilities using TypeSafe Jev.
0343GitHub·Research & dataOriginal source ↗
choxos@choxos
Systematic-review extraction app where Jev answers structured review forms and points reviewers to quoted source lines.
0779GitHub·Research & dataOriginal source ↗
keltokhy@keltokhy
Selects source-linked evidence within a token budget using Jev relevance judgments plus local diversity-aware selection.
0807GitHub·Research & dataOriginal source ↗
vehas@vehas
Charts: TypeSafe Jev evaluated on Thai standardized exams vs 110 other models.
1199GitHub·Research & dataOriginal source ↗
smasato@smasato
Jev (TypeSafe) 性能評価プロジェクト — 日本郵便 KEN_ALL をマスタに、AI SDK 経由の Jev が住所のあいまい一致にどこまで使えるかを検証.
1195GitHub·Research & dataOriginal source ↗
Menny1337@Menny1337
TypeScript experiments, evaluations, and latency benchmarks for TypeSafe's Jev model.
1192GitHub·Research & dataOriginal source ↗
zsavage8@zsavage8
Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev.
1187GitHub·Research & dataOriginal source ↗
shibadogcap@shibadogcap
AI benchmark on Japan's 2026 Common Test: Jev vs luna-none vs luna-low (static dashboard).
1184GitHub·Research & dataOriginal source ↗
TokenTrim@TokenTrim
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
1178GitHub·Research & dataOriginal source ↗
anisselbd@anisselbd
Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark.
1174GitHub·Research & dataOriginal source ↗
anessbelbati@anessbelbati
Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap.
1116GitHub·Research & dataOriginal source ↗
RINNECODER@RINNECODER
Independent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification.
1113GitHub·Research & dataOriginal source ↗
AkashPriyadarshii@AkashPriyadarshii
High-throughput synthetic & pretraining dataset sifter powered by TypeSafe AI Jev (api.typesafe.ai). Stream, filter, and score Parquet & JSONL datasets at 1,500+ rows/sec using System One typed decisions (Choice, Score, Noul).
1110GitHub·Research & dataOriginal source ↗
zhuyansen@zhuyansen
Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.
1109GitHub·Research & dataOriginal source ↗
mahlernim@mahlernim
Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence.
1106GitHub·Research & dataOriginal source ↗
zhihz@zhihz
Local bilingual probability decisions from context, questions, and candidate answers. Independent research preview inspired by TypeSafe Jev.
1096GitHub·Research & dataOriginal source ↗
rorshopping@rorshopping
Unofficial study: Jev-style parallel typed decisions on stock 1.5B-8B models on an Apple Silicon laptop. Benchmarks, research notes, and a Hugging Face Space demo.
1093GitHub·Research & dataOriginal source ↗
r-ms@r-ms
Mini-Jev: what a Jev-style typed-decision interface looks like on a frozen Qwen3-4B — read the option letter's logits instead of generating JSON. Preregistered experiment, results, teaching bench.
1092GitHub·Research & dataOriginal source ↗