devwithjev
— reading now— views
Submit a build

Jev Search and Rerank Evaluation

The project evaluates whether a TypeSafe Jev reranker beats embedding search on the Agent Skills Hub catalog. It uses graded relevance evaluation over 9,831 pairs and 164 zh/en queries, with judge-circularity bias measured.

zhuyansen@zhuyansen
Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured.
Sep 19, 2026GitHubView on GitHub

Also filed under Research & data