benchmark author demo
Jev as a reranker: graded-relevance eval
Evaluates Jev reranking against bge-m3 embedding search over an Agent Skills Hub catalog, with judge-bias and robustness analysis.
Notes
33,047 skills, 9,831 labelled pairs, 164 queries: fusing bge-m3 with Jev lifts NDCG@10 from 0.774 to 0.864; standalone Jev rerank scores 0.028 below the embedding baseline.
author demo — Result shown by its author. About this label