article source-backed
Evaluating Jev across 37 public datasets
An independent evaluation of pinned jev-1.13.0 across classification, routing, inference, moderation, legal clauses and rubric scoring. The paper uses frozen task templates, reports calibration and selective prediction, and releases its harness and raw model responses.
Notes
The authors report 346,009 answered requests across 37 datasets for under $10. Their analysis separates Choice calibration from binary threshold behavior and includes comparisons on identical requests with two open-weight reference models.
source-backed — Public repository, docs or live artifact. About this label