jevland

article source-backed

Evaluating Jev across 37 public datasets

An independent evaluation of pinned jev-1.13.0 across classification, routing, inference, moderation, legal clauses and rubric scoring. The paper uses frozen task templates, reports calibration and selective prediction, and releases its harness and raw model responses.

The authors report 346,009 answered requests across 37 datasets for under $10. Their analysis separates Choice calibration from binary threshold behavior and includes comparisons on identical requests with two open-weight reference models.

source-backed — Public repository, docs or live artifact. About this label

← Back to the directory