benchmark source-backed
jev-acento: a paired Spanish-language accuracy and calibration audit
A preregistered audit compares English and Spanish state and instruction combinations for Jev on paired human-labelled datasets. The repository publishes the runner, frozen prompts, response-derived results and a CLI for repeating the protocol on other labelled data.
Notes
Run 20260921-es-v1 records 19,200 calls over 3,200 paired XNLI, PAWS-X, MASSIVE and Belebele items using jev-1.13.0; Spanish state lowers reported accuracy by 3.0–6.4 percentage points across those datasets.
source-backed — Public repository, docs or live artifact. About this label