jevland

benchmark source-backed

Jev × LexGLUE: typed decisions across seven legal tasks

A reproducible zero-shot evaluation freezes the LexGLUE test splits and maps legal classification tasks to Choice or per-label Noul questions. Its runner saves completed requests, checks model-version changes and writes micro/macro-F1, calibration, usage and cost reports.

For 23,607 test examples over seven tasks, the author reports Jev 1.13's arithmetic mean micro-F1 of 69.9 and an observed API cost of $4.02; validation-tuned thresholds are reported separately.

source-backed — Public repository, docs or live artifact. About this label

Link reachable ·

HTTP 200 · HTTP response only. Checked destination ↗ · Last reachable

What these checks cover →

Report a problem →

← Back to the directory