jevland

benchmark author demo

Jev on two public safety corpora

A thread benchmarking Jev on public prompt-injection data, thresholding calibrated probabilities in code, with a dashboard and calibration analysis in follow-ups.

96.5% on all 662 messages of deepset/prompt-injections with no tuning, 325 ms p50.

author demo — Result shown by its author. About this label

← Back to the directory