benchmark author demo
Jev on two public safety corpora
A thread benchmarking Jev on public prompt-injection data, thresholding calibrated probabilities in code, with a dashboard and calibration analysis in follow-ups.
Notes
96.5% on all 662 messages of deepset/prompt-injections with no tuning, 325 ms p50.
author demo — Result shown by its author. About this label
Destinations
X post only — no separate service, repository or docs URL.