jevland

benchmark source-backed

llm-rankers: checking Jev's calibration with TREC judgments

The ielab document-ranking toolkit gained a jev module, with code that uses TREC human relevance judgments to test whether Jev's probabilities are actually calibrated.

Code for the calibration check; results in the repository.

source-backed — Public repository, docs or live artifact. About this label

← Back to the directory