Item Response Theory for AI Safety
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. Authors: Joshua Fonseca Rivera, Neil Shah, David Demitri Africa.
Why it matters
Read this for the paper's specific claim in Artificial Intelligence / Machine Learning: Language models differ in how safely they behave and these differences are measured by safety benchmarks.