KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models
Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Authors: Sparsh Roy, Samuel Girmachew, Nishita Chavan.
Why it matters
Read this for the paper's specific claim in Artificial Intelligence / Machine Learning: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups.