When Repository Labels Are Not Image-Level Truth: A Supervision Auditing Framework for Chest Radiograph AI

August 10, 2026 ยท Grace Period ยท ๐Ÿ› MICCAI 2026

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Yesika Alexandra Agudelo-Londoรฑo, Jhon Wilmer Pino-Romรกn, Brahian Carrera Rodrรญguez, Josรฉ Miguel Castaรฑeda-Bedoya, Juan Pablo Gรณmez-Lรณpez, Aura C. Puche-Sarmiento, Niharika S. D'Souza, Juan Sebastian Osorio-Valencia, Jon E. Duque-Grajales, Jazmรญn Ximena Suรกrez-Revelo, Jorge Mario Vรฉlez-Arango, Gabriel Castrillรณn arXiv ID 2608.10084 Category eess.IV: Image & Video Processing Cross-listed cs.CV Citations 0 Venue MICCAI 2026
Abstract
Public chest X-ray repositories are widely used to train medical AI systems, yet their labels are typically extracted from radiology reports rather than verified directly on images. As a result, repository labels are often treated as image-level ground truth without validating whether they reflect what is actually visible in the radiograph. We introduce Repository Supervision Auditing (RSA), a framework that evaluates repository-derived labels against expert image-level annotations before model development. Using cardiomegaly in MIMIC-CXR as a case study, RSA compares repository labels with radiologist-reviewed image annotations, characterizes disagreement sources, and builds a curated cohort for deployment-oriented evaluation. Repository-derived cardiomegaly labels showed near-zero agreement with expert image-level assessment, identifying only 1% of expert-confirmed cases. Most discrepancies resulted from non-mention rather than explicit report negation, with expert-confirmed cardiomegaly identified in nearly half of studies assigned a repository-derived No Finding label. Using the resulting expert-curated cohort, a DenseNet121 model achieved a test ROC-AUC of 0.853. These findings show that repository labels may not reliably represent image-level truth and highlight supervision auditing as a critical step for developing trustworthy medical imaging AI.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Image & Video Processing