Performance-interpretability trade-offs and generalization in deep learning for pneumonia detection: A benchmarking study

dc.contributor.authorGómez-Vela, Francisco Antonio
dc.contributor.authorLópez Fernández, Aurelio
dc.contributor.authorDivina, Federico
dc.contributor.authorGarcía Torres, Miguel
dc.date.accessioned2026-09-18T10:45:14Z
dc.date.available2026-09-18T10:45:14Z
dc.date.issued2026-09-10
dc.description.abstractChest X-ray imaging remains the most widely used radiological modality for pneumonia screening. While recent advances in deep learning have demonstrated strong diagnostic performance, the deployment of such models in real-world settings requires not only high accuracy but also robustness and interpretability across different model designs and populations. In this work, we present a comprehensive benchmarking study of multiple deep learning architectures for pneumonia detection. A distinctive methodological feature of this study is the use of two demographically distinct datasets: a pediatric chest X-ray dataset for training, and an independent adult population dataset for external validation. All models were evaluated under a standardized cross-validated protocol. Beyond predictive metrics, we conduct an extensive eXplainable Artificial Intelligence (XAI) analysis assessing both qualitative and quantitative properties, including explanation stability and localization fidelity. Results show that model design choices significantly influence both predictive performance and explanation behavior. In particular, certain architectures consistently achieved superior performance while producing more focused and stable explanations. The performance-interpretability ranking is preserved under external validation on a demographically distinct adult cohort, providing evidence that performance-interpretability coupling is robust to population shift and supports the generalizability of the proposed framework beyond the pediatric training distribution. This work contributes a reproducible, end-to-end methodology that jointly optimizes performance and interpretability, offering practical guidance for selecting and deploying explainable deep learning models in clinical pneumonia screening. All code and data are publicly available at: https://github.com/SynergIA-Lab/pneumoniacnn
dc.description.sponsorshipUniversidad Pablo de Olavide de Sevilla, Departamento de deporte e informática
dc.format.mimetypeapplication/pdf
dc.identifier.citationAppl Intell 56, 414 (2026)
dc.identifier.doi10.1007/s10489-026-07398-5
dc.identifier.urihttps://hdl.handle.net/10433/27429
dc.language.isoen
dc.publisherSpringer
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 Internationalen
dc.rights.accessRightsopen access
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectDeep learning
dc.subjectMedical image analysis
dc.subjectExplainable artificial intelligence
dc.subjectChest X-ray
dc.subjectModel interpretability
dc.subjectModel generalization
dc.titlePerformance-interpretability trade-offs and generalization in deep learning for pneumonia detection: A benchmarking study
dc.typejournal article
dc.type.hasVersionVoR
dspace.entity.typePublication
person.affiliation.nameUniversidad Pablo de Olavide
person.affiliation.nameUniversidad Pablo de Olavide
person.affiliation.nameUniversidad Pablo de Olavide
person.affiliation.nameUniversidad Pablo de Olavide
person.identifier.orcid0000-0001-7376-5790
person.identifier.orcid0000-0001-5986-5437
person.identifier.orcid0000-0002-0964-9506
person.identifier.orcid0000-0002-6867-7080
relation.isAuthorOfPublicationd1d327f0-daff-46c1-af17-bd2b79390ed7
relation.isAuthorOfPublication5205a971-aeb9-4488-a278-e61cadd3b544
relation.isAuthorOfPublication82e2c456-c4b8-494e-b3d9-f6c84c8cf9a5
relation.isAuthorOfPublication4ce19614-9553-49b0-9b6e-09817f551658
relation.isAuthorOfPublication.latestForDiscoveryd1d327f0-daff-46c1-af17-bd2b79390ed7

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Pneumonia.pdf
Size:
3.01 MB
Format:
Adobe Portable Document Format