Search papers, labs, and topics across Lattice.
The paper introduces XIMED, a dual-loop evaluation framework for XAI methods in medical imaging, specifically chest X-ray classification. It evaluates LIME and SHAP explanations using both predictive model-centered metrics (sensitivity to model changes, feature identification) and human-centered metrics (trust, diagnostic agreement) with 97 medical experts. The study found that while SHAP significantly impacted diagnostic changes and both methods identified critical features, both LIME and SHAP negatively impacted contra-indicative agreement, with SHAP proving more effective in facilitating correct diagnostic changes when initial diagnoses were correct.
SHAP explanations can significantly influence diagnostic changes in medical experts, but both SHAP and LIME can negatively impact agreement on contra-indicative reasoning.
In this study, a structured and methodological evaluation approach for eXplainable Artificial Intelligence (XAI) methods in medical image classification is proposed and implemented using LIME and SHAP explanations for chest X-ray interpretations. The evaluation framework integrates two critical perspectives: predictive model-centered and human-centered evaluations. Predictive model-centered evaluations examine the explanations’ ability to reflect changes in input and output data and the internal model structure. Human-centered evaluations, conducted with 97 medical experts, assess trust, confidence, and agreements with AI’s indicative and contra-indicative reasoning as well as their changes before and after provision of explainability. Key findings of our study include explanation of sensitivity of LIME and SHAP to model changes, their effectiveness in identifying critical features, and SHAP’s significant impact on diagnosis changes. Our results show that both LIME and SHAP negatively affected contra-indicative agreement. Case-based analysis revealed AI explanations reinforce trust and agreement when participant’s initial diagnoses are correct. In these cases, SHAP effectively facilitated correct diagnostic changes. This study establishes a benchmark for future research in XAI for medical image analysis, providing a robust foundation for evaluating and comparing different XAI methods.