Search papers, labs, and topics across Lattice.
The paper introduces Heart2Mind, a Contestable AI (CAI) system for psychiatric disorder prediction using wearable ECG data, designed to allow clinicians to inspect and revise algorithmic outputs. The system employs a Multi-Scale Temporal-Frequency Transformer (MSTFT) to analyze R-R intervals from ECG sensors, combining time and frequency domain features. Results on the HRV-ACC dataset show MSTFT achieves 91.7% accuracy, and human-centered evaluation demonstrates that experts and the CAI system can effectively collaborate to confirm correct decisions and correct errors through dialogue.
Clinicians can now contest AI-driven psychiatric diagnoses using wearable ECG data, thanks to a novel system that flags inconsistent predictions and facilitates collaborative refinement.
Psychiatric disorders affect millions, yet diagnosis depends on subjective assessments and uneven access to care. To address these challenges, there is a growing need for Contestable AI (CAI), a framework that extends beyond Explainable AI (XAI) by allowing clinicians to inspect, question, and revise algorithmic outputs, thereby reducing automation bias and strengthening accountability. We present Heart2Mind1, a human-centered CAI system for psychiatric disorder prediction that provides objective evidence while preserving clinical oversight. Heart2Mind collects R-R interval (RRI) time series from Polar H9/H10 wearable ECG sensors via a Cardiac Monitoring Interface and analyzes them using a Multi-Scale Temporal-Frequency Transformer (MSTFT) that combines time-domain and frequency-domain features. For contestability, the Contestable Diagnosis Interface integrates model explanations with dialogue. Self-Adversarial Explanations compare attention-based and gradient-based explanation maps to flag inconsistent predictions, and a collaboration chatbot helps users verify and challenge outputs. On the HRV-ACC dataset, MSTFT achieved 91.7% accuracy under leave-one-out cross-validation, outperforming benchmark methods. Human-centered evaluation with the Human-CAI Consensus Rate showed experts and CAI could confirm correct decisions and correct errors through readable, efficient dialogues ( \(FKGL\approx 15\) , median 8.3 minutes, 4 turns). These results support low-cost wearable CAI screening with objective biomarkers, safeguards, and an interactive path for clinicians to refine recommendations.