Search papers, labs, and topics across Lattice.
This paper introduces EndoCA, a benchmark designed to evaluate the consistency between complex answers and their atomic components in endoscopic visual question answering (VQA). The authors assess 11 vision-language models (VLMs) and find that while some achieve high accuracy on complex answers, their performance on atomic answers and overall consistency is significantly lower. To address this inconsistency, they propose Atomic-Support Reconciliation (ASR), a training-free method that enhances answer accuracy by utilizing atomic answers for contextual support and selective answering strategies.
Despite high complex-answer accuracy, many VLMs struggle with atomic consistency, revealing a critical gap in endoscopic VQA performance.
Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries. Such complex answers may be scored as correct even when the same model fails on associated atomic questions. We introduce EndoCA, a paired complex-atomic answer consistency benchmark for evaluating whether complex answers remain consistent with same-image atomic answers. EndoCA contains two suites: EndoCA-Core evaluates compact question-complexity patterns commonly seen in practical endoscopic VQA, and EndoCA-Diagnostic supports controlled analysis across increasing question complexity. We evaluate 11 VLMs spanning open, medical, endoscopy-adapted, and closed-source models on EndoCA. Some VLMs achieve high complex-answer accuracy, yet their atomic-answer accuracy and complex-atomic answer consistency remain substantially lower. To reduce this complex-atomic inconsistency, we introduce Atomic-Support Reconciliation (ASR), a training-free mechanism that uses model-generated atomic answers as contextual premises for answer revision and consistency-guided selective answering. On four selected publicly available models, ASR-Revise improves paired complex-atomic correctness with modest changes in complex-answer accuracy, while ASR-Selective improves accuracy on answered cases by allowing the model to abstain from less reliable cases. Together, EndoCA and ASR provide a consistency-aware benchmark and a training-free mechanism for answer reconciliation and selective answering in endoscopic VQA.