Search papers, labs, and topics across Lattice.
This study introduces BrailleBench, a comprehensive benchmark designed to evaluate the comprehension of Braille by large language models (LLMs) across various criteria, including mathematics and multi-hop question answering. By aligning 5,570 instances from five datasets and employing a deterministic, expert-reviewed methodology, the research highlights significant disparities in LLM performance between print-English and Braille, particularly noting that Grade 2 Braille comprehension is notably fragile. The findings underscore the need for improved AI systems that can effectively support Braille users, revealing critical gaps in current models' capabilities.
LLMs exhibit a troubling performance gap in Braille comprehension, with Grade 2 Braille proving especially challenging, highlighting urgent needs for inclusive AI design.
Although Large language models (LLMs) mediate access to knowledge and computational assistance, their capabilities should benefit vulnerable groups in the same way. However, it is unclear whether existing AI systems are inclusive enough for blind and deafblind users to access the same functionality through Braille, whose indicators, contractions, and digital representations introduce distinct requirements for model comprehension. To this end, we introduce BrailleBench, a benchmark for evaluating LLMs in Braille comprehension from different Criteria. BrailleBench aligns 5,570 instances from five datasets, including mathematics, commonsense, and multi-hop question answering across English and Braille Grades 1 and 2. Different configurations are designed to understand whether the systems can comprehend Braille-authored content, express answers in Braille, and complete end-to-end Braille interaction. To ensure the quality and prevent evaluation bias, the benchmark is built through a deterministic, expert-reviewed pipeline via a self-created Braille Toolkit without using any data instances generated by LLMs. We evaluate six representative LLMs from various aspects. The results reveal a persistent gap between print-English capability and Braille accessibility. Braille understanding and expression are asymmetric, where Grade 2 is especially fragile on the input side compared to Grade 1, and fully Braille requests further reduce performance. The experimental observations provide valuable guidance for the development of future Braille AI systems. All related resources in BrailleBench are publicly available for future research.