Search papers, labs, and topics across Lattice.
This paper introduces the IEEE SLT 2026 SmartGlasses Challenge, aimed at benchmarking egocentric multi-talker speech recognition and understanding using audio-language models. The challenge evaluates two tracks鈥擠yadic Dialogue Understanding and Multi-party Meeting Understanding鈥攖hrough a comprehensive dataset of 106 hours of real-world, four-channel egocentric speech recordings. Key findings reveal that significant speaker overlap severely impacts Time-Stamped Speaker-Attributed Automatic Speech Recognition (TSA-ASR) performance, while current models struggle with paralinguistic acoustic understanding in complex spoken language understanding (SLU) scenarios.
Heavy speaker overlap drastically hinders speech recognition accuracy in smart glasses, revealing critical limitations in current audio-language models.
Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentric platform for continuous audio sensing and assistance. However, speech recognition and understanding in this setting remain challenging because of dynamic acoustic conditions, speaker overlap, and the spatial ambiguity introduced by wearer-centered recording geometry. To support systematic evaluation in this setting, we introduce the IEEE SLT 2026 SmartGlasses Challenge for egocentric multi-speaker speech processing. The challenge consists of two tracks, Dyadic Dialogue Understanding and Multi-party Meeting Understanding, and jointly evaluates Time-Stamped Speaker-Attributed Automatic Speech Recognition (TSA-ASR) and Spoken Language Understanding (SLU). It is built on a 106-hour four-channel egocentric speech dataset containing 714 sessions collected in real-world scenarios. This paper describes challenge tasks, dataset construction, submissions, and summarizes the main findings from the shared evaluation. The results show that heavy speaker overlap remains a major factor affecting TSA-ASR performance, while paralinguistic acoustic understanding continues to be difficult for current audio-language models in complex SLU settings. Further details can be found on the official challenge website.