Search papers, labs, and topics across Lattice.
3
0
4
3
Visual textualization and image-native modeling outperform traditional text-only methods in predicting item difficulty, revealing the critical role of visual evidence in assessment design.
Achieving over 82% output correctness, this new benchmark and model redefine the standards for audio-visual target speaker extraction by effectively integrating visual cues.
LLMs can achieve state-of-the-art audio-visual speech recognition by sparsely aligning modalities and refining with visual unit guidance, substantially boosting robustness in noisy environments.