Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Despite strong downstream performance, SLMs struggle with instruction-following due to weak alignment between speech and text representations, revealing a critical gap in their training methodology.
Retaining seemingly redundant tokens can actually enhance model performance, defying traditional assumptions about token importance in visual processing.