Search papers, labs, and topics across Lattice.
3
0
3
1
LALMs can achieve high agreement with human evaluators while still relying on misleading shortcuts, risking the integrity of speech evaluations.
AI-generated music detectors reveal hidden vulnerabilities when tested against unseen generator sources, with token performance varying dramatically based on training data.
Forget MOS: a new preference-based metric, AnimeScore, finally cracks the code for automatically evaluating "anime-like" speech with 90.8% AUC.