Search papers, labs, and topics across Lattice.
Thanks:
6
0
9
2
Semantic speech tokens can be refined to enhance intelligibility and consistency across speakers, with significant implications for voice conversion and TTS applications.
DemoPSD effectively reduces privileged information leakage while enhancing exploration, leading to superior generalization in large language models.
Frequency adaptation can dramatically enhance source-free time-series domain adaptation, aligning target signals with source distributions more effectively than previous methods.
VL-DINO outperforms existing models in zero-shot object detection by effectively integrating vision-language knowledge, achieving a remarkable 38.1 AP on the LVIS benchmark.
ReasonAlloc reallocates KV cache resources in real-time, achieving superior reasoning efficiency with minimal overhead.
No single speech model excels at all editing tasks, revealing critical gaps in current Speech LLM capabilities.