Search papers, labs, and topics across Lattice.
7
0
8
5
Overall, the results suggest that, under parameter-constrained fine-tuning, improvements in graph attention layers depend more on capacity allocation and branch interaction than on simply adding more learnable parameters.
Xiaomi-CocktailASR-1 achieves state-of-the-art performance, effectively addressing the cocktail party problem through a unified architecture that balances multispeaker and single-speaker recognition accuracy, along with rejection capability.
Achieving a staggering reduction in Word Error Rate from 12.15% to 2.79%, MiDashengLM-Gen sets a new standard for text-to-audio generation.
Mi-Memory achieves over 93% accuracy in preserving user context across diverse AI interactions, redefining how memory can govern personal AI experiences.
Agents trained with SEE-generated trajectories not only navigate complex multi-step procedures but also achieve unprecedented task success rates in real-world applications.
Harness TTS achieves up to 35.6-point improvements in instruction-following win rates, revolutionizing expressive speech synthesis for voice assistants.
Finally, a single model generates realistic and coherent audio scenes from text, rivaling specialized models and even approaching real-world recordings.