Search papers, labs, and topics across Lattice.
7
0
11
31
Video LLMs can significantly improve their QA performance by integrating spatio-temporal evidence, bridging the gap between accuracy and visual perception.
Achieving 4.7x to 8.2x higher throughput for trillion-parameter MoE models could redefine the limits of large-scale model training.
Whisper's speech-centric training leaves audio-LLMs tone-deaf to music and environmental sounds, but a simple fine-tune can fix that.
Forget slow, end-to-end models: building real-time voice agents hinges on a cascaded streaming pipeline, as demonstrated by a new tutorial achieving sub-second latency.
Forget text prompts: vector prompt interfaces are the key to unlocking scalable and stable LLM customization.
Real-time voice agents can bypass slow vector DB lookups with a dual-agent architecture that pre-fetches relevant documents into a sub-millisecond semantic cache.
A unified Vision-Language Model and Diffusion architecture unlocks surprisingly effective optical flow forecasting from noisy web data, enabling language-conditioned robot control and video generation.