Search papers, labs, and topics across Lattice.
10
0
8
4
AV-Flamingo outperforms existing models on complex audio-visual tasks, revealing that size isn't everything when it comes to reasoning capabilities.
AutoVSR achieves up to 59.45% higher accuracy in generating symbolic expressions from circuit schematics, revolutionizing the way we interpret circuit behavior.
Physically aligned video models can boost robotic manipulation success rates by over 50% compared to traditional methods.
Teacher-forcing consistency models can accelerate autoregressive video generation by ten times, revolutionizing the training landscape for streaming applications.
SC3-Eval achieves a remarkable 0.929 Pearson correlation in evaluating robot policies, revealing critical insights into their real-world performance.
Fluent demonstrations may mislead robot learning, but a new representation method recovers critical motion insights, boosting performance significantly.
V2A models prioritize text captions over visual cues when generating audio, resulting in physically plausible but often temporally misaligned sounds.
Video LLMs can ace individual traffic video questions but still fail spectacularly at subtle counterfactual reasoning, revealing a critical blind spot for safety-critical applications.
Unified benchmarks reveal the state-of-the-art in simultaneously addressing multiple real-world image degradations like blur, low-light, and rain.
Audio-language models can now reason about 30-minute-long audio clips with timestamp-grounded intermediate steps, unlocking a new level of fine-grained understanding.