Search papers, labs, and topics across Lattice.
3
0
5
9
A 14M parameter model that outperforms larger transformers while being 12 times smaller, reshaping the efficiency landscape of audio-visual processing.
USAD 2.0 achieves state-of-the-art audio understanding by seamlessly integrating self-supervised and supervised learning techniques, scaling to one billion parameters.
Forget finetuning: TTA-Vid adapts video reasoning models to new datasets *during inference* using test-time reinforcement learning, achieving state-of-the-art results without any labels.