Search papers, labs, and topics across Lattice.
4
0
10
Depth integration in audio-visual segmentation leads to over 10% performance gains, revealing a critical yet overlooked modality in multimodal perception.
Fine-tuning a model on rigorously synthesized tasks can outperform larger models by leveraging high-fidelity data, achieving a new benchmark in terminal agent performance.
Jointly training MTP and RL doesn't have to hurt: a simple coefficient calibration scheme unlocks performance gains on mathematical reasoning tasks.
Today's visual generation models are often evaluated on the wrong things, leading to inflated performance claims that mask critical failures in spatial reasoning, temporal consistency, and causal understanding.