Search papers, labs, and topics across Lattice.
4
2
7
10
Atomic visual perception in MLLMs is largely unsolved, with no model surpassing 60% accuracy on a new benchmark designed to isolate perceptual capabilities.
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
Achieve near-dense Video-LLM performance on long videos with up to 57% fewer FLOPs by adaptively selecting which video cubes and tokens to process.
Forget slow visual token concatenation: LaVi modulates LLM features directly with visual context, slashing FLOPs by 94% while boosting speed and accuracy.