Search papers, labs, and topics across Lattice.
8
0
10
Query-grounded visual sampling can boost long video understanding in LVLMs by over 11% without the need for extensive training.
Task-adaptive control of multimodal focus can boost evidence-grounded reasoning performance by over 6 points in critical scenarios.
Overcome the static train-then-freeze paradigm with a new test-time adaptation framework that significantly improves camouflaged object detection in unseen environments.
Quantizing VLAs for robots doesn't have to trash performance: DA-PTQ recovers near full-precision accuracy even at low bit-widths by explicitly minimizing kinematic drift during sequential control.
Revitalizing the latent edge-sensitivity of DINO with SAM's structural priors yields state-of-the-art open-vocabulary segmentation, especially in cluttered scenes.
Forget relying on immediate observations: Keyframe-Chaining VLA lets robots nail long-horizon tasks by remembering and chaining together only the *important* past states.
By intrinsically guiding action refinement through sparse imagination, SC-VLA achieves SOTA performance in robot manipulation tasks, outperforming existing methods by a significant margin in both simulation and real-world settings.
Topological data analysis reveals that structurally stable and compact soft prompts lead to better downstream performance, enabling a new loss function (TSLoss) that improves convergence and tuning performance.