Search papers, labs, and topics across Lattice.
5
0
10
11
Visual evidence can now actively drive long-horizon multimodal search, transforming how agents interact with complex information.
Camera-radar fusion gets a boost: ConFusion's heterogeneous query interaction surpasses existing methods, achieving state-of-the-art 3D object detection performance on nuScenes.
Stop reinventing the wheel: OpenWorldLib offers a unified framework and codebase for advanced world models, finally bringing standardization to a fragmented field.
LLMs can be coaxed into admitting "I don't know" more often, slashing hallucination rates, simply by tweaking prompts to reward humility and truthfulness.
Current multimodal agents are surprisingly bad at web browsing, achieving only 36% accuracy on a new benchmark designed to test deep, multi-modal reasoning across web pages.