Search papers, labs, and topics across Lattice.
7
0
8
2
This work examines what structure is supplied by the game or workflow, what AI learns or produces, which capabilities and artifacts transfer across settings and roles, and what evidence supports the claims and identifies cross-role connections.
MLLMs can recognize urban scenes but fail to maintain reliable navigation and goal-directed behavior over extended exploration in complex environments.
Lifecycle-wide perception is crucial for evidence-grounded scientific discovery, allowing OmniScientist to outperform traditional AI systems in research quality and depth.
State-of-the-art multimodal models falter in interpreting implicit social cues, revealing a critical gap in AI's understanding of human communication.
Culture-specific emotional perception can be effectively integrated into MLLMs, but current models struggle to achieve even 50% accuracy on this nuanced task.
Overcome the scarcity of 4D training data by cleverly borrowing spatial understanding from 3D models and temporal dynamics from video models.
Forget unimodal tasks鈥擴niM throws down the gauntlet for truly unified multimodal AI, demanding models juggle any combination of text, image, audio, video, code, documents, and 3D inputs and outputs in a single, interleaved stream.