Search papers, labs, and topics across Lattice.
11
0
10
7
Over 1,500 submissions revealed stark differences in model performance across diverse domains, highlighting the challenges of generalizing egocentric video understanding.
MixDiffusion allows for the integration of multiple control conditions in text-to-image generation, breaking the constraints of traditional single-condition models.
Achieving an 80.9% success rate in spatial reasoning tasks, RoboSpatialBrain reveals the critical role of selective reasoning activation in embodied AI.
Personalized content generation can be dramatically improved by transforming user interaction history into actionable instructions, leading to more relevant and visually appealing outputs.
DramaDirector achieves unprecedented fidelity in short drama generation by leveraging real cinematographic geometry, outperforming traditional video generation methods.
TailorMind outperforms traditional content generation methods by creating personalized multimodal outputs without needing existing content pools, achieving significant gains in coherence and novelty.
Skill-Compositional Experts can dramatically enhance a robot's ability to learn new tasks without forgetting old ones, tackling a critical barrier in embodied AI.
ChartLens combines data correction and summary refinement to achieve state-of-the-art performance in chart understanding, outperforming existing solutions.
Mamba strikes again, enabling VLA models to learn more robust manipulation policies that generalize better to real-world scenarios and require less training data.
Escaping the endless cat-and-mouse game of deepfake detection may be possible by shifting from static pattern recognition to physics-inspired dynamical stability analysis, where real images are stable and deepfakes are not.
Robotic manipulation gets a serious upgrade: ConsisVLA-4D boosts performance by up to 41.5% and speeds up inference by 2.4x, all while ensuring your robot understands the scene in 3D *and* how it changes over time.