Search papers, labs, and topics across Lattice.
Xiaomi EV
18
0
12
6
Future-aware geometry tokens enable autonomous vehicles to make safer and more efficient driving decisions, outperforming existing models.
SPG-Layout achieves a breakthrough in 3D scene synthesis by generating physically plausible layouts in non-Manhattan environments, outperforming existing methods.
FlowCIR slashes training resource requirements by 90% while boosting robustness against negation in zero-shot image retrieval tasks.
AVTok achieves superior audio-video synchronization and reconstruction, setting a new standard for unified multimodal generation.
LISA accelerates training and enhances output quality in visual-condition generation by aligning side network features with likelihood scores, all without extra inference costs.
Advanced RAG methods like GraphRAG and Agentic RAG can reduce token usage by up to 53%, but they don't always enhance generation quality as expected.
SPAR bridges the critical gap between semantic perception and pixel-level generation, achieving unprecedented quality in visual outputs without external supervision.
CP4D achieves photorealistic 4D scene generation by seamlessly integrating static environments with dynamic objects, outperforming existing methods in visual fidelity and physical consistency.
OneVLA unifies navigation and manipulation tasks into a single framework, enabling robots to seamlessly interpret commands and interact with their environments like never before.
Fine-grained 3D object grounding gets a boost: SSR3D-LLM uses latent spatial reasoning steps to iteratively refine candidate rankings, outperforming single-pointer methods and setting a new standard for unified 3D-LLMs.
Decoupling radial and angular dynamics in vision-language model adaptation unlocks significant gains in few-shot performance, outperforming existing flow matching methods.
Endowing VLMs with intrinsic 3D geometric awareness and physical interaction cues via XEmbodied substantially boosts performance on spatial reasoning and embodied tasks, surpassing existing 2D image-text pretrained models.
Latent reasoning can beat explicit Chain-of-Thought – but only if you force it to learn causal dynamics via a visual world model, not just language.
Image-goal navigation gets a boost from hierarchical reasoning, using vision-language models for high-level planning and online RL for low-level execution, significantly reducing wandering and improving success in complex environments.
Achieve state-of-the-art hyperspectral image denoising by adaptively balancing data fidelity and noise priors, outperforming existing methods that overemphasize image priors.
Achieve high-fidelity image editing without sacrificing source fidelity by straightening the latent trajectory and adaptively blending source and target velocities.
Autonomous driving models no longer need to compromise between spatial perception and semantic reasoning: UniDriveVLA's expert decoupling unlocks state-of-the-art performance across a range of driving tasks.
Ditch language descriptions: this new driving model leverages dense 3D geometry for superior autonomous driving performance and cross-camera generalization.