Search papers, labs, and topics across Lattice.
CASIA *Equal contribution
14
32
15
5
PixelEyes achieves precise visual localization by separating reasoning from perception, drastically reducing the redundancy in multi-turn visual searches.
Structured supervision can boost VLA model performance by over 50% in complex robotic tasks, transforming how we approach fine-tuning in manipulation.
Effective service zone design can outperform battery upgrades in profitability, especially under varying demand conditions.
Token-oriented inference optimizations can cut production costs and boost efficiency, transforming large model services from merely callable to fully operable.
HiMPO assigns less-entangled credit to memory updates, significantly boosting long-horizon agent performance while minimizing blame leakage from errors.
LaWAM achieves up to 24x lower latency than traditional pixel-space models while maintaining state-of-the-art performance in robot control tasks.
Achieving photorealistic 3D human avatars from a single image in under a second could revolutionize virtual reality and gaming applications.
Forget object-centric prompts: Function2Scene designs 3D indoor scenes directly from natural language descriptions of *how* the space will be used, not just *what* furniture to put there.
A 440MB multilingual translation model now rivals commercial APIs, opening the door for performant on-device translation.
LLMs can achieve state-of-the-art unsupervised multimodal entity linking by reasoning over diverse evidence types, including graph-based neighborhood information.
ControlFoley lets you generate audio from video with unprecedented control over text descriptions and reference audio, even when those inputs conflict.
Forget complex memory architectures: simple retrieval and generation, when carefully tuned for signal density, can outperform sophisticated methods in conversational agents.
LLMs can be jailbroken with 90% success by subtly "salami slicing" harmful intent across multiple turns, even against state-of-the-art models like GPT-4o and Gemini.
The largest open-source image generative model to date, HunyuanImage 3.0, achieves state-of-the-art performance using a Mixture-of-Experts architecture and native Chain-of-Thoughts schema.