Search papers, labs, and topics across Lattice.
9
3
11
44
Switching between autoregressive and diffusion modes allows Nemotron-Labs-Diffusion to achieve unprecedented throughput and efficiency in language modeling.
Dynamic token editing in image synthesis could redefine how we approach high-resolution generative models.
A single generalist model outperforms specialized systems, achieving over 35% improvement in real-world robotic task success.
SR-REAL's dual-path reasoning framework allows spatial VLMs to excel in both linguistic deduction and 3D geometric inference, significantly enhancing performance on complex spatial reasoning tasks.
Fine-tuning on the new ProCUA-SFT dataset boosts UI-TARS 7B's performance from a dismal 8-10% to an impressive 45.0% on OSWorld tasks, highlighting the critical role of high-quality training data.
Achieving six times the inference throughput of current LLMs while maintaining accuracy, Nemotron 3 Ultra redefines performance benchmarks for agentic reasoning tasks.
GRAIL achieves an impressive 84% success rate in real-world object pick-up tasks using only synthetic data, revolutionizing humanoid robot training.
Grounding boosts spatial reasoning in VLMs: explicitly linking language to 2D and 3D scene elements lets models decompose complex spatial problems and improve performance even on non-grounded tasks.
MLLMs can now handle 4K videos up to 100x faster thanks to AutoGaze, which selectively attends to only the most informative patches.