Search papers, labs, and topics across Lattice.
Affiliation:
4
3
7
8
VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.
Lightning Weave is introduced, a post-training framework that extracts and composes independently learned capabilities in a single student through on-policy distillation and achieves a state-of-the-art accuracy-efficiency frontier.
Trial Parallelism accounts for over 65% of reasoning computation in LLMs, and harnessing it can lead to significant speedups in problem-solving.
Switching between autoregressive and diffusion modes allows Nemotron-Labs-Diffusion to achieve unprecedented throughput and efficiency in language modeling.