Search papers, labs, and topics across Lattice.
28 papers from NVIDIA Research on Architecture Design (Transformers, SSMs, MoE)
MicroEvo achieves a staggering 10.6x increase in search efficiency while improving Pareto-front quality by 36.2%, revolutionizing microarchitecture design exploration.
ShimGen not only matches but surpasses manually-designed protocols in consistency, revealing critical performance gains in heterogeneous memory systems.
Hierarchical local attention in TextNCA reveals that the arrangement of attention windows can dramatically influence language modeling performance, even more than the model's iterative nature.
MMA-Former's innovative attention mechanism enables spatially adaptive feature extraction, significantly improving PNI prediction accuracy from 3D MRI scans.
Gauntlet outperformed human reviewers in technical critique of computer architecture papers, revealing that LLMs can achieve significant analytical depth through structured multi-agent collaboration.
MAESTRO prunes MoE models more effectively by leveraging the interdependencies of expert routing, achieving superior performance retention and consistency across tasks.
Switching between autoregressive and diffusion modes allows Nemotron-Labs-Diffusion to achieve unprecedented throughput and efficiency in language modeling.
Achieving up to 1.90X speedup in video generation without sacrificing fidelity, ScalingAttention redefines efficiency in Diffusion Transformers.
Achieving six times the inference throughput of current LLMs while maintaining accuracy, Nemotron 3 Ultra redefines performance benchmarks for agentic reasoning tasks.
MSA slashes per-token attention compute by over 28x while maintaining competitive performance, revolutionizing how LLMs can handle ultra-long contexts.
Forget scaling laws: a single looped transformer block, iterated explicitly, crushes billion-parameter feed-forward networks at multi-view 3D reconstruction.
Get up to 1.79x faster ViT inference on high-resolution images without sacrificing accuracy by surgically replacing full-attention blocks with cheaper alternatives *after* pre-training.
Ditch slow, token-by-token box generation: LocateAnything's Parallel Box Decoding (PBD) boosts VLM grounding speed and accuracy by decoding entire bounding boxes at once.
Forget everything you thought you knew about linear attention: decoupling erase and write operations unlocks significantly better long-context retrieval.
Forget assuming NaNs and single-bit flips are the main culprits in GPU silent data corruption; this study reveals they're surprisingly rare, demanding a rethink of fault modeling.
Looping language models isn't just for single agents anymore: Recursive Multi-Agent Systems (RecursiveMAS) show that agent collaboration itself can be scaled through recursion, yielding faster and more efficient problem-solving.
Forget GPU-centric designs: AMMA slashes attention latency by 15x and energy consumption by 7x with a memory-centric architecture for long-context LLMs.
Multimodal models can now achieve state-of-the-art performance in real-world tasks like document understanding and audio-video comprehension with significantly reduced inference latency thanks to novel token-reduction techniques.
Forget clunky animation pipelines – MotionBricks lets you assemble real-time, high-quality character motions like LEGOs, even controlling robots.
Open-vocabulary 3D instance segmentation just got 100x faster, thanks to a new transformer architecture that ditches region proposals and fragmented masks.
Squeeze up to 3.2x more performance from your long-context LLMs by intelligently splitting attention computation between CPU and GPU.
Speech-to-speech translation can now convey laughter and tears with human-like fidelity, thanks to a surprisingly data-efficient approach leveraging LoRA experts.
Nemotron 3 Super proves you can achieve comparable accuracy to existing 120B models, but with significantly higher inference throughput, by combining Mamba, Attention, and Mixture-of-Experts.
Gaussian Splatting gets a high-frequency boost: Neural Harmonic Textures unlock significantly more detail in primitive-based 3D reconstructions without sacrificing speed.
Stop wasting precious GPU memory: this new cache-semantic hash table library achieves up to 3.9 billion key-value lookups per second, outperforming standard approaches by up to 9.4x.
Training trillion-parameter Mixture-of-Experts models just got a whole lot faster: Megatron Core now achieves >1 PFLOP/GPU on NVIDIA's latest hardware.
Forget monolithic LoRAs: LoRWeB dynamically mixes a basis set of LoRAs to unlock SOTA generalization in visual analogy tasks.
Achieve state-of-the-art depth completion by adapting 3D foundation models at test time with minimal parameter updates, outperforming task-specific encoders that often overfit.