Search papers, labs, and topics across Lattice.

Research division of NVIDIA focusing on GPU-accelerated AI, computer graphics, robotics, and autonomous systems.
100
0
0
Trial Parallelism accounts for over 65% of reasoning computation in LLMs, and harnessing it can lead to significant speedups in problem-solving.
PhysCaP enables robots to actively infer hidden physical properties, achieving superior manipulation efficiency without additional sensors.
Linear masking can lead to critical components being entirely omitted from dynamical models, but a new algebra-based scoring method recovers these components with significantly fewer coordinates.
FedCoRe can recover nearly half of the performance lost due to missing ECG or CXR data in federated healthcare models, transforming how we handle incomplete patient information.
UniProbe slashes object hallucinations in LVLMs by 55% during generation, all while operating with minimal latency.
GatorOnco outshines traditional models by achieving expert-level treatment planning performance while enhancing readability and completeness in colorectal cancer care.
WorldTrace redefines memory management in video world models, achieving up to 19.5% better episodic recall without any retraining.
ShimGen not only matches but surpasses manually-designed protocols in consistency, revealing critical performance gains in heterogeneous memory systems.
MicroEvo achieves a staggering 10.6x increase in search efficiency while improving Pareto-front quality by 36.2%, revolutionizing microarchitecture design exploration.
The accuracy gap in multilingual reasoning can swing dramatically by up to 57 points based on output-token caps, challenging conventional evaluation methods.
Hierarchical local attention in TextNCA reveals that the arrangement of attention windows can dramatically influence language modeling performance, even more than the model's iterative nature.
Geometry-guided supervision enables robots to generalize actions across viewpoints, achieving robust manipulation even with limited camera coverage.
Coordinate-wise memory control through CMG enables unprecedented accuracy in forecasting quantum dynamics, outperforming traditional methods by over 91%.
The rise of LLM agents threatens to erase the clarity of authorship and accountability in collaborative knowledge work, raising urgent questions about intellectual integrity.
Achieving state-of-the-art video generation with significantly improved diversity and speed, PDD redefines the efficiency of diffusion models.
Fine-tuned robotic policies can be biased towards certain instruction factors, but a new bias-aware data strategy can significantly enhance their performance with fewer demonstrations.
Jointly optimizing visual token selection and LLM computation can drastically reduce inference costs while enhancing performance in multimodal tasks.
AITE autonomously discovers wireless communication algorithms that not only outperform existing solutions but also dramatically reduce computational latency.
Scaling visuomotor context to 8K timesteps enables robots to master complex tasks and adapt in real-time, outperforming previous models by a staggering margin.
AeroAct achieves real-time language-conditioned quadrotor navigation by predicting flight actions directly from visual and language inputs, without the need for future video generation.
Structured feedback can boost LLM agent success rates by up to 44 percentage points, revealing the critical role of admissible alternatives in the repair process.
Runtime attacks on GPU kernels can be effectively mitigated with WarpGuard, the first framework to provide joint attestation for CPU-GPU workloads.
LoSA-Net outperforms existing models in predicting perineural invasion by effectively preserving crucial boundary details in 3D MRI scans.
Gauntlet outperformed human reviewers in technical critique of computer architecture papers, revealing that LLMs can achieve significant analytical depth through structured multi-agent collaboration.
MMA-Former's innovative attention mechanism enables spatially adaptive feature extraction, significantly improving PNI prediction accuracy from 3D MRI scans.
Adaptive routing in transformer-based models can drastically enhance the efficiency of PNI prediction while maintaining high accuracy in capturing subtle imaging features.
NAIS can autonomously conduct complex biomedical research while maintaining rigorous governance and human oversight, achieving results on par with traditional expert-led studies.
ARDY achieves real-time, controllable 3D human motion generation that outperforms existing methods in both fidelity and flexibility.
Audex achieves state-of-the-art audio understanding and generation while maintaining the reasoning prowess of its text-only foundation, all through a unified architecture.
CGVQ achieves a remarkable 20% reduction in bits per pixel while maintaining visual quality, revolutionizing Gaussian-based image compression.
Current avatar systems are more diverse than ever, yet foundational prior learning is often overlooked in discussions of photorealistic digital humans.
Current VLMs struggle with specialized domains, failing to adapt effectively in both zero-shot and ICL scenarios, revealing critical gaps in their spatio-temporal reasoning abilities.
ROSA revolutionizes robot factory operations by boosting productivity up to 12.06x through innovative shared GPU-pool serving and factory-focused scheduling.
STRATA achieves 50x better energy efficiency than conventional storm-resolving models while delivering realistic global weather simulations at kilometer-scale resolution.
Claw-like agents are vulnerable to severe security breaches, with malicious plugins achieving a 100% success rate in attacks.
Tri-serve redefines energy efficiency in multimodal inference by addressing hidden power inefficiencies, achieving a 22% boost without latency trade-offs.
ASR models can exhibit drastically different performance depending on user preferences, revealing hidden quality disparities in traditional benchmarks.
Policies trained in SimFoundry's automated environments achieve up to 40% higher success rates in real-world tasks by leveraging affordance-preserving scene variations.
Physically aligned video models can boost robotic manipulation success rates by over 50% compared to traditional methods.
NeuMatEx outperforms PBR techniques by extracting complex neural materials with unprecedented visual fidelity and precision from multi-view images.
Intent-aware training can dramatically enhance safety classifier performance, revealing that faithful intent modeling is a powerful supervision signal.
Distinct model capabilities reveal that relational context significantly influences mental health assessments, with Claude-3-Haiku and GPT-4o leading in classification and trigger detection, respectively.
SR-PPO achieves significant gains in reasoning tasks by effectively assigning credit to individual tokens from a single rollout, transforming how we approach reinforcement learning in language models.
Current multimodal conversational models miss critical emotional cues, revealing a significant gap in their ability to engage in nuanced human-like dialogue.
FAR-LIO cuts odometry latency by 38.4% while improving accuracy, setting a new standard for real-time performance in autonomous racing.
AI data centers can now act as dynamic grid assets, reducing peak electricity demand while ensuring priority workloads remain unaffected.
SHERLOC boosts code repair agents' effectiveness by improving fault localization accuracy while slashing token usage by over 23%.
GRAFT enables robots to manipulate unseen objects with just one demonstration by leveraging geometric similarities, outperforming traditional semantic retrieval methods.
Abandoning arithmetic logic for string similarity allows LLMs to achieve unprecedented accuracy in deducing logical rules from binary strings.
Achieving up to 1.90X speedup in video generation without sacrificing fidelity, ScalingAttention redefines efficiency in Diffusion Transformers.
Bagpiper-TTS can seamlessly transform natural language requests into high-quality speech across diverse applications, outperforming traditional TTS systems.
Self-correcting models can achieve unprecedented fidelity and plausibility in generative tasks by actively learning from their own alignment errors.
Coding agents can now autonomously refine robotic manipulation policies to achieve a staggering 99% success rate on complex tasks, revolutionizing real-world robotics.
Post-training on synthesized safety-critical scenarios can dramatically enhance the reliability of autonomous driving systems, reducing failures in rare but critical events.
Diffusion-Proof not only surpasses AR LLMs in theorem proving but also solves challenging problems that state-of-the-art models fail to address.
LLMs can pass Taiwanese lawyer qualification exams but struggle with precise legal citations, revealing critical gaps in their legal reasoning capabilities.
Tail latency in LLM serving can be cut by up to 50% without relying on length predictions, reshaping how we optimize inference performance.
Combining learning and geometric optimization, this framework achieves a 60.9% grasp success rate, outperforming traditional methods by a significant margin.
SR-REAL's dual-path reasoning framework allows spatial VLMs to excel in both linguistic deduction and 3D geometric inference, significantly enhancing performance on complex spatial reasoning tasks.
A single-line code change can restore diversity and fidelity in video generation models, outperforming even the original teacher models.
Long-form speech generation can now achieve remarkable coherence and naturalness without the need for extensive retraining on long-form datasets.
Tactile-reactive policies can boost robotic manipulation success rates by over 30% through innovative data collection and a new Mixture-of-Transformers architecture.
Agents can become "addicted" to visible rewards, sacrificing safety for short-term gains, raising alarms about AI alignment in real-world applications.
Extracting action signals from 32,041 hours of human video enables CAIP to outperform leading vision encoders in robotic manipulation tasks by over 30%.
cuTile Rust achieves 7 TB/s for element-wise operations on the NVIDIA B200 GPU, all while ensuring memory safety in GPU kernel programming.
SPARC reduces noisy labels by leveraging task structure, enabling robots to learn from more reliable demonstrations and outperforming traditional methods in real-world applications.
MSA slashes per-token attention compute by over 28x while maintaining competitive performance, revolutionizing how LLMs can handle ultra-long contexts.
VLMs trained on the new 4DP-QA dataset show marked improvements in understanding complex 4D scenes, revealing the critical role of disentangling motion dynamics.
Naively scaling test-time compute is wasteful; strategically allocating it with DIRECT can enhance embodied agent performance while slashing latency by up to 65%.
VIPIR achieves orders-of-magnitude higher throughput for private information retrieval while slashing communication and memory overheads, revolutionizing large-scale database privacy.
Decoupling modality processing in VLA models leads to a staggering 95.2% success rate in complex manipulation tasks, far surpassing traditional synchronous approaches.
DEHP dramatically boosts the success rates of high-precision robotic tasks by dynamically adjusting execution horizons based on task complexity.
City-scale reconstructions with over 1 billion Gaussian splats reveal a breakthrough in multi-GPU efficiency and detail, surpassing current state-of-the-art methods by more than 25 times.
VoLoAgent outperforms traditional manipulation systems by seamlessly integrating planning, monitoring, and recovery in real-time, transforming how robots handle complex tasks in dynamic environments.
Retaining the right evidence before a query can boost long-horizon agent performance by over 70% in F1 score, transforming how we think about memory management in AI.
NVSHMEM's innovative device-side symmetric-memory model could redefine GPU communication strategies, pushing the boundaries of hardware performance.
A real-time generative world model can synthesize complex driving scenarios that traditional simulators struggle to capture, enabling safer and more effective evaluation of autonomous vehicle policies.
Cosmos 3 sets a new benchmark for omnimodal models, outperforming existing state-of-the-art in both Text-to-Image and Image-to-Video tasks.
Static evidence selection fails under budget constraints, revealing the need for adaptive strategies in retrieval-augmented systems.
Achieving zero-shot generalization in robotic grasping across diverse gripper designs could revolutionize how robots interact with their environments.
Reward models trained only on success are fundamentally misaligned with human values, leading to dangerous over-rewarding of poor robot behaviors.
Disagreement among security scanners reveals that 81.9% of flagged skills are identified by only one scanner, challenging the reliability of single-scanner assessments in agent-skill security.
Steering imaginations in video world models can reveal critical failure points in robotic actions that traditional methods might overlook.
Current vision-speech agents are surprisingly bad at mimicking the subtle, real-time audio-visual cues that make human conversation feel natural.
Achieving a 40x speedup in training for deformable simulations could revolutionize real-time applications in robotics and animation.
Forget scaling laws: a single looped transformer block, iterated explicitly, crushes billion-parameter feed-forward networks at multi-view 3D reconstruction.
Masking just 5% of attention heads in vision-language models tanks performance on long-context tasks, revealing a surprisingly sparse and critical set of "multimodal retrieval heads" that attend to both text and images.
Stop wasting compute on redundant code generation attempts: CPPO boosts pass@$K$ by explicitly coordinating exploration across diverse algorithmic strategies.
Get up to 1.79x faster ViT inference on high-resolution images without sacrificing accuracy by surgically replacing full-attention blocks with cheaper alternatives *after* pre-training.
Ditch the clunky tool-use pipelines: STORM teaches video-language models to reason about space and time using *internalized* latent trajectories, slashing inference costs while boosting accuracy.
Relight 3D assets 25x faster with a feed-forward network that distills relightable representations from large reconstruction models, sidestepping expensive per-scene optimization.
Forget assuming NaNs and single-bit flips are the main culprits in GPU silent data corruption; this study reveals they're surprisingly rare, demanding a rethink of fault modeling.
Hierarchical power allocation in datacenters can achieve near-perfect satisfaction ratios, even with oversubscription, by using a novel three-phase QP/LP optimization policy.
Speculative decoding, typically used post-RL, can be integrated directly into RL training loops to accelerate LLM rollout generation by up to 2.5x.
VLN agents can navigate more accurately in zero-shot settings by "looking forward, now, and backward," mimicking human navigational strategies.
Near-field lighting? No problem: 8DNA pre-bakes complex light transport into neural representations, outperforming prior methods with faster inference and lower training costs.
Forget GPU-centric designs: AMMA slashes attention latency by 15x and energy consumption by 7x with a memory-centric architecture for long-context LLMs.
Looping language models isn't just for single agents anymore: Recursive Multi-Agent Systems (RecursiveMAS) show that agent collaboration itself can be scaled through recursion, yielding faster and more efficient problem-solving.
Multimodal models can now achieve state-of-the-art performance in real-world tasks like document understanding and audio-video comprehension with significantly reduced inference latency thanks to novel token-reduction techniques.
See where your citations are coming from with a single command, thanks to CiteRadar's open-source platform that automatically generates interactive maps and detailed researcher profiles from your Google Scholar ID.