Search papers, labs, and topics across Lattice.
43
0
22
14
GlaKG achieves near-perfect classification while providing a transparent reasoning framework that links biomarker evidence to clinical rules, addressing the black-box nature of traditional deep learning models in healthcare.
ProCon achieves unprecedented anomaly detection accuracy without the need for training or pseudo-anomaly supervision, redefining the capabilities of memory-based methods.
Bridge-WA achieves superior task performance by predicting where and how the world will change, enabling robots to focus on relevant scene dynamics rather than irrelevant visual details.
Hypic slashes time-to-first-token by 2.45x and doubles throughput for hybrid-attention LLMs, all while preserving near-full accuracy.
Runtime diagnoses from multi-faceted bug reproduction tests can significantly boost patch generation effectiveness, leading to a 75.7% resolution rate on verified issues.
ICMPG achieves a groundbreaking balance between semantic fidelity and physical realism in motion synthesis, outperforming traditional methods in both standard and zero-shot scenarios.
SR-PPO achieves significant gains in reasoning tasks by effectively assigning credit to individual tokens from a single rollout, transforming how we approach reinforcement learning in language models.
Sentence-level contextual entrainment can skew inference probabilities, but selectively disabling just a few attention heads can mitigate this effect without sacrificing performance.
TokenMinds reveals that combining discrete SID-based user tokens with dense embeddings can significantly enhance user modeling in recommender systems at scale.
Strong proprietary models falter in grounding their predictions, revealing a critical flaw in current VideoQA systems that could reshape evaluation standards.
Preserving skill-level attention structures in MLLMs can dramatically reduce forgetting while adapting to new tasks without relying on replay mechanisms.
ARP not only aligns visual observations with action representations but also refines execution precision, leading to unprecedented performance in robotic manipulation tasks.
Malicious instructions hidden in images can bypass existing skill scanners, exposing a critical vulnerability in LLM-based systems.
Current vision-language models struggle with process understanding in robotic manipulation, but targeted post-training can yield significant improvements.
Natural backdoor vulnerabilities are not just a theoretical concern; they are prevalent in CodeLMs and can significantly compromise code security.
Superficial reasoning in video temporal grounding can be transformed into high-quality, time-aware insights with the right optimization framework.
Generating realistic 3D environments from satellite imagery in under 10 minutes could revolutionize how we visualize and interact with our planet.
The new REO framework reveals that the true challenge in differential equation discovery lies not just in recovering equations, but in leveraging them to reshape scientific understanding.
Text world models can transform LLM-based agents from reactive responders into proactive planners, enhancing their performance in complex interactive tasks.
Transforming the KV cache from a monolithic structure into a dynamic, head-aware system could revolutionize LLM serving efficiency and scalability.
Achieving nearly 50% Recall@1 in video retrieval without any training marks a significant leap in efficiency and effectiveness for complex user queries.
MLLMs can be manipulated to produce harmful outputs from benign inputs, exposing a critical vulnerability in their safety mechanisms.
MAAD not only automates architecture design but also enhances the quality of outputs through a collaborative agent framework and advanced LLM integration.
APEIRIA bridges the gap between interpretable neuro-symbolic reasoning and the flexibility of multi-modal language models, achieving superior performance in 3D spatial reasoning.
Disinformation detection gets a major upgrade with ExTax, a framework that doesn't just flag fake news, but explains *how* it manipulates you through persuasion, emotion, and narrative.
Finance LLM agents can now block unauthorized actions mid-trajectory without sacrificing performance, thanks to a novel inline safety harness that adaptively routes verification between lightweight and advanced LLM judges.
Solving Poisson equations just got faster and more stable: NPSolver trains neural operators without solution labels by iteratively refining predictions with preconditioned conjugate gradient steps.
Reconstructing high-fidelity 3D heart models from noisy radar data is now possible, thanks to a novel mesh deformation approach that leverages physics-informed learning.
Counterintuitively, letting radar cardiac sensors learn to mimic ECGs first yields far better performance on downstream tasks like blood pressure regression and waveform segmentation than directly training on those tasks.
Code dataset watermarking gets a stealthy upgrade: PuzzleMark hides watermarks in variable names based on code complexity, making them nearly undetectable while guaranteeing perfect verification.
Federated learning can overcome data silos, but struggles when clients have different label relationships; FedHarmony shows how to harmonize these differences, leading to better performance.
Today's best vision-language models are surprisingly bad at reading scientific figures, failing to match expert-level reasoning on a new benchmark of experimental images.
Forget fully connected relation graphs: CasLayout's sparse relation modeling unlocks enhanced controllability and realism in 3D indoor scene synthesis.
Simple, artist-friendly quad meshes can now be automatically generated on 3D shapes using a diffusion model trained on a continuous surface representation, sidestepping the complexity of discrete mesh optimization.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.
MLLMs are better at understanding videos than directly grounding text queries within them, and a self-correction training loop can close the gap.
LLMs disperse similar prompts instead of clustering them, leading to significant prompt sensitivity that challenges stability and reliability.
Ditching caches for compiler-managed data streams, Li Auto's M100 architecture achieves higher utilization than GPUs on autonomous driving tasks, hinting at a new path for efficient AI inference.
GitHub abuse is more widespread and varied than previously thought, demanding a unified detection approach to safeguard software supply chains.
RL fine-tuning of discrete diffusion models can be made dramatically more stable and effective by treating the final denoised sample as the action and reconstructing trajectories using the forward diffusion process.
LLMs still struggle to understand the meaning of common phrases, idioms, and compound words, revealing critical gaps in semantic reasoning.
Imagine creating high-fidelity, navigable 3D worlds from just a text prompt or a single image – HY-World 2.0 makes it a reality.
PID controllers can now be enhanced with RL-derived gains, allowing for robust performance in uncertain environments without losing interpretability.