Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
VLMs struggle to provide accurate annotations in video games, revealing significant gaps in their understanding of dynamic environments.
Annotating video game datasets with VLMs can transform the way RL agents learn by simplifying reward extraction and conditioning.
NeuPAT recovers nearly all lost language capabilities in multimodal LLMs while ensuring robust performance across diverse tasks.
LSAR transforms audio retrieval by mapping audio directly into a sparse lexical space, enabling efficient and interpretable retrieval without the pitfalls of transcription.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
NeuPAT recovers nearly all lost language capabilities in multimodal LLMs while ensuring robust performance across diverse tasks.
LSAR transforms audio retrieval by mapping audio directly into a sparse lexical space, enabling efficient and interpretable retrieval without the pitfalls of transcription.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
Grounding tasks can be unified into a single correspondence prediction problem, leading to a 48% boost in performance on complex datasets.
MirrorWorld achieves unprecedented fidelity in mirror reflection generation by effectively modeling both what to reflect and how to spatially arrange it.
AutoPrune enables LLMs to autonomously design visual-token pruning strategies, achieving a remarkable 9.9x reduction in FLOPs without sacrificing performance.
AtlasVLA outperforms multi-view baselines by over 17% in long-horizon tasks, showcasing the power of proactive reasoning in embodied AI.
WorldTrace redefines memory management in video world models, achieving up to 19.5% better episodic recall without any retraining.
Capek 0.5 reveals that organizing embodied capabilities by functional roles can significantly enhance performance in complex execution tasks.
A multi-agent forensic reasoning framework outperforms leading closed-source models in deepfake detection by leveraging diverse analytical perspectives on forgery cues.
Artist-grounded image generation can achieve unprecedented fidelity by explicitly controlling for artistic intent, rather than relying on shortcuts that distort user vision.
Visual tool-use in multimodal LLMs may create an illusion of effectiveness, with many models achieving gains that are not causally justified.
PaDoc achieves a remarkable 94.24 F1 score among end-to-end parsers while being the fastest parser at five concurrency levels, revolutionizing document parsing efficiency.
Achieving 74.8% accuracy on a new temporal reasoning benchmark, ChronoVision redefines how multimodal models can tackle complex visual tasks.
Self-PreTraining boosts transformer accuracy in medical time series by up to 6 percentage points, even with limited data.
Bridging implicit and explicit relational biases boosts skin lesion diagnosis accuracy by nearly 3% using graph-based methods.
VLMs struggle to provide accurate annotations in video games, revealing significant gaps in their understanding of dynamic environments.
Annotating video game datasets with VLMs can transform the way RL agents learn by simplifying reward extraction and conditioning.
Ground condition classification achieved with 95% accuracy using proprioceptive sensors, enabling real-time gait adaptation in autonomous robots.
KVAE tokenizers achieve superior generative performance, setting a new benchmark for multimodal models in audio, image, and video synthesis.
SR-JEPA reveals that a predictive pathway can infer missing entity representations in 3D scenes with remarkable accuracy, transforming how we approach latent state learning.
Depth information can drastically improve object counting accuracy in crowded scenes, achieving over 60% reduction in counting errors.
PRISM achieves state-of-the-art results in unpaired image translation by intelligently preserving important features while allowing for targeted changes, outperforming existing methods.
VLMs may excel in scoring but often lack meaningful visual grounding, revealing critical limitations in their zero-shot control capabilities.
GSBF achieves beamforming without the need for instantaneous channel state information, dramatically reducing latency and complexity in MIMO systems.
Multimodal inputs can enhance reasoning in executive decisions, but their indiscriminate use may paradoxically undermine resource allocation effectiveness.
Detectability of findings in 3D CT scans hinges more on physical characteristics than model architecture, revealing a critical bottleneck in diagnostic performance.
Existing models miss critical visual evidence in metaphor understanding, but M$^3$R-Reasoner closes the gap, outperforming larger models in both accuracy and justification metrics.
Integrating VPR with FF3D models boosts localization accuracy, overcoming the limitations of traditional visual methods.
Uncertainty-guided feedback can dramatically improve the reliability of reward models in visual diffusion, leading to superior optimization and quality outcomes.
EmoWorld achieves up to 37% improvement in emotional alignment while decoupling atmosphere, semantics, and temporal progression in video generation.
CFGPNet achieves up to 97.8% mAP in multispectral object detection, setting a new standard for accuracy and efficiency in challenging conditions.
Event-driven techniques reveal hidden dynamics, enabling EvReflection to outperform traditional methods in reflection removal by a significant margin.
HoloWorld achieves a groundbreaking integration of indoor and outdoor urban generation, enhancing spatial coherence and visual identity across entire cityscapes.
Noise-aware residual correction boosts the realism of autoregressive audio-visual generation, tackling issues of identity drift and desynchronization head-on.
Curia-MAE achieves superior performance in 3D medical image segmentation with a frozen encoder, challenging the need for extensive fine-tuning even in data-scarce environments.
VLMs can transform under-resourced historical languages by automating data extraction at unprecedented scales, as demonstrated by the mapping of Armenian commercial advertisements in Paris.
HALO's innovative dual-prior approach eliminates attention drift, achieving unprecedented clarity and color accuracy in low-light remote sensing imagery.
Retailers can now generate highly accurate virtual try-ons that reflect true garment fit, reducing misleading representations in online shopping.
Achieving seamless identity replacement in videos, Vorch-IR can handle multiple subjects and backgrounds without requiring precise pose matching.
TruthLens reveals that a simple self-evaluation mechanism can dramatically enhance the accuracy of object detection in LVLMs, outperforming existing methods by a significant margin.
Decoupling coordinate frame selection from box regression leads to a remarkable 11% accuracy boost in 3D visual grounding tasks.
Models that generate convincing anomaly descriptions often fail to accurately track the corresponding instances, exposing a critical gap in video anomaly understanding.
Achieving real-time audio-video generation at 27.12 FPS, Vorch-Streamer tackles the dual challenges of exposure bias and causal speech generation in long-form content.
Tile-based background refinement in 360-degree telepresence can dramatically enhance perceived detail and interactivity, surpassing traditional video resolution techniques.
Real-time MRI-guided interventions are now possible with a master-slave robotic system that allows for unprecedented control and precision in needle-based procedures.
Coordinating global and local reasoning in long-video understanding leads to a significant 2.9 point improvement over traditional frame selection methods.
Aligning robot scene geometry with ARGUS enables manipulation policies to learn 4-6 times faster from diverse viewpoints, transforming their generalization capabilities.
Adversarial prompts can hijack VLM-controlled robots with a success rate of up to 29%, exposing a critical vulnerability in their operational integrity.
KILVO outperforms state-of-the-art fusion methods, achieving high accuracy and robustness even in the face of sensor failures.
Running a powerful omni-modal search engine entirely on-device could redefine user privacy and performance in local data retrieval.
A single global weight outperforms personalized modality weighting in multimodal recommenders, raising questions about the validity of personalization claims at scale.
Real-time character animation is now feasible with Wan-Animate-2, which achieves high-fidelity results without the pitfalls of traditional motion representation methods.
EviSelect achieves a 3.9x speedup in long video understanding by dynamically selecting relevant frames based on the MLLM's internal attention evidence, cutting visual token selection by half.
HOPE achieves accurate pressure estimation from monocular videos, enabling robust predictions of hand-object interactions without the need for specialized sensors or extensive labeled data.
A single model can now seamlessly handle over 10 diverse audio-visual tasks without the need for task-specific architectures.
Self-supervised learning can unlock high-quality data extraction from bar charts without the need for extensive labeled datasets.
Achieving 98.0% success in cross-embodiment manipulation without manual action alignment could redefine how we approach robot control across diverse platforms.
By combining multi-step latent self-prediction with observation-level dynamics, OG-SPR achieves superior performance in visual control tasks, revealing the limitations of traditional predictive methods.
Label-free reliability in vision-language models has a computable blind spot that can be systematically characterized and detected.
A new dataset and model achieve a staggering 4.98% symbol error rate for classical music, setting a high bar for audio-to-score transcription in popular music.
SkillMemo transforms robotic manipulation by enabling models to leverage reusable skill structures, leading to unprecedented compositional generalization in complex tasks.
GAUGE achieves superior multimodal classification by fine-tuning evidence modulation at a granular level, even when faced with incomplete data inputs.
Video language models falter dramatically in counting transient events, with less than 0.2% accuracy in high-frequency scenarios.
BioKD achieves a remarkable 68.01% accuracy in trial-wise arousal recognition, showcasing the power of reliability-aware knowledge distillation in emotion recognition tasks.
Expert-validated data and a compact model make Bangla Sign Language recognition feasible on personal devices, enhancing accessibility for the deaf community.
Structured video-grounded semantics can boost synthetic IMU generation, leading to a 19.86% improvement in tail-class recognition over traditional methods.
Current agentic AI systems fall short, with none demonstrating more than two out of nine critical maturity criteria, exposing a significant gap in their capabilities.
Smart-home agents struggle to differentiate between real commands and misleading ambient noise, with traditional detectors and MLLMs both failing in complementary ways.
CogVis redefines change detection by decoupling temporal perception from semantic categorization, achieving unprecedented efficiency and accuracy across multiple benchmarks.
Unified Agent outperforms existing multi-agent systems by effectively maintaining a compact state across devices and time, revolutionizing user-agent interactions.
Grounding generative editing in physics dramatically enhances shadow removal quality, revealing that classic vision principles are still vital in AI-driven image editing.
ECG-LENS outperforms existing ECG report generation systems by integrating multi-lead signal modeling with advanced clinical context, achieving unprecedented report quality.
Implicit semantic guidance in UniVVT outperforms traditional geometric preprocessing, setting a new standard for high-fidelity video virtual try-on.
Routing supervision falters precisely when it's most needed, as weaker agents yield fewer labels, limiting the potential for optimal mode selection.
Operation laundering in vision encoders can be effectively mitigated, revealing hidden boundaries in learned assignments that traditional methods obscure.
Prior-SG enables robots to redefine their spatial understanding in real-time, achieving zero-shot flexibility in semantic region segmentation even in the absence of physical boundaries.
DARAD not only adapts to new remote sensing data but also preserves the integrity of historical retrieval performance, a dual capability that sets a new standard in continual learning.
Integrating historical predictions into geospatial models boosts crop classification accuracy by over 1.6 percentage points, correcting significant recall biases.
Automated tooth-level mapping from smartphone images could revolutionize access to dental care in resource-limited settings.
Achieving an FID of 1.45 in just 600 epochs, Energy-Guided Flow Matching redefines efficiency in high-quality image generation without the need for extensive model adaptations.
Fine-tuned foundation models can significantly outperform human-inspired methods in compositional analysis, but at the expense of interpretability and generalization.
ConceptADapt achieves significant performance gains in few-shot anomaly detection by recalibrating feature statistics through a novel dynamic attention mechanism, even with minimal training data.
Selective convergence in self-supervised learning allows models to prioritize salient audio-visual cues, leading to superior performance in multi-source localization without manual annotations.
MAVISEG reveals that diffusion transformers can retain and utilize more structured visual information than traditional methods, leading to significant improvements in segmentation accuracy.
STAIL achieves significant performance gains in medical imaging tasks while drastically minimizing storage requirements and privacy concerns associated with traditional rehearsal methods.
TTA can boost accuracy but often at the cost of calibration, and ZAEC is the key to restoring reliable confidence without labeled data.
Open-vocabulary segmentation can achieve spatially coherent and context-aware predictions without any training, thanks to SCI-CLIP's innovative use of a segment-centric inference framework.
Robust-WAM achieves superior out-of-distribution generalization in robot control by seamlessly integrating semantic foresight into action predictions while leveraging extensive VGM pretraining.
Achieving state-of-the-art HDR reconstruction quality, DOME-HDR reveals how dual-output synthesis can enhance both SDR and HDR imaging from bracketed inputs.
A single ranking can adapt to any frame budget, improving accuracy and reducing latency without retraining the model.
Pretraining on relevant scientific images can significantly boost quality assessment performance, outperforming larger datasets.
Uncertainty-aware segmentation can drastically improve performance in medical imaging, with DistMedVL achieving superior results using only 6.3M parameters.
MultiMoQ achieves smoother viewport playback by increasing goodput and reducing latency, even under challenging network conditions.
Reducing token overhead by pruning redundant communication edges allows multi-agent systems to achieve better performance without the computational burden.
Ambient temperature fluctuations can be harnessed to create robust adversarial attacks that consistently deceive multimodal perception systems.
$\omega$-0 enables humanoid robots to seamlessly integrate movement and manipulation, outperforming traditional models by predicting coordinated actions directly from sensory inputs.
F$^2$Agent achieves over 20% better annualized returns than existing models by dynamically capturing inter-modality dependencies and resisting market noise.
Hierarchical post-training can significantly enhance robotic manipulation by enabling agents to better navigate complex tasks through effective subgoal decomposition.
Grounded language comprehension, rather than free-form reasoning, is the key to unlocking superior performance in Vision-Language-Action models.