Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
Breakthrough prediction accuracy of 97% and zero dural injuries in ex vivo cranial surgeries showcase a new era of robotic precision in high-stakes procedures.
Regime-aware adaptive fusion can boost Bitcoin price prediction accuracy by over 30% compared to traditional methods.
UltraViT achieves a groundbreaking 1.7x speed increase for on-device vision-language model encoding without sacrificing performance.
OmniScope reveals that treating audio and video relevance separately can drastically enhance performance in omnimodal models, achieving remarkable efficiency gains without sacrificing accuracy.
Agents struggle to act effectively in 3D scenes, with none of the eleven evaluated VLMs achieving consistent performance across diverse tasks.
Regime-aware adaptive fusion can boost Bitcoin price prediction accuracy by over 30% compared to traditional methods.
UltraViT achieves a groundbreaking 1.7x speed increase for on-device vision-language model encoding without sacrificing performance.
OmniScope reveals that treating audio and video relevance separately can drastically enhance performance in omnimodal models, achieving remarkable efficiency gains without sacrificing accuracy.
Agents struggle to act effectively in 3D scenes, with none of the eleven evaluated VLMs achieving consistent performance across diverse tasks.
ID-V2V enables seamless video edits while preserving facial identity and performance, revolutionizing the way we approach generative video models.
Codebook collapse is tackled head-on, enabling high-fidelity visual representation quantization that scales effectively with vocabulary size.
Scaling native multimodal pre-training reveals that text-heavy data mixtures require larger models for optimal efficiency, challenging conventional resource allocation strategies.
WMH-aware supervision can significantly boost lesion detection while preserving healthy tissue segmentation, addressing a critical gap in glioma MRI datasets.
TableVerse-100K offers a groundbreaking dataset of 100,000 realistic environments that could redefine how robots learn to manipulate objects in complex, cluttered settings.
Bridging the reasoning gap, X$^3$-OPD enables audio-language models to outperform their text-based counterparts in logical reasoning tasks.
Clustering semantically aligned tasks can drastically improve performance in cooperative multi-tasking, reducing negative transfer and enhancing accuracy.
Different views of the same problem reveal hidden reasoning paths, enabling VLMs to achieve unprecedented accuracy in multimodal reasoning tasks.
Achieving 92.4% accuracy in cognitive impairment detection, this multimodal framework outperforms traditional single-modality approaches and sets a new benchmark for generalization across datasets.
MTL not only slashes training costs but also boosts prediction accuracy across diverse game scenarios, challenging the dominance of single-task models.
BoE transforms candidate selection by leveraging partial verification, significantly enhancing outcomes in vision-language tasks where complete evaluations are unattainable.
GS-Agent transforms natural language into intricate 4D worlds, showcasing a new era of automated creative content generation that rivals traditional manual methods.
Thinkink transforms ideation by enabling seamless interaction with LLMs through handwritten notes and sketches on a shared canvas, enhancing creativity and engagement.
MSBraM achieves unprecedented performance in EEG analysis by mastering the complex interplay of local and global neural dynamics.
Agents relying solely on in-scene cues struggle significantly, achieving only 7.4% success in easy navigation tasks, underscoring the complexity of long-horizon navigation without external guidance.
Achieving an AUROC of 0.878 on the CHB-MIT benchmark, this multimodal EEG model redefines the landscape of seizure detection by enabling robust performance across diverse datasets and conditions.
Unlearning can inadvertently reinforce biases if demographic requests are imbalanced, but FAUN ensures fairness while effectively removing sensitive data from MLLMs.
Player-centric modeling in basketball video analysis reveals that traditional methods miss critical interactions, leading to significant performance improvements with the new PlayNet framework.
Harmful videos paired with benign queries can exploit a critical vulnerability in Video LLMs, leading to nearly half of all attacks succeeding despite high content recognition accuracy.
UnDA achieves superior cross-modal knowledge transfer in medical imaging, even in the absence of paired clinical data, by dynamically suppressing noise from uncertain predictions.
T-STAR reveals that over 1.1 million instance masks and 3.8 million spatio-temporal triplets can significantly improve the understanding of dynamic geospatial scenes in satellite video.
Audiovisual dynamics can enhance identity recognition, achieving a 68.2% reduction in false accepts compared to traditional methods.
Spatial reasoning hallucinations in MLLMs can be drastically reduced by integrating geometric evidence, challenging the effectiveness of traditional mitigation methods.
TrainsBiolab reveals the complexities of manipulating transparent objects in cluttered scenes, providing a goldmine of data that could transform how robots perceive and interact with their environments.
Compressing visual tokens in Omni-LLMs can cut input costs by over half while boosting accuracy beyond traditional methods.
Achieving state-of-the-art stability in Angle of Polarization estimation from a single RGB image could revolutionize how we approach polarization-based applications without specialized hardware.
Unified image restoration methods can now be rigorously compared across real-world conditions, with 20 teams showcasing innovative solutions that push the boundaries of restoration accuracy.
Fine-grained hallucination diagnosis can dramatically enhance the reliability of multimodal language models by revealing the types of hallucinations they produce and how to correct them.
GroupVideo achieves unprecedented fidelity in multi-character video generation, resolving identity confusion and unnatural motions that plague existing methods.
Emotion-oriented MLLMs can now achieve unprecedented accuracy in visual emotional intelligence, revealing critical insights into their limitations and capabilities.
GeoThreat achieves superior transferability and controllability in adversarial attacks on LVLMs, redefining how we assess model robustness in remote sensing contexts.
Modeling action unit relationships can dramatically enhance facial expression recognition across diverse domains, leading to substantial performance gains.
Ms.Forcing achieves a 39.6% speedup in streaming video generation while significantly enhancing quality by intelligently adapting spatial granularity to noise levels.
Webly supervised multi-label recognition can achieve state-of-the-art performance with a novel dual-branch learning approach that tackles label noise head-on.
A lightweight vision-only framework achieves superior face anti-spoofing accuracy without the complexity of multimodal fusion.
The challenge reveals that explainability in deepfake detection is not just a nice-to-have; it’s essential for user trust and effective verification in real-world applications.
MagicMakeup achieves unprecedented regional control in makeup transfer, ensuring high fidelity and identity preservation even across diverse styles and poses.
FA-LAM achieves unprecedented quality in animatable head reconstruction by disentangling the complex tasks of animation and reconstruction, leading to superior performance in fine facial details.
Despite advances in MLLMs, they still struggle with dynamic reasoning, falling far short of human capabilities in interpreting continuous visual cues.
Breakthrough prediction accuracy of 97% and zero dural injuries in ex vivo cranial surgeries showcase a new era of robotic precision in high-stakes procedures.
CCBR empowers users to shape their recommendations through text-driven interactions, outperforming traditional models while maintaining high interpretability.
SDM reveals that pretrained image models can be transformed into powerful tools for understanding complex video dynamics, outperforming traditional methods with less supervision.
Prior Collapse in video editing models can be overcome, leading to state-of-the-art one-shot editing performance without sacrificing generative quality.
VLM-IE3D achieves state-of-the-art performance in 3D tasks by seamlessly integrating implicit and explicit geometric representations from RGB inputs.
UniD achieves competitive performance in dense scene prediction by effectively bridging domain gaps without the need for overlapping annotations or costly pseudo-labeling.
DAPM sets a new benchmark in monocular depth estimation for UAVs, achieving unprecedented accuracy in dynamic aerial environments.
Continuous alignment of text and visual representations in DINOde leads to significant performance gains in open-vocabulary semantic segmentation, outperforming traditional methods.
CRWKV redefines recurrent state propagation in underwater image enhancement, achieving superior visual quality by dynamically adapting to scene-specific degradation patterns.
FedSEPT achieves superior local adaptation without sacrificing global generalization, even under stringent privacy conditions.
ResponseGuard outperforms traditional reasoning-based guards in harmfulness detection while operating at a fraction of the time cost, reshaping our understanding of safety in AI responses.
M$^3$-Gen reveals how multimodal integration can generate interpretable gene expression profiles, bridging the gap between clinical imaging and molecular data.
PC-Edit achieves superior image editing quality by directly capturing semantic differences through prompt-contrastive attention, enabling precise object replacement without sacrificing background fidelity.
Incorporating optical priors into generative models can dramatically enhance the fidelity of super-resolution fluorescence microscopy images.
Shift-IISR achieves superior consistency in infrared super-resolution by leveraging dual-path diffusion techniques that enhance both global distribution and local structure.
Knowledge retrieval and reasoning are the main bottlenecks in KI-VQA, but a new diagnostic benchmark reveals deeper issues in visual grounding and object identification.
C-PTQ reveals that harmonizing task sensitivity with quantization error can dramatically enhance MLLM performance without additional complexity.
Current pathology VLMs can achieve high accuracy without visual input, revealing critical flaws in how we assess their multimodal capabilities.
S3GNet achieves unprecedented accuracy in hyperspectral salient object detection by effectively distinguishing between intrinsic material properties and misleading spectral variations.
EmoAgent-R1 reveals that dynamic agent specialization can dramatically enhance emotion reasoning in multimodal contexts, outperforming traditional fixed-prompt approaches.
Captions can be made 48% more complete and hallucination reduced by 45% without retraining models, thanks to a novel prominence-aware rectification approach.
Radiological findings in frozen vision-language models are encoded by just 10 channels, revealing a surprising efficiency that challenges traditional assumptions about model complexity.
Richer visual representations don't always lead to better human alignment; sometimes, simpler methods like temporally averaged images outperform full video in identifying engaging moments.
HyWorldVLA achieves unprecedented noise robustness in autonomous driving by seamlessly integrating pixel-level grounding with latent representation learning.
Transforming text-to-video retrieval into a distribution alignment task leads to significant performance gains and better uncertainty calibration in ranking.
MAGE-Vein reveals that demographic biases, not the modality itself, were the primary culprits behind past failures in finger vein age estimation.
WhereEdit reveals that spatially aware editing can dramatically improve one-step image editing quality, outperforming traditional methods that rely on global transformations.
A robotic craniotomy system that autonomously adapts to unknown tissue properties, achieving high precision in temperature estimation and surgical execution.
Correctable visual attention can drastically enhance robot manipulation performance, especially in unpredictable environments.
This new evaluation framework can reliably distinguish between meaningful audio caption variations and genuine corruptions, transforming how we assess automated audio captioning systems.
VCSD achieves up to a 4.5% performance boost over traditional methods by distilling knowledge without requiring external teachers or visual cues.
GraphVid achieves superior video quality and controllability with significantly less training data, revolutionizing how we can interact with multi-object video generation.
Explicitly modeling shared world states in multi-agent settings can drastically enhance video generation quality and logical consistency.
Image-generation models can outperform text-output VLMs in spatial tasks when evaluated through visual answers, revealing a critical shift in how we assess spatial intelligence in AI.
Oxygen-TryOn achieves unprecedented realism in virtual try-on by synthesizing images across diverse fashion categories, outperforming both proprietary and open-source models.
Trace achieves a 4.06 percentage point boost in visual reasoning accuracy for large language models, showcasing the power of structured task generation in RLVR.
AAMFM achieves unprecedented accuracy in functional antibody design by effectively pairing antibody sequences with antigen contexts, setting a new benchmark in the field.
SoftReason achieves seamless integration of high-dimensional perception and symbolic reasoning, enabling end-to-end differentiable learning in complex reasoning tasks.
Condition Dropout transforms RGB-D models by enabling them to maintain performance even when one modality is absent, effectively turning a limitation into an advantage.
Achieving high-quality full-length song generation from diverse inputs, this framework outperforms existing methods in musicality and fidelity.
Bridging the lab-to-store gap in humanoid robotics is less about new architectures and more about smart data design and integration strategies.
StreamHOI achieves real-time HOI video generation with impressive interaction fidelity, hitting 17.6 FPS and just 0.75 seconds latency.
Audio-Zero reveals that fine-grained auditory reasoning can be achieved without any external labels, transforming how we approach audio model training.
Achieving video editing speeds 155–171 times faster than traditional methods without sacrificing quality could revolutionize real-time content creation.
Tiny aerial targets can now be accurately identified in streaming videos, thanks to a new dataset and a model that preserves critical visual details while managing memory constraints.
Surface accuracy metrics can mislead researchers about the true reliability of multimodal search systems, with silent failures lurking beneath the surface.
Achieving 86% accuracy in automated symbol generation could revolutionize PCB design by drastically reducing manual errors and time investment.
Vision-language models exhibit a surprising sensitivity to input order, with image-first prompting outperforming question-first by a significant margin, revealing a critical circuit-level failure.
Hypergraph visualization in RAG systems boosts performance by effectively utilizing the visual perception capabilities of multimodal large language models.
FashionAM revolutionizes fashion image retrieval by directly linking multimodal queries to visual embeddings, eliminating the pitfalls of textification.
Jointly optimizing visual token selection and LLM computation can drastically reduce inference costs while enhancing performance in multimodal tasks.
Importance-Aware Sampling (IAS) reveals that not all patches in VIS-IR data are created equal, leading to substantial performance gains in multi-sensor perception tasks.
Captions generated by PercepCap are not only more accurate but also grounded in explicit spatio-temporal perception, revealing the underlying reasoning behind each description.
Multi-view aggregation boosts building health assessment accuracy by over 14%, but neighborhood context only adds a modest improvement, revealing its limited role in inspection.
Achieving 94% accuracy in real-time EEG electrode detection could revolutionize point-of-care diagnostics by streamlining cap placement.
OffNadirLoc reveals that leveraging structure-aware contextual weighting can dramatically enhance UAV-to-satellite geo-localization performance under challenging off-nadir conditions.
ETPDesigner transforms static theater programs into immersive, interactive experiences, enabling real-time conversations with virtual characters.