Search papers, labs, and topics across Lattice.
100 papers published across 2 labs.
Achieving real-time 3D hand pose estimation without the need for camera parameters could revolutionize applications in AR and robotics.
Visual insensitivity in multimodal LLMs can be effectively mitigated with a new framework that enhances alignment and reduces hallucinations.
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
FIRM-Video reveals that a checklist-driven approach can significantly enhance the reliability of text-to-video reward modeling, achieving state-of-the-art performance in evaluation metrics.
Vision-language models struggle with spatial reasoning, achieving only 32% accuracy in grounding tasks, revealing a critical flaw in candidate localization.
FIRM-Video reveals that a checklist-driven approach can significantly enhance the reliability of text-to-video reward modeling, achieving state-of-the-art performance in evaluation metrics.
Vision-language models struggle with spatial reasoning, achieving only 32% accuracy in grounding tasks, revealing a critical flaw in candidate localization.
Turn-taking prediction in Turkish conversations can be significantly improved using a novel multimodal dataset and a hybrid rule-based approach that captures complex interaction cues.
TLive-Omni achieves superior live-commerce understanding by seamlessly integrating multi-modal inputs and optimizing for real-time response quality.
Current omni-modal models can achieve only moderate success in interactive video assistance, revealing critical gaps in their understanding of user interactions and visual cues.
EviRank redefines multimodal image re-ranking by transforming complex queries into structured evidence, achieving unprecedented accuracy and efficiency.
InfinityEdit allows for real-time, unbounded video editing that adapts to live streams, maintaining quality and coherence despite multiple edits.
Compressing VLMs to a mere 3.7 GB without sacrificing performance could revolutionize mobile AI applications.
EXPL-FR reveals that you can achieve interpretable face recognition without any text training, using a simple adapter to bridge vision and language spaces.
Explicit modular spatial verification can dramatically enhance diagnostic accuracy in medical imaging, outperforming traditional end-to-end models by a staggering margin.
A single architecture can achieve state-of-the-art results in pelvic imaging across multiple modalities, addressing critical data integrity issues in women's health research.
ArmorOCR not only enhances adversarial OCR perception but also preserves competitive performance on standard OCR tasks, bridging a critical gap in model robustness.
Achieving real-time 3D hand pose estimation without the need for camera parameters could revolutionize applications in AR and robotics.
Rethinking the VLM training pipeline allows for the creation of specialized models that outperform larger counterparts while using far fewer resources.
Existing methods falter on a new benchmark that captures the full complexity of basketball game events, revealing critical gaps in current visual understanding approaches.
Mix&Fix-Net bridges the monitoring gap for small vessels by integrating AIS and vision data, achieving superior trajectory prediction accuracy.
Achieving over 96% success in real-world door traversal tasks, this framework redefines how robots can learn from a single video input.
Structured affinity enables Deep AINs to learn visual memory without replay, achieving impressive accuracy while retaining earlier class information.
Feature migration in Vision Transformers is predominantly an early-stage phenomenon, with deeper layers stabilizing faster than shallow ones.
Training only the projector can match or exceed the performance of fully fine-tuned multimodal models while avoiding capability drift.
Achieving unprecedented accuracy in 3D object detection by refining both object discovery and model training through innovative dual-guidance techniques.
Robots using EAFG can identify crucial objects before planning, leading to a dramatic increase in successful task completion rates.
Models trained on incomplete multimodal data can significantly improve generalization to unseen combinations, achieving over 5% higher accuracy than existing approaches.
Semantic structured partitioning in SCPaT leads to enhanced forecasting accuracy by intelligently modeling interactions among heterogeneous temporal patterns.
TESTNAV achieves up to 2.15x faster exploration of compositional perturbation spaces while ensuring realistic and severe input failures are prioritized.
Subtle timing in subtitle presentation can amplify video jailbreak effectiveness against LVLMs, achieving unprecedented attack success rates.
Core-KAN achieves superior performance in vision tasks by synthesizing spatial filters at arbitrary resolutions without the computational burden of per-location kernel generation.
Thoughtful resistance in AI content creation can significantly elevate educational quality, proving that educators and AI can be powerful allies rather than adversaries.
Iterative proxy correction boosts sentiment analysis accuracy by refining initial proxies and adapting to incomplete multimodal inputs.
Visual insensitivity in multimodal LLMs can be effectively mitigated with a new framework that enhances alignment and reduces hallucinations.
Integrating LLMs into autonomous driving systems can enhance decision-making without sacrificing control or safety, even in unpredictable environments.
LLMs systematically favor text over numbers in evidence arbitration, revealing a critical failure mode in decision-making systems that rely on heterogeneous data sources.
Achieving 100% requirement coverage and Gherkin generation accuracy, this testing pipeline could redefine how automotive software is validated across decentralized systems.
Turning circle-based control barrier functions enable autonomous surface vehicles to navigate complex environments without predefined paths, achieving superior safety and efficiency.
Jointly forecasting surgical instrument trajectories and visual states could revolutionize surgical motion planning by providing a more coherent understanding of action-scene dynamics.
EXIMO achieves unprecedented sample efficiency in robotic policy finetuning by leveraging a vision language model to decompose complex tasks.
Achieving a 0.499 face similarity score, WithEveryone revolutionizes group image generation by ensuring identity preservation for up to ten individuals without direct face copying.
Open-vocabulary word-level information can be reliably decoded from EEG during silent reading, revealing a scalable approach to understanding inner speech.
Streetscape qualities crucial for pedestrian-friendly urban design are found to be alarmingly scarce in suburban areas, revealing a significant gap in urban planning.
G-MARK reveals that grounding multi-agent reasoning in provenance-aware knowledge graphs can drastically enhance occlusion reasoning and decision-making in cooperative driving.
Iterative perception can boost document VQA accuracy by over 25%, proving that smarter evidence acquisition trumps mere model size.
Trusting individual predictions from VLMs can be achieved without fine-tuning, revealing critical insights into model failures that standard self-consistency checks overlook.
Optimizing latent visual representations can boost multimodal reasoning performance by over 9% on complex tasks.
G-CARL not only improves the accuracy of medical report interpretations but also ensures they are tailored to patient queries, outperforming traditional methods in both factuality and user satisfaction.
Conversational surveys combined with multimodal LLMs can enhance travel behavior predictions, achieving over 71% accuracy by leveraging visual context.
RuleMaze reveals that separating perception, execution, and rule verification can dramatically enhance MLLMs' ability to follow complex natural-language instructions in spatial planning tasks.
DECOWAM achieves a 21.7% reduction in action prediction error while maintaining robust task performance, showcasing the power of embodiment-aware factorization in mobile manipulation.
Integrating prompt-conditioned channel attention can boost segmentation accuracy by over 23% in challenging medical imaging tasks.
Real-world tennis serving by humanoid robots is now possible without motion capture, thanks to a novel adaptive framework that learns directly from video.
DARS achieves superior performance in instruction-based image editing by transforming outcome-level feedback into actionable, localized supervision for both planning and rendering stages.
SEFS achieves superior artistic stylization by leveraging low-resolution image crops, enhancing content consistency while avoiding unwanted style transfer artifacts.
Under extreme visual degradation, adaptive fusion of sonar and visual data boosts underwater detection accuracy by over 33%, revealing the critical role of modality reliability.
Sarcastic-aware contrastive regularization enables the model to discern nuanced sarcasm, outperforming traditional methods that struggle with modality inconsistencies.
Jointly leveraging selection-based and reinforcement-learning-based alignment can transform medical image captioning into a more trustworthy diagnostic tool.
Enhancing AI clones with listening behaviors can dramatically elevate user perceptions of authenticity and engagement.
UniLang enables pretrained LLMs to seamlessly integrate machine-native symbols, outperforming traditional models in diverse structured prediction tasks.
A unified framework that simultaneously enhances interaction understanding and generation, achieving state-of-the-art results in multimodal human-human interaction analysis.
DPC-Net achieves unprecedented image restoration quality by seamlessly integrating semantic understanding with low-level visual cues, outperforming existing methods on multiple benchmarks.
Eyeglasses removal from video can now be achieved with unprecedented fidelity and temporal stability, thanks to a physics-grounded approach that preserves identity and expression.
Achieving leading performance in image generation with only 6 billion parameters, Swift-Image redefines the efficiency frontier for compact models.
DreamHand achieves a groundbreaking 40% reduction in error for 3D hand trajectory recovery in occluded environments, setting a new standard for egocentric video analysis.
G3Ego reveals that integrating gaze into graph construction can significantly enhance the efficiency and accuracy of egocentric action recognition.
UPAL achieves a remarkable 4x speedup and 10x smaller memory footprint while maintaining state-of-the-art performance in multi-view feature extraction.
Micro-drones can autonomously navigate hazardous environments without GPS, preserving vital sensor data even when communication is lost.
MUST-PET achieves superior lesion segmentation and reconstruction accuracy, even with limited labeled data, by harnessing the power of multimodal self-supervised learning across diverse PET-CT scans.
By combining images with text queries, ID-VTG significantly improves the accuracy of video grounding in scenarios with visually similar entities.
Achieving accurate pose estimation with just one or two feature correspondences could revolutionize localization in consumer devices.
AutoLumNet achieves state-of-the-art exposure correction by uniting monotonicity, optimal transport, and local adaptivity in a single trainable model.
Injecting diffusion representations into CLIP pipelines boosts CZSL performance, revealing the untapped potential of generative models in zero-shot learning tasks.
Transforming static avatars into dynamic, realistic representations could redefine the standards for avatar realism in virtual environments.
Achieving near fully supervised accuracy in tumor segmentation using only image-level labels, even in the presence of client-specific missing modalities, is a game changer for federated learning in healthcare.
Achieving a Dice score of 0.6796, AsymFeX outperforms existing methods by effectively utilizing brain symmetry for accurate stroke lesion segmentation across imaging modalities.
TextRefine achieves superior text editing in product posters by ensuring high fidelity and optimal placement, overcoming common pitfalls of existing models.
Modern autonomous driving systems hinge on learned representations that prioritize safety and compliance, not just raw performance metrics.
OrthoSkillVLA preserves prior skills in pretrained VLA models while seamlessly integrating new ones, outperforming traditional methods in both simulated and real-world scenarios.
CVSD-Reg achieves a remarkable 97.7% success rate on challenging LiDAR datasets, outperforming existing methods by up to 44% without relying on camera data.
LFPR boosts bounding box accuracy by over 3% on average without any target annotations, revealing a complex interplay between referent selection and boundary precision.
Hierarchical tactile modeling boosts robot manipulation success rates by over 40% in contact-rich tasks, revealing the power of structured tactile forecasting.
Participants transformed their dance experiences by leveraging a sound-based device that creates immersive soundscapes from movement, enhancing spatial awareness and improvisation.
A 14M parameter model that outperforms larger transformers while being 12 times smaller, reshaping the efficiency landscape of audio-visual processing.
NAPE achieves state-of-the-art performance in audio representation learning by simplifying the pre-training process to a single autoregressive prediction task.
Achieving high-quality 4D human reconstruction from casual monocular videos could revolutionize applications in virtual reality and gaming.
Transferable presentation attack representations can be learned without relying on facial content, achieving high performance across standard benchmarks.
Achieving a $5.15\times$ speedup in text-to-3D generation without sacrificing quality could revolutionize real-time 3D content creation.
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
Annotations can be transformed into powerful oracle rollouts, dramatically enhancing the efficiency of reinforcement learning for video MLLMs.
SCORE achieves a remarkable 53.23% Top-1 accuracy in EEG-to-image retrieval, outperforming existing methods by over 17 percentage points, even without target labels.
Anchoring neural and visual representations can boost brain-to-image retrieval accuracy by over 9% with minimal stimulus repetitions, challenging the reliance on trial averaging.
Denoising in diffusion models acts as a dynamical Bayesian classifier, revealing that posterior probabilities can focus on a single cluster under specific conditions.
Small, human-imperceptible image perturbations can drastically mislead Vision Language Models, revealing a significant security flaw in multimodal AI systems.
Achieving competitive skin disease classification while ensuring fairness across skin tones, MIFR aligns clinical and dermoscopic data in a shared embedding space.
Treating missing data as a contextual signal, MARCUS dramatically improves rent prediction accuracy, slashing MAE by over 51% in some cases.
Adaptive similarity margins in HN-CLIP boost retrieval accuracy by up to 4.3% while training 2.4x faster than leading methods.
Hallucinations in LVLMs can be reduced by over 21% without sacrificing performance, thanks to a novel method of calibrating visual evidence during decoding.
Counterfactual images generated without reliance on specific classifiers can significantly reduce bias and improve interpretability in medical imaging tasks.
Pair-Aware Discriminative Reasoning in UMER reveals critical distinctions between semantically similar candidates, elevating retrieval accuracy in multimodal tasks.
DentAgent outperforms senior specialists by 17.3 percentage points in multi-label diagnosis, revolutionizing multimodal dental reasoning with traceable evidence integration.
DRB can optimize reasoning budgets to improve LLM performance while cutting costs, achieving better results than traditional maximum-budget approaches.
Current OCR systems struggle with complex handwritten inputs, with performance plummeting on multi-line formulas and generative models often hallucinating errors.
MR-IQA-2 reveals that decoupling reasoning from rating can significantly enhance the reliability of image quality assessments, achieving human-level alignment without sacrificing interpretability.