Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
Achieving real-time, high-accuracy point correspondences without 3D scene knowledge could revolutionize image processing in dynamic environments.
FlashDrive slashes VLA inference latency by 4.7x while maintaining accuracy, making real-time autonomous driving feasible with a single GPU.
MOSS-VL achieves a staggering 66.0 average score in proactive alerting, far surpassing the best baseline by nearly 30 points.
PRISM achieves a remarkable 12.94 percentage point improvement in noisy audio classification without any additional training, redefining adaptation strategies for ATMs.
Current leading models fail to achieve over 60% success in constructing 3D worlds from user prompts, but VibeWorlder models leverage RL training to outperform them significantly.
MOSS-VL achieves a staggering 66.0 average score in proactive alerting, far surpassing the best baseline by nearly 30 points.
PRISM achieves a remarkable 12.94 percentage point improvement in noisy audio classification without any additional training, redefining adaptation strategies for ATMs.
Current leading models fail to achieve over 60% success in constructing 3D worlds from user prompts, but VibeWorlder models leverage RL training to outperform them significantly.
CPI-Bench reveals significant performance gaps among image editing models, offering a more nuanced evaluation that aligns with real-world user experiences.
SPARGen achieves competitive performance across diverse spatial tasks by unifying perception and reasoning into a single generative framework, challenging the need for separate architectures.
Unlocking the potential of multimodal AI to interpret complex scientific images could revolutionize how we access and understand experimental data.
Subtracting information from the student rather than adding it to the teacher can yield performance improvements that rival those achieved with privileged data.
A hybrid CNN-bi-LSTM architecture decodes motor imagery EEG signals with unprecedented accuracy, even in the presence of significant noise.
MAG leverages abundant unlabeled multi-modal data to boost few-shot learning performance, achieving significant gains in label-scarce environments.
GeoCache achieves over 2x speedup in multi-view texture diffusion without sacrificing fidelity by exploiting cross-view geometric redundancy.
Certified caching can boost edge image classification speed by 1.65x without compromising reliability.
UniTraffic-Agent achieves top-tier performance in complex traffic video reasoning tasks, showcasing the potential of structured workflows in multimodal AI applications.
FlashDrive slashes VLA inference latency by 4.7x while maintaining accuracy, making real-time autonomous driving feasible with a single GPU.
CASA not only sets a new benchmark in automatic speaking assessment but also clarifies how acoustic and content features interact, paving the way for more interpretable AI assessments.
Achieving superior 3D scene reconstruction from a single measurement, this method combines Gaussian Splatting with advanced vision model priors to tackle the complexities of Snapshot Compressive Imaging.
Achieving a 12.1% boost in robust concept removal, MapRoute++ redefines the boundaries of visual concept unlearning.
Achieving state-of-the-art hand gesture recognition with a unified model that combines the strengths of CNNs and transformers while using fewer resources is a game changer for real-time applications.
SketchSense achieves remarkable gains in image inpainting by intelligently interpreting imperfect sketch guidance, outperforming traditional methods in both quality and structural accuracy.
RbFT-Net reveals that correcting radar measurements before fusion can significantly enhance depth completion accuracy in autonomous systems.
A simple Vanilla SFT model outperforms complex reasoning methods in social audio-visual question answering, revealing the inefficiencies of current approaches.
Contrast-based measures can significantly enhance the objectivity of evaluating historical manuscript restorations, outperforming traditional image quality metrics.
QuISE effectively neutralizes typographic attacks on VLMs without any model-specific adjustments, achieving impressive accuracy recovery rates.
Training on the PassGen framework allows robots to anticipate human intentions earlier and more accurately, transforming human-robot collaboration.
Identifying intra-image predictive subsets can significantly boost visual classification performance, especially in challenging data shift scenarios.
Non-thinking inference in hybrid-thinking MLLMs suffers from a staggering increase in response-pattern failures, revealing a critical misalignment that can undermine user trust.
Achieving over 90% mAP in text-based person anomaly retrieval reveals the power of heterogeneous vision-language ensembles and selective multimodal reasoning.
HounsWorld achieves state-of-the-art performance in multimodal patient-state inference by seamlessly integrating CT imaging and clinical language, transforming how we interpret medical data.
The SPARED framework not only boosts detection accuracy but also enhances the quality of reasoning behind verdicts, making AI-generated image detection more robust and explainable.
TennisVAR redefines sports video analysis by grounding tactical reasoning in stroke-level evidence, enabling deeper insights into match strategies.
Transforming feature relationships through geometric guidance leads to consistent performance gains in both image classification and reasoning tasks across multiple architectures.
VOS-Agent outperforms existing methods by leveraging specialized agents for different target types, achieving state-of-the-art results in video object segmentation.
Transforming unresolved failures into a powerful learning signal, FIRE-VLA reduces mean L2 error in autonomous driving models by nearly 19% while maintaining policy efficiency.
HumanoidVLN reveals that navigation success in humanoid robots is significantly influenced by physical embodiment, achieving a 43.55% success rate with state-of-the-art models.
P2Fusion achieves state-of-the-art fusion quality by adaptively regulating modal competition through learnable dynamic regulators, outperforming existing methods in 14 out of 20 evaluation metrics.
GATO-Vid achieves superior spatial localization in text-to-video generation without the computational costs of traditional gradient-based methods.
Moving beyond anchor boxes, this method leverages signed distance functions to achieve superior instance segmentation performance, particularly for irregularly shaped objects.
Current MLLMs achieve only a 75% compilation success rate in scientific figure editing, but targeted training can boost performance to over 83%.
Generated retinal images can inherit critical clinical information, but they may not perform well against real-world classifiers, exposing a crucial representation gap.
DSCC is the first method to effectively ground long-form captions in multimodal models by integrating visual anchors during training, achieving unprecedented precision and length in caption generation.
TabSOM not only outperforms existing tabular-to-image methods but also enhances interpretability by revealing feature relationships and interactions.
TraVEL boosts motion-centric video retrieval performance by up to 9.8 points in mAP, proving that trajectory-guided learning can outperform traditional methods without complex rule systems.
ARMDIL outperforms traditional routers by dynamically selecting the best-suited model for each image, significantly enhancing cross-domain generalization.
UltraIR outperforms traditional machine-learning approaches in IR spectroscopy, achieving high accuracy with minimal labeled data across diverse chemical analysis tasks.
High reasoning accuracy in VLMs doesn't equate to reliability, as shown by GPT-5.2's 96% hallucination rate despite top performance metrics.
VLMs can recognize when to abstain from making a decision, yet they fail to express this restraint, achieving a mere 0.292 on a new metric designed to measure this capability.
CaC2-treated fruits exhibit distinct spectral signatures that can be detected with 95% accuracy, revealing a critical tool for food safety in the fruit industry.
Network features dominate in predicting 6G beamforming performance, outpacing other feature groups in critical metrics.
A hybrid cloud-edge system achieves near-perfect diagnostic recall while slashing operational costs and latency in rural healthcare settings.
A single adversarial texture can compromise the performance of Vision-Language-Action models across multiple tasks, revealing alarming shared vulnerabilities.
Achieving over 30 PSNR in sign language video synthesis with a GAN framework that balances stability and detail could revolutionize communication for the hearing impaired.
LongEarth-R1 outperforms all existing models on long-sequence Earth observation tasks, revealing the critical importance of structured reasoning in complex spatial analyses.
IMU-based sensing outperforms egocentric vision in detecting freezing of gait, but the latter reveals crucial contextual insights that could transform clinical assessments.
Coverage-driven token pruning can significantly enhance the efficiency of 3D VLMs without sacrificing reasoning capabilities.
Current MLLMs struggle with long-term memory, achieving only 71.8% accuracy on a benchmark where humans score 94.2%.
EEG-PRIME achieves zero-shot transfer capabilities in EEG decoding, outperforming traditional models without the need for domain-specific calibration.
Dual-role Identifiers in DrIG not only enhance retrieval accuracy but also tackle local optima challenges, revolutionizing multimodal information retrieval.
Vision-language models struggle to leverage visual evidence in medical VQA, with only one model surpassing human performance on a subset of questions.
VLSR's localize-then-reason approach boosts throughput by 9.6X, revolutionizing how we analyze molecular properties from images.
The model Gemma 3 4B IT reveals a striking distinction in its representation of falsehoods and impossibilities, challenging our understanding of how AI interprets language.
StreamTTT outperforms existing streaming VLMs by enhancing real-time perception and long-term memory without sacrificing performance.
Static task vectors can effectively capture demonstration-induced changes, but their success hinges on the shared structure across queries—complexity is only needed when queries diverge significantly.
Detecting partially forged videos is now feasible with a novel framework that leverages static images for enhanced supervision and accuracy.
FineX boosts fine-grained action recognition accuracy by over 7% on challenging datasets, leveraging a unique fusion of visual and skeletal representations.
HPSD enables TI2V models to internalize high-quality visual cues, resulting in a remarkable boost in text-to-video performance while simultaneously enhancing image-to-video generation.
DINOv2-style pretraining outperforms other SSL methods in resource-limited settings, but combining it with video objectives reveals critical trade-offs in performance.
ProPose not only bridges the gap in pose estimation for diverse limb types but also enhances accuracy for long-tail prosthetic joints through innovative structure-aware loss functions.
Achieving state-of-the-art performance in 3D perception tasks, GeoUP reveals that integrating geometry-grounded representations can significantly enhance autonomous driving systems.
Achieving real-time, high-accuracy point correspondences without 3D scene knowledge could revolutionize image processing in dynamic environments.
Princigram achieves a new standard in scientific diagram generation, ensuring physical accuracy through a structured reasoning framework that traditional models lack.
A novel framework that harmonizes semantic reasoning and predictive dynamics, achieving unprecedented performance in autonomous driving tasks.
A unified framework that effectively fuses RGB and event data can drastically improve person re-identification accuracy across different camera views.
Directly manipulating internal representations allows for effective concept erasure in MM-DiTs without the burden of model tuning.
Task progress in vision-language-action models can be read directly from their internal representations, even before task-specific training, revealing a surprising depth of interpretability.
Spatially grounded tokens in LocusGS lead to coherent Gaussian distributions that dramatically improve rendering quality in 3D scene synthesis.
SCOPE achieves a remarkable 1.99× speedup in video attention while enhancing fidelity, challenging the effectiveness of traditional sparse attention methods.
RippleNet reveals that focusing on low-SNR forgery traces can significantly enhance the detection of AI-generated images, outperforming traditional methods.
A single generalist model can achieve state-of-the-art performance in Referring Expression Comprehension while generalizing across diverse datasets without fine-tuning.
GDI transforms defect classification by generating single-defect samples, leading to a remarkable 63.6% boost in F1-Score for rare defects.
Mr3D-VL outperforms existing models in cross-modal reasoning for brain tumor imaging, achieving a BERTScore of 0.856 and setting a new standard for interpretability in mpMRI applications.
Achieving over 98% accuracy in classifying infected mosquitoes from video frames showcases the power of combining vision and language models for biological analysis.
Temporal GRPO reveals that aligning reinforcement learning updates with task stages can significantly boost both efficiency and success rates in complex vision-language-action tasks.
Action-derived visual attention can boost robot task success rates by over 28% without relying on external labels.
AirForesight achieves superior UAV navigation by seamlessly integrating current spatial knowledge with future trajectory predictions, outperforming traditional methods that lack explicit scene grounding.
Achieving a 12.2% boost in navigation success rates, SAP-Nav redefines how agents can dynamically interact with and understand their environments without prior training.
VoxAudio revolutionizes vocalized audio synthesis by embedding intelligible speech seamlessly within complex soundscapes, outperforming traditional methods that compromise on clarity and control.
Achieving a 5.28% improvement in segmentation accuracy while running 4.7% faster than existing methods, DiCoR redefines efficiency in referring remote sensing image segmentation.
EgoPHI transforms egocentric vision by enabling precise 3D force estimation from a single image, bridging the gap between contact localization and physical interaction reasoning.
Multi-view MRI inputs can enhance spatial localization but may compromise temporal reasoning, revealing critical limitations in current foundation models for clinical use.
MLLMs can leak sensitive personal information when visual evidence is lacking, but the Dynamic Relational Unlearning Framework (DRUF) significantly mitigates this risk without sacrificing performance.
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
SMA achieves superior spatial reasoning in frozen VLMs by converting verified experiences into transferable lessons, outperforming traditional methods without the need for parameter updates.
Latent SDS can produce noisy artifacts, but PixSDS offers a novel solution that preserves semantic content while reducing these distortions.
Realistic predictions in robotic manipulation now come with precise arm control, ensuring the right actions yield the expected outcomes.
AutoDesign outperforms existing design systems by aligning with human design principles and achieving superior poster generation quality through recursive self-improvement.
Achieving state-of-the-art performance in document parsing, NaviDC-OCR tackles geometric distortions and structural reasoning challenges that plague existing models.
A single foundation model can effectively leverage heterogeneous cardiac signals to achieve superior performance across multiple tasks, defying the limitations of modality-specific training.
V-RAE achieves a remarkable 2.13 rFVD on K600, outperforming traditional video VAEs by retaining significantly more semantic information in its latent representations.
Evaluations reveal that current models struggle with long-range narrative integration and cultural reasoning, highlighting a critical gap in video understanding capabilities.
Lifecycle-wide perception is crucial for evidence-grounded scientific discovery, allowing OmniScientist to outperform traditional AI systems in research quality and depth.
Achieving over 96% accuracy in exercise quality assessment could revolutionize remote rehabilitation by enabling effective, therapist-free patient monitoring.