Search papers, labs, and topics across Lattice.
100 papers published across 1 lab.
FLARE transforms VLAs from brittle performers into resilient agents capable of autonomously recovering from common execution failures in robotic manipulation.
MedFG-VQA achieves superior performance in medical VQA tasks while being lightweight enough for practical clinical use, thanks to innovative memory and graph attention techniques.
EditaLive achieves real-time character video editing with state-of-the-art performance, preserving facial expressions while eliminating latency issues in live streaming.
MLLMs can recognize urban scenes but fail to maintain reliable navigation and goal-directed behavior over extended exploration in complex environments.
Editing long videos with multiple instructions can be done without hallucinations or loss of temporal continuity, thanks to a novel agentic framework that combines LLMs and VLMs.
EditaLive achieves real-time character video editing with state-of-the-art performance, preserving facial expressions while eliminating latency issues in live streaming.
MLLMs can recognize urban scenes but fail to maintain reliable navigation and goal-directed behavior over extended exploration in complex environments.
Editing long videos with multiple instructions can be done without hallucinations or loss of temporal continuity, thanks to a novel agentic framework that combines LLMs and VLMs.
Task-specific image editing can boost MLLM performance by over 29%, but not all tasks benefit equally—some are left behind.
QuantumBoostNet outperforms traditional models in cardiac ultrasound view identification, showcasing the potential of hybrid classical-quantum approaches in medical imaging.
HALO achieves a 13.7 percentage point boost in zero-shot open-set accuracy with 10x fewer parameters than leading models, redefining efficiency in human activity recognition.
MM-Spectrum achieves substantial performance gains in molecular structure elucidation by intelligently routing multimodal data, revealing the power of tailored expert mechanisms in handling spectral heterogeneity.
SAGE achieves state-of-the-art forecasting accuracy by integrating multimodal semantic knowledge without the computational burden of large language models.
Transcript-based shortcuts in dialogue models lead to a staggering drop in accuracy, revealing a critical flaw in current evaluation methods.
PDPO revolutionizes robot crowd navigation by generating action chunks that enhance safety and efficiency in dense human environments.
PETs that achieve similar classification accuracy can perform drastically differently across various vision tasks, revealing hidden vulnerabilities in their effectiveness.
Transforming ECG signals into GADF images reveals critical inter-lead dependencies, leading to superior classification performance in coronary artery disease detection.
AI's aesthetic categorization of art diverges significantly from human emotional responses, revealing a unique interpretative framework that could transform media organization.
Strong-modality collapse can degrade dominant modalities by over 18%, but Inverted Asymmetric Fusion preserves performance while enhancing weaker modalities.
Transforming static charts into editable SVGs, Chart2SVG unlocks new possibilities for interactive data visualization and manipulation.
Achieving a record 38.8% mIoU on the SemanticKITTI hidden test with a single-sweep, single-sample approach, this work redefines the limits of LiDAR scene completion.
CLAP achieves zero-shot deployment of physical simulators across diverse robot embodiments, outperforming traditional models in complex environments.
LeVJEPA achieves up to 20.8x less pretraining compute while surpassing the performance of leading video representation methods, reshaping the landscape of video-based learning.
By shifting capacity from complex actors to deep critics, LAC achieves state-of-the-art performance with dramatically lower inference latency.
SCG enables Vision Transformer encoders to adaptively grow in complexity based on task demands, achieving significant efficiency gains without sacrificing representation quality.
Aggressive 4-bit quantization can cripple MLLM performance, but a novel residual reconstruction technique recovers lost capabilities with minimal overhead.
TransMeme achieves a 33.1% improvement in meme transcreation quality by expertly balancing cultural adaptation and multimodal coherence.
Property-specific discrepancies in PPG-to-rPPG recoverability reveal that not all physiological signals are equally preserved under varying observation conditions.
Automated fluorescence measurements reveal plant stress with unprecedented precision, transforming how we assess agricultural health.
Four new event-based datasets could redefine the landscape of spiking neural network research by providing the high-quality data needed for robust object classification.
OmniUE achieves an unprecedented 83.7% improvement in visual-interactive benchmarks, setting a new standard for multimodal embedding performance.
T2S transforms open-vocabulary semantic segmentation by generating precise seed points from text, leading to superior segmentation without the need for training.
Achieving a 3.1x speedup in inference time while retaining 93.8% of performance, PACE revolutionizes how VLMs handle visual token efficiency.
Explicitly incorporating state awareness into task planning with MM-LLMs leads to a 32.8% increase in action executability, revolutionizing human-robot collaboration.
Pixel-level table compression can dramatically reduce token usage while enhancing accuracy in document question answering, challenging conventional methods.
OCSD reduces path error by over 30% and generates more realistic long-term human motion forecasts by effectively integrating object cues and social interactions.
MedFG-VQA achieves superior performance in medical VQA tasks while being lightweight enough for practical clinical use, thanks to innovative memory and graph attention techniques.
Query-aware evidence forests can drastically reduce memory lifecycle costs while improving accuracy, setting a new standard for multimodal agent memory management.
Real-time queue management can slash waiting times by 30% and boost throughput by nearly 20% in border control systems.
Instruct-to-Act reveals that decoupling planning from control can enhance action execution speed and flexibility in complex environments without sacrificing performance.
MedREAL achieves a remarkable 68.49% gIoU and 70.47% cIoU, setting a new benchmark for interpretable medical image analysis that aligns reasoning with pixel-level accuracy.
Achieving high-fidelity video virtual try-on in real time, LiveVVT reduces latency by 26x while enhancing generation quality.
Visually plausible charts generated by LLMs often mask significant data-level hallucinations, revealing a critical gap in current AI capabilities.
AesCanvas reveals that aesthetic specialization does not guarantee contextual suitability, challenging the assumptions about model performance in image assessment.
Merging existing expert capabilities can yield significant performance boosts, but choosing the right fusion method can make all the difference in multi-domain reinforcement learning.
Generative image retrieval just got a major upgrade—PailitaoGR boosts performance by 13.8% by mastering target focus and auxiliary evidence utilization.
CoGeo-GS achieves superior multi-object removal in 3D scenes by integrating concept-driven tagging with geometry-aware completion, outperforming traditional methods in both quality and stability.
Multimodal models struggle with consistency, showing significant judgment discrepancies between text and speech inputs, especially in Arabic contexts.
Authentic telemedicine conversations in Bengali reveal that existing models can significantly improve their performance on clinical reasoning tasks with the right data.
Cross-lingual word-to-speech mappings can be effectively learned from visual grounding without the need for transcriptions or extensive model training.
Prioritizing semantic anchors over easy tokens can drastically reduce error accumulation in multimodal language models.
VIG-Sampler boosts multimodal generation quality by leveraging visual attention, outperforming traditional methods with fewer decoding steps.
Audience preferences can mislead creative strategies, with pooled estimates sometimes suggesting the opposite of what drives engagement.
Reducing model parameters by over 70% while boosting tumor detection rates reveals a critical trade-off in medical imaging segmentation tasks.
G2D boosts zero-shot image classification accuracy by up to 27.42 percentage points by effectively combining generative verification with discriminative retrieval.
GeoMAD achieves superior anomaly detection by seamlessly integrating geometric correspondence with distributional consistency, all while maintaining efficiency in 2D feature-space learning.
Masking just 20 Visual Retrieval Heads in VLMs can lead to an 80-point drop in grounding accuracy, revealing their critical role in visual information extraction.
MILO redefines 3D human-object interaction reconstruction by leveraging Large Reconstruction Models, achieving unprecedented accuracy from just a single image.
NeRF-based methods outperform traditional techniques in generating high-fidelity holograms of complex laboratory objects, transforming educational visualization.
Integrating depth information with visual data leads to a dramatic boost in 3D awareness, outperforming traditional methods on key benchmarks.
Sidecar boosts character consistency in visual storytelling by seamlessly infusing missing identity semantics into prompts, all without the need for retraining.
CODE outperforms prior methods in Open World Object Detection by effectively balancing known and unknown object detection through innovative calibration and suppression techniques.
VLMs exhibit a significant performance gap in visual text understanding, with even the best models falling short of human accuracy in error correction tasks.
TetherMem enables video generation models to dynamically adapt scenes while keeping subjects consistent, achieving a significant leap in overall video quality and scene progression.
Ancient-Bench reveals that even advanced models fail to solve the challenges of recognizing ancient Chinese texts, underscoring a critical gap in AI capabilities.
Vision generative AI models could revolutionize edge applications, but only if we rethink how they are designed in relation to hardware constraints from the outset.
AnatoProto not only surpasses state-of-the-art models in fetal ultrasound detection but also reveals that combining anatomy-weighted pooling with prototype loss can dramatically enhance recall.
Calibration in medical vision-language models is crucial, and MVC-Bench reveals that a simple train-time calibration method can outperform existing approaches in most scenarios.
Video-OPSD reveals that focusing on privileged visual evidence can drastically improve the efficiency and effectiveness of self-distillation in Video-LLMs.
AVTP achieves a 2x inference speedup in LVLMs while maintaining up to 96.1% accuracy, revolutionizing multi-image processing efficiency.
Gait recognition can now leverage natural language queries, achieving unprecedented accuracy while preserving identity integrity.
Emotion recognition in visual intelligence is heavily skewed towards linguistic cues, revealing a critical gap in purely visual affect recognition capabilities.
Real-time emotion predictions can be made more reliable by selectively invoking multimodal reasoning only when necessary, improving accuracy without sacrificing speed.
RubricRM adapts evaluation criteria dynamically, leading to significant performance gains in visual generative tasks compared to static reward models.
UniGeo achieves a remarkable 13.59-point improvement in retrieval accuracy for text-guided drone geo-localization by leveraging a unified multimodal framework.
FAN-LoRA outperforms existing methods by explicitly decoupling frequency components, leading to significant enhancements in medical image segmentation accuracy.
By integrating uncertainty modeling into SDF learning, NeuDonatello significantly boosts the accuracy of 3D surface reconstruction from RGB images, even in challenging scenarios.
LVLMs exhibit a dramatic decline in performance when faced with shuffled meme panels, underscoring a critical gap in order-sensitive reasoning capabilities.
Attackers can now induce specific failure behaviors in VLA models with unprecedented precision, revealing a new dimension of vulnerability in AI systems.
SRM-FND achieves superior fake news detection by enhancing reasoning quality through self-reflection, outperforming traditional methods in reliability and interpretability.
Benchmark-driven results can obscure the true complexities of real-world imaging problems, leading to misguided research priorities.
Preserving cross-modal alignment flow can dramatically enhance the performance of multimodal models, countering the effects of catastrophic forgetting during fine-tuning.
Achieving minutes-long coherence in video generation, Ring Forcing reconciles the trade-off between historical fidelity and generative diversity.
Grounding glass surface detection in 3D geometry leads to state-of-the-art performance and improved scene reconstruction, challenging the limitations of traditional 2D methods.
Prioritizing content novelty over recency, RECAP-Forcing revolutionizes memory management in long video generation, leading to enhanced visual coherence.
MASON boosts VLM accuracy on compositional layout understanding by over 12% while using only 30% of the training data compared to standard methods.
A sub-million-parameter robot manipulation policy outperforms larger models by leveraging predictive coding for real-time state correction.
FLARE transforms VLAs from brittle performers into resilient agents capable of autonomously recovering from common execution failures in robotic manipulation.
DPA-I2P achieves a remarkable 45% reduction in pose estimation errors, setting a new standard for image-to-point cloud registration in autonomous driving.
Real-time object navigation can see up to an 11% boost in success rates by addressing inference latency and asynchronous stepping in model design.
Video-FLAIR boosts accuracy by over 5 points while slashing token usage from 417 to just 95 by dynamically selecting reasoning strategies for multimodal queries.
A striking task-dependent robustness gap reveals that while ASR thrives on direct audio retrieval, AQA falters due to bottlenecks in mediated information access.
FlashVLA achieves over 30 Hz control frequency with smooth asynchronous execution, revolutionizing real-time robotic manipulation.
Current methods falter in efficiently rearranging scenes with occlusions, exposing a critical gap in embodied agent capabilities.
GRAFT boosts robot manipulation success rates by 25 percentage points while slashing the computational costs of online learning.
Achieving over 97% success in long-horizon robot manipulation, TemporalFlow-VLA reveals the critical role of execution history in action prediction.
A novel soft gripper design maintains stable grasping forces over time, achieving unprecedented accuracy by compensating for stress relaxation in real-time.
CLIPPER slashes planning rollout times by up to 28.9 times while keeping coverage nearly identical to traditional methods, revolutionizing urban micromobility strategies.
PrismRec revolutionizes micro-video recommendation by treating video content as a core driver of user preferences, achieving a remarkable 22.65% improvement over existing methods.
ProRetrieval outperforms state-of-the-art models by synthesizing executable retrieval programs that seamlessly combine structured and unstructured data queries.
RetrievalRouter achieves a 2.5% accuracy boost while being 12.4 times faster than traditional static retrieval pipelines, revolutionizing how we approach document retrieval.
Current MLLMs falter in following complex video instructions, revealing a significant oversight in their evaluation metrics.
Rubric-based scoring reveals that vision-language models can significantly improve their grounding in visual evidence, enhancing reasoning and instruction adherence.
PPE achieves a staggering 83.3% recall in predicting Ebola outbreak zones, outperforming traditional models by over 10 percentage points.
Humans can pinpoint systematic flaws in AI-generated images, revealing critical insights into the limitations of current text-to-image models.