Search papers, labs, and topics across Lattice.
100 papers published across 9 labs.
Artists can now seamlessly texture 3D models with AI-generated images while preserving intricate details and UV layouts, revolutionizing the asset creation process.
Current multimodal large language models struggle with OCT image understanding, falling short even with specialized adaptations.
JoyNexus slashes GPU time and boosts service efficiency by enabling concurrent multi-tenant training with isolated workloads and shared resources.
AV-Flamingo outperforms existing models on complex audio-visual tasks, revealing that size isn't everything when it comes to reasoning capabilities.
S1-Omni consolidates fragmented AI capabilities into a single model, outperforming leading benchmarks and domain-specific models in scientific reasoning tasks.
Artists can now seamlessly texture 3D models with AI-generated images while preserving intricate details and UV layouts, revolutionizing the asset creation process.
Current multimodal large language models struggle with OCT image understanding, falling short even with specialized adaptations.
JoyNexus slashes GPU time and boosts service efficiency by enabling concurrent multi-tenant training with isolated workloads and shared resources.
AV-Flamingo outperforms existing models on complex audio-visual tasks, revealing that size isn't everything when it comes to reasoning capabilities.
S1-Omni consolidates fragmented AI capabilities into a single model, outperforming leading benchmarks and domain-specific models in scientific reasoning tasks.
Current MLLMs struggle with active visual observation, scoring as low as 3.5% on tasks designed to test this critical cognitive function.
Current models falter in executing cross-modal editing instructions, revealing significant gaps in audio-visual consistency and fidelity.
SUFLECA achieves a remarkable 33.4% category accuracy and 42.3% instance accuracy in CAD-to-image alignment, outpacing traditional methods while cutting down computational costs.
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
Video can be reimagined as a dynamic interplay of stable contexts and evolving events, revolutionizing real-time interaction capabilities in AI.
AlphaWiSE achieves superior continual adaptation in multimodal models without sacrificing cross-modal alignment, outperforming traditional methods.
A lightweight model achieves 0.110 higher MCI-class F1 than a heavy baseline while providing intrinsic spatial explainability through attention maps.
Multi-axis max@K boosts image diversity and fairness in text-to-image generation, enhancing representation without sacrificing quality.
Achieving competitive performance with multimodal models trained solely on test environment data challenges the necessity of large-scale internet datasets for effective learning.
Bypassing reasoning tokens can cut inference costs by over 60% while maintaining precision in multimodal document QA.
Equivariance-based self-supervised learning can unlock semantically rich representations in symbolic music that outperform traditional methods by a wide margin.
Habitat-fit demotion and multi-scale aggregation emerged as critical factors in achieving competitive performance for multi-species plant identification from complex imagery.
Existing multimodal systems falter in repository-level localization, with the best performance still falling short of reliable accuracy thresholds.
VideoSEMA outperforms heavier models while maintaining efficiency, achieving top accuracy even as image resolution scales up.
Achieving state-of-the-art perceptual quality in deblurring, JADE-GS outperforms traditional methods by effectively merging physics-based and learned priors.
Summaries generated from dialogues can now capture emotional dynamics alongside semantic content, enhancing our understanding of conversational nuances.
Future feature foresight and sparse point tracking together can transform how VLA models navigate complex environments, leading to unprecedented performance in visuomotor tasks.
Text-centered multimodal fusion boosts ambivalence recognition accuracy by over 4% compared to traditional text-only models, challenging the need for complex ensembles.
Multimodal fusion significantly enhances emotion recognition accuracy in vehicles, revealing critical insights into driver state monitoring.
VLMs struggle with strategic decision-making in soccer, showing a preference for safer plays over optimal actions, which could reshape our understanding of their reasoning capabilities.
Achieving 98.9% precision in gesture detection, this framework redefines how robots can interact in resource-constrained environments.
TopoAgent's innovative use of dynamic graph evolution allows for noise-resistant scientific reasoning that outperforms traditional linear models.
Relationships with conversational AIs evolve through both gradual accumulation of familiarity and sudden relational turning points that can be anticipated through user behavior.
A modern multimodal assistant can achieve impressive performance on legacy hardware, revealing surprising insights about weight precision and long-context processing.
Answer-level improvements in multimodal VQA don't guarantee trustworthy clinical reasoning, revealing critical gaps in current evaluation practices.
Autistic individuals exhibit a distinct gaze pattern in art viewing, diverging from both artists and neurotypicals in attention dynamics and consistency.
Achieving competitive saliency detection accuracy with weak supervision could redefine the landscape of annotation efficiency in computer vision.
Real-world datasets reveal that tracking tomato ripeness can significantly enhance agricultural automation and phenotyping techniques.
Roadwork zones can confuse VLMs, but WorkDrive's causal reasoning framework cuts trajectory prediction errors by up to 12% in these challenging environments.
Visualizing navigation structures transformed practitioners' approach to accessibility, shifting their mindset from compliance to design innovation.
Achieving a mean error of just 4.89° / 1.63 m, this method redefines the efficiency and accuracy of integrating LiDAR and camera data for autonomous systems.
TVB achieves unprecedented accuracy in BEV segmentation by leveraging variational inference and attention mechanisms to fuse multi-camera data effectively.
Experience reuse in navigation can significantly boost performance, with VTM-Nav outperforming traditional methods by leveraging a hierarchical memory structure.
CosFly-VLA slashes tracking errors by over 34% in complex urban environments, even when targets are occluded.
Future tactile states can be predicted more effectively from intermediate action features, transforming how we approach tactile supervision in robotic manipulation.
Reflex achieves a 2.58× inference speedup while maintaining high-frequency stability, revolutionizing real-time control in robotics.
Removing implicit cues from teleoperation cripples force-awareness in manipulation tasks, but torque proxies can restore and even enhance performance.
Female visitors learned significantly better with mixed-agent tour guides, revealing gender-specific benefits in interactive educational settings.
VQ-Touch achieves superior tactile data generation while drastically reducing the need for expensive sensors, reshaping the landscape of robotic perception.
Embedding reference tokens at semantic positions allows for unprecedented precision in multi-reference video editing, setting a new benchmark for instruction quality.
Mask-free virtual try-on can now preserve intricate textures and support multiple garments simultaneously, revolutionizing digital fashion experiences.
Current MLLMs fail to provide adequate support for visually impaired individuals, particularly in anticipating navigation-critical events in real-time.
Adversarial examples in vision-language models can be detected by their tendency to stray further from the data manifold, revealing a critical vulnerability in multimodal AI systems.
False negatives in medical imaging can be mitigated by a new contrastive learning approach that leverages semantic similarities, leading to a 22.6% boost in diagnostic accuracy.
Optimizing negative prompts can drastically enhance image quality in diffusion models, reducing artifacts and improving semantic accuracy.
Token-based representations can rival traditional supervised methods in detecting animal vocalizations, challenging the dominance of established architectures.
Natural paper revisions can be harnessed to train AI agents for precise and context-aware editing of complex scientific diagrams.
SceneBind revolutionizes scene understanding by seamlessly integrating what and where across multiple modalities, achieving state-of-the-art retrieval and grounding capabilities.
VLMs can achieve 70% accuracy in building typology classification, but they often miss the broader context that human experts consider crucial.
Many autonomous GUI-agent failures can be repaired when users can see and edit plans in real-time, transforming how we interact with automation systems.
Systematic misalignments in MLLM-generated captions can be detected with 63.8% accuracy, revealing a critical flaw in image-text pairing that has been largely overlooked.
Current multimodal language models falter in scientific visualization literacy, with only Gemini surpassing human performance in select areas.
Tiramisu's tiered transformer architecture revolutionizes newspaper image understanding by explicitly modeling document hierarchy, outperforming traditional methods in reconstructing complex layouts.
Pretraining MIL networks with knowledge distillation from foundation models boosts performance and stability, especially in few-shot scenarios.
Action QFormer boosts navigation success rates from 18.8% to 56.3% by intelligently reorganizing multimodal information under action supervision.
Collaborative LLMs can significantly enhance the accuracy and clarity of MRI report generation, outperforming traditional methods in brain oncology.
Stance detection in TikTok political discourse reveals striking differences in user engagement and opinion dynamics across major political figures.
Truncating unnecessary output sequences at the first denoising step can boost DMLLM throughput by up to 31 times while enhancing accuracy on complex tasks.
Multi-modal person re-identification can significantly enhance identity matching in challenging environments, with a new Transformer-based framework setting the stage for future breakthroughs.
Achieving over 84% accuracy in myocardial infarction localization, MCF-Net revolutionizes echocardiography analysis by fusing motion cues with visual features across multiple views.
CI-Diff successfully decouples unusual attributes from common associations, enabling accurate image synthesis for rare concepts that previous models struggled to render.
Physics-informed diffusion not only enhances the realism of sign language generation but also ensures that the motions adhere to anatomical constraints, bridging the gap between semantics and biomechanics.
Landmark bias can lead to significant inaccuracies in geo-localization, but HoloGeo effectively mitigates this issue through evidence-driven reasoning, outperforming existing models.
Mobile agents can now navigate complex GUIs with unprecedented efficiency, thanks to a novel data-environment co-scaling framework.
QuReC achieves unprecedented image restoration performance by combining query-specific guidance with robust local-global feature aggregation.
Budget-aware training can boost low-turn performance by over 10% while maintaining scalability across diverse tasks.
Achieving over 80% mAP in detecting small, occluded cotton squares could revolutionize precision agriculture practices.
FilmGPT redefines video editing by turning chaotic footage into polished sequences without generating new frames, leveraging learned cinematic grammar instead.
Erasing unwanted concepts from visual generative models is now possible without sacrificing the quality of non-target outputs, thanks to Uni-AdaVD's innovative approach.
URVC transforms video coding by enabling real-time adaptability to user preferences and motion complexity, achieving unprecedented efficiency in dynamic environments.
MLLMs can now achieve over 8% improvement in spatial reasoning for egocentric scenes by leveraging a novel Ego-element Graph for enhanced perception.
AeroAct achieves real-time language-conditioned quadrotor navigation by predicting flight actions directly from visual and language inputs, without the need for future video generation.
Robots can now learn new tasks in real-time while retaining past skills, thanks to a groundbreaking dual-timescale adaptation mechanism.
Directly injecting 3D scene tokens into VLMs boosts navigation success rates and enables seamless transfer to real-world scenarios without retraining.
VLM-driven agents often succeed in tasks while neglecting critical process-level safety, highlighting a dangerous oversight in current evaluations.
Occluded-object pose tracking can be revolutionized by a kinematic-aware approach that smartly fuses haptic and visual signals, achieving up to 15 times better performance in real-world manipulation tasks.
Targeted illumination attacks can reduce VLA model task success rates to zero, exposing a critical flaw in current defense strategies that misinterpret color information.
Robots can now maintain accurate relative pose estimates even when visual overlap is absent, thanks to a novel communication-efficient framework.
Achieving competitive video generation performance with under 1% of the trainable parameters typically needed could revolutionize how we approach fine-tuning in large-scale models.
FreqLF achieves competitive disparity estimation accuracy while eliminating the need for memory-intensive cost-volume construction.
TanGO achieves state-of-the-art 3D editing by allowing fine-grained control over individual tokens, drastically reducing semantic artifacts.
Introspective attention modulation can significantly enhance the safety of T2I models without sacrificing quality, outperforming existing methods like concept erasure.
UPrompt achieves a remarkable 7.3 rSum improvement over existing models on MSCOCO, redefining the standards for cross-task generalization in vision-language tasks.
Event-based imaging can achieve coherent large-scale spatial reconstruction while effectively filtering out fine-scale textures, challenging conventional image processing methods.
Stitch-Inferencer achieves real-time endoscopic video segmentation by transforming how we manage occlusions and field-of-view limitations, outperforming traditional methods without the need for extensive retraining.
Integrating diverse visual priors can lead to significant improvements in spatial reasoning tasks, with ViPS setting new benchmarks in MLLM performance.
SSRL not only bridges the modality gap in USVI-ReID but also filters out pseudo-label noise, leading to superior performance compared to existing methods.
Simplifying the attack pipeline for VLPMs can lead to a remarkable boost in transferability, outperforming complex methods with less resource consumption.
Achieving a 76.2% relative gain in multi-step reasoning success rates, HDR redefines the capabilities of video models in real-time applications.
Achieving a 57.6% success rate on RoboCasa365, Xiaomi-Robotics-1 sets a new standard for vision-language-action models in real-world robotic manipulation.
Combining video and route descriptions can dramatically enhance geo-localization accuracy, achieving significant improvements over existing methods.
Achieving 99.3% accuracy in domain classification, this multi-expert OCR system adapts to diverse Manchu writing styles with minimal training data.
Joint modeling of tumor growth and dropout using EB-VAE reveals critical genetic indicators that could transform personalized cancer treatment strategies.
Automated analysis of multimodal cardiac imaging can match expert physician assessments, achieving a balanced accuracy of 0.76 in identifying disease-related abnormalities.
HIVE-3D transforms single-image inputs into high-resolution 3D scenes, setting a new benchmark in quality and detail.