Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
Inverted self-attention reveals hidden causal relationships in multivariate time series, outperforming traditional methods by reducing spurious correlations.
The Dynamic Exploration Graph achieves superior efficiency in dynamic nearest neighbor searches, outperforming existing algorithms while maintaining static dataset performance.
A frozen pixel diffusion model can self-guide its generation process using its own samples, achieving over 50% improvement in image quality with minimal additional compute.
Achieving real-time affordance segmentation on embedded devices while balancing performance and energy efficiency could revolutionize wearable robotics.
ReBA achieves over fivefold improvement in load balancing for Vision-Language MoE without sacrificing accuracy, revealing a critical interplay between token mix and load profiles.
ReBA achieves over fivefold improvement in load balancing for Vision-Language MoE without sacrificing accuracy, revealing a critical interplay between token mix and load profiles.
Achieving a 79.9% acceptance rate, Poplar transforms the landscape of human-centric image dataset creation with its structured and quality-controlled synthesis pipeline.
DreamTraj sets a new benchmark in 6-DoF trajectory prediction by leveraging language and a single image, outperforming traditional methods that rely on privileged inputs.
Achieving high-quality low-light imaging without the need for clean RGB data could revolutionize how we approach image enhancement in challenging environments.
WCM transforms the way reinforcement learning models handle temporal dynamics, leading to unprecedented generalization in robotic manipulation tasks.
Multi-dimensional Evaluation-Verification Reward transforms multi-reference image editing by providing a structured approach to evaluate and enhance visual consistency, yielding superior results over existing models.
Converged diffusion loss in visual generation improves linearly with structured language, leading to a new training paradigm that outperforms both open-weight and closed-weight models.
A frozen pixel diffusion model can self-guide its generation process using its own samples, achieving over 50% improvement in image quality with minimal additional compute.
Training-Distribution Hallucination is a critical challenge in robot manipulation, but ST-WAM's innovative use of DINOv3 features dramatically boosts performance under visual shifts.
Understanding image similarity just got clearer: this framework reveals how specific concepts like texture and shape drive similarity scores, outperforming traditional methods.
PhiZero reveals that using physical language for world modeling can significantly enhance reasoning and simulation capabilities compared to traditional pixel-based methods.
ACE-Data-0 reveals that existing models struggle significantly with complex interactions, exposing critical gaps in embodied AI performance.
ReToken boosts vision-language model performance by over 20% on key benchmarks while fitting training and inference within a single GPU's memory constraints.
GVR-Coder not only generates high-quality diagrams from complex texts but also incorporates a novel feedback mechanism that significantly enhances visual aesthetics and structural integrity.
Higher-resolution satellite imagery can enhance solar irradiance retrieval accuracy under cloudy conditions, but fails to solve clear-sky challenges.
Contextual embeddings can revolutionize how we evaluate expressive MIDI performances, aligning closely with human perception while overcoming traditional metric limitations.
A humanoid robot can dodge 95% of thrown balls using only depth data from a head-mounted camera, showcasing the potential of perception-aware safety in dynamic environments.
Isolating global reasoning from local evidence in multimodal retrieval can dramatically boost QA accuracy and evidence recall.
Bridging neural perception with symbolic reasoning, this framework reveals how visual defects translate into actionable severity scores, enhancing interpretability in sewer assessments.
Identifying when perception, rather than reasoning, is the failure point can boost multimodal reasoning performance by over 2.5 points on key benchmarks.
Despite the rise of multimodal models, even the most advanced struggle with fine-grained visual tasks in pathology, revealing critical gaps in their understanding.
Achieving an order-of-magnitude reduction in memory usage, MonoVoc enables efficient open-vocabulary 3D scene understanding directly from monocular video.
Latent objects can revolutionize how Video-LLMs manage memory, achieving a 10-point performance boost while slashing memory usage by 50%.
Supervising latent representations at the reasoning-process level leads to a breakthrough in multimodal reasoning, outperforming traditional methods by a significant margin.
MUL-T achieves state-of-the-art performance in tissue image analysis with a fraction of the parameters, redefining efficiency in spatial cellular architecture modeling.
MMLDSum-LLM outperforms existing models by significantly enhancing key information coverage and cross-modal consistency in long-document summarization.
Reflective Retrieval Memory transforms how agents retrieve and utilize past experiences, leading to superior performance in long-horizon multimodal reasoning tasks.
Query-grounded visual sampling can boost long video understanding in LVLMs by over 11% without the need for extensive training.
Detailed long captions can reduce depth estimation errors by up to 25% in challenging visual conditions, transforming how we approach monocular depth estimation.
Foundation models can revolutionize hand-object interaction by systematically categorizing and leveraging diverse priors to enhance robot learning and task performance.
A two-stage search approach can slash query costs by up to 15 times while maintaining high accuracy in identifying changes in satellite images.
Ghosting artifacts are dramatically reduced and rendering quality improved in UAV scene reconstructions through adaptive feature aggregation tailored to local dynamics.
MIND achieves a breakthrough in medical image fusion by integrating intent-driven diagnostics, resulting in significantly improved segmentation accuracy for brain tumors.
Achieving unprecedented accuracy in 3D shape reconstruction, this method captures fine details even in the most challenging visual conditions.
BladeYOLO not only detects subtle wind turbine blade defects but does so with remarkable accuracy even when trained on limited annotations, outperforming existing methods by a significant margin.
ReGenVC achieves artifact-free video reconstruction at one-tenth the bitrate of traditional codecs, enabling real-time streaming without frame drops.
ViP-Rig enables artists to achieve precise control over skeletal structures and skinning weights through intuitive visual prompts, outperforming traditional geometry-based methods.
VCP-DCN reveals that leveraging depth-specific features can dramatically enhance the detection of camouflaged objects, outperforming traditional RGB-D methods.
CoRE-UIR achieves a remarkable 1.05 dB PSNR improvement while being 11.83 times faster than the leading image restoration method, BaryIR.
Even with stable feature representations, AI detectors can still fail due to drifting decision boundaries, leading to a new failure mode called Dual Degradation.
Direct visual adaptation in few-shot medical image segmentation outperforms prompt-based strategies, revealing critical insights into model performance.
SLQA reveals that models can achieve deeper comprehension of sign language by answering complex questions, rather than merely translating or recognizing signs.
FA-RDP achieves superior success rates in contact-rich manipulation while preserving diverse action modes, revolutionizing how we approach multimodal decision-making in robotics.
Clustering performance improves dramatically with DAS-PMVC, which effectively tackles the partial view alignment problem that plagues traditional multi-view methods.
FasTac reduces depth estimation error by over 85% while processing tactile data more than three times faster than conventional methods.
Longitudinal medical visual reasoning is critically underexplored, with existing models failing to grasp temporal nuances, as evidenced by their poor performance on the new LoMeVQA benchmark.
Tactile supervision can boost robot manipulation success rates by over 37% compared to traditional visual-only models.
End-effector traces boost cross-embodiment transfer in robot manipulation, enhancing real-world task performance by 28% when leveraging simulation data.
Live music recordings can now be effectively separated using new datasets that account for venue acoustics and audience noise, transforming MSS capabilities.
VocalRender achieves a remarkable $0.42$ improvement in naturalness over the strongest baseline, revolutionizing singing voice synthesis for real-world composition.
VIG-RL achieves a new state-of-the-art in Verified Image Grounding by dynamically integrating visual evidence into text responses, outperforming static retrieval methods.
The Dynamic Exploration Graph achieves superior efficiency in dynamic nearest neighbor searches, outperforming existing algorithms while maintaining static dataset performance.
Switching document formats can lead to accuracy drops of over 53%, revealing a hidden vulnerability in LLM workflows that demands urgent attention.
Achieving photorealistic 3D head avatars from a single image, S-Avatar sets a new standard for realism and consistency in virtual reality applications.
A deep learning model achieves up to 99% accuracy in insect classification by leveraging a hierarchical taxonomy and a massive dataset from camera traps.
Supervised learning outperforms self-supervised methods in flood monitoring, revealing critical insights for data-limited environments.
LAST achieves 95.4% accuracy with only 12.5% of visual tokens, revolutionizing edge-cloud inference efficiency.
Achieving robust 3D spatial reasoning without any training, ViewMind3D redefines the landscape of 3D question answering.
Malicious audio instructions can stealthily hijack multimodal agents, achieving a 69.10% success rate in real-world scenarios.
Chimera achieves 7.3x compute efficiency over traditional models while enabling zero-shot extrapolation from short video clips to significantly longer sequences.
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
SpatialCLI enables VLMs to achieve a staggering 84.6% accuracy on spatial reasoning tasks, far surpassing existing models.
LLM-generated feature programs enable a data-efficient approach to scar classification, outperforming traditional methods while keeping patient data secure and decisions auditable.
AAPT boosts GUI agent success rates by 58% in decision-critical moments, eliminating delays without sacrificing accuracy.
Achieving real-time affordance segmentation on embedded devices while balancing performance and energy efficiency could revolutionize wearable robotics.
Inverted self-attention reveals hidden causal relationships in multivariate time series, outperforming traditional methods by reducing spurious correlations.
Dual-Ambiguity Rectification can enhance image restoration performance by disentangling complex degradation cues, leading to cleaner outputs and fewer artifacts.
EndoCLIP can match expert endoscopists in classifying lesions, showcasing the transformative potential of linking clinical findings to images in colonoscopy.
HyperClaim outperforms traditional methods by achieving up to 87.3% accuracy in detecting video misinformation through advanced hypergraph reasoning that captures complex cross-modal interactions.
High-fidelity captions generated for disaster response images reveal a surprising 78.65 semantic agreement, exposing inconsistencies in human annotations.
Current pathology foundation models can rival fully trained architectures in detecting mitotic figures, showcasing their potential beyond mere classification tasks.
Simple image transformations can undermine even the most advanced AI-based content moderation systems, exposing significant vulnerabilities.
Outperforming GPT-5.4, IndustryForge-27B achieves a remarkable 33.65 percentage point lift in CAD-specific tasks, setting a new standard for multimodal models in industrial applications.
Three frontier models achieved near-perfect accuracy in generating ASP theories, while GPT-5's performance varied dramatically, revealing the critical influence of model design on task success.
Transforming non-verbal signals into a unified semantic space allows LLMs to grasp complex human emotions more effectively than ever before.
LVLMs may not be as adept at reasoning as previously thought, with new benchmarks revealing significant gaps in their performance.
DualAnchor not only preserves language priors but also bridges the lexical fidelity gap in sign language translation, leading to significantly improved translation quality.
Achieving high-fidelity 3D generation with just 1.5% of the training data could revolutionize resource allocation in 3D modeling.
GenEvA achieves up to 10.1 points higher accuracy in long-video understanding while using less than 0.4% additional video tokens.
"Late-blooming" tokens can be crucial for deep-layer reasoning, and our method ensures they aren't discarded prematurely, achieving over 77.8% token reduction without sacrificing performance.
Integrating street-level imagery with satellite data boosts crop classification accuracy, transforming agricultural monitoring practices.
Reinforcement learning empowers Vision-Language Models to effectively reason about AI-generated image edits, achieving high accuracy with minimal supervision.
TARS achieves robust camera and viewpoint control in video re-shooting without relying on 3D reconstruction or paired data, enabling the plausible synthesis of previously unseen regions.
Scaling VLMs doesn't guarantee bias mitigation; in fact, larger models can falter significantly on complex bias evaluations.
FaithEyes reveals that self-judging mechanisms in VLMs can drastically improve tool use fidelity, leading to more reliable multimodal reasoning.
FarmSeeker redefines farmland segmentation by actively querying spatio-temporal information, leading to significantly improved accuracy and stability in challenging environments.
Transforming low-quality facial images into high-resolution representations while enhancing re-identification accuracy could redefine standards in facial recognition technology.
Embedding face and voice features into a convex hull can drastically reduce false associations, leading to unprecedented accuracy in cross-modal tasks.
Achieving 70.8% recall of radar response regions, mmRadarTwin reveals critical insights into the challenges of indoor mmWave radar perception.
Achieving photorealistic head avatars from a single image without external tracking, SpiD sets a new standard for speed and fidelity in digital human synthesis.
Event cameras can boost video compression performance, yielding up to 22% better efficiency by refining motion estimates in challenging conditions.
Single-Patch Text Spotting achieves unprecedented accuracy in scene text spotting by leveraging a single visual token per instance, outperforming leading models in the field.
A late-fusion ensemble approach not only topped the Concept Detection leaderboard but also highlighted the potential of training-free methods to match fine-tuned performance at a fraction of the cost.
A novel gating mechanism that integrates contextual information from both encoder and decoder dramatically boosts performance on minority class detection in nucleus segmentation tasks.
Hallucinations in LVLMs can be cut by over 43% without sacrificing grounded object coverage, thanks to a novel verifier-guided approach.
Routing previously encoded evidence in HR-VQA can yield up to 9.9-point improvements on benchmarks while slashing inference time by over 97%.
Agricultural navigation systems can now better handle ambiguous traversability, improving success rates by 15% in challenging environments.
Recovering full-body meshes from head poses is now over 50 times faster without sacrificing quality, thanks to EgoGVAE's innovative approach.
Achieving a 42.67X reduction in data payload for robotic visual communication could revolutionize how robots collaborate in complex environments.
Reliable robotic agency emerges not from bigger models, but from the strategic orchestration of modular components that enhance VLA capabilities.