Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
Achieving 100% event detection accuracy in extreme conditions, this framework transforms how we localize workpieces in hot forging environments.
Agreement among perturbed inputs can mislead accuracy assessments, as fine-tuning on consensus can paradoxically reduce performance.
Grounding tasks can be unified into a single correspondence prediction problem, leading to a 48% boost in performance on complex datasets.
Ego-OSCAR delivers a game-changing, budget-friendly solution for crowdsourced egocentric data collection, complete with a rich dataset and open-source tools.
MirrorWorld achieves unprecedented fidelity in mirror reflection generation by effectively modeling both what to reflect and how to spatially arrange it.
Grounding tasks can be unified into a single correspondence prediction problem, leading to a 48% boost in performance on complex datasets.
Ego-OSCAR delivers a game-changing, budget-friendly solution for crowdsourced egocentric data collection, complete with a rich dataset and open-source tools.
MirrorWorld achieves unprecedented fidelity in mirror reflection generation by effectively modeling both what to reflect and how to spatially arrange it.
Cross-view correspondence, not geometry, is the critical bottleneck in achieving reliable 3D localization of rib fractures from CT projections.
YOLO-PEFT transforms the fine-tuning landscape for real-time detectors by replacing trial-and-error with structured, auditable planning, achieving notable performance gains.
SimWAM achieves state-of-the-art performance in autonomous driving while eliminating the need for future video generation at inference, drastically reducing latency.
A multi-agent forensic reasoning framework outperforms leading closed-source models in deepfake detection by leveraging diverse analytical perspectives on forgery cues.
Artist-grounded image generation can achieve unprecedented fidelity by explicitly controlling for artistic intent, rather than relying on shortcuts that distort user vision.
EffectLearner achieves unprecedented video object removal quality by effectively reasoning about complex object-induced effects in dynamic scenes.
De-identified medical images may still reveal patient identities, challenging the assumption that such scans are truly anonymous.
Bridging implicit and explicit relational biases boosts skin lesion diagnosis accuracy by nearly 3% using graph-based methods.
Super-resolution techniques can significantly erase small white matter lesions in MRI scans, challenging assumptions about their reliability in clinical settings.
SR-JEPA reveals that a predictive pathway can infer missing entity representations in 3D scenes with remarkable accuracy, transforming how we approach latent state learning.
Achieving 100% event detection accuracy in extreme conditions, this framework transforms how we localize workpieces in hot forging environments.
Agreement among perturbed inputs can mislead accuracy assessments, as fine-tuning on consensus can paradoxically reduce performance.
Depth information can drastically improve object counting accuracy in crowded scenes, achieving over 60% reduction in counting errors.
PRISM achieves state-of-the-art results in unpaired image translation by intelligently preserving important features while allowing for targeted changes, outperforming existing methods.
VLMs may excel in scoring but often lack meaningful visual grounding, revealing critical limitations in their zero-shot control capabilities.
Action segmentation accuracy skyrockets with D-CLOT, achieving up to +12.7 F1 by resolving representation-prototype inconsistencies in unsupervised settings.
Detectability of findings in 3D CT scans hinges more on physical characteristics than model architecture, revealing a critical bottleneck in diagnostic performance.
Integrating VPR with FF3D models boosts localization accuracy, overcoming the limitations of traditional visual methods.
UQ-Loc achieves significant gains in localization accuracy by integrating uncertainty estimation directly into the SCR process, challenging the notion that deterministic predictions are sufficient.
EmoWorld achieves up to 37% improvement in emotional alignment while decoupling atmosphere, semantics, and temporal progression in video generation.
ALTER transforms how we generate CT reports by accurately modeling longitudinal changes across multiple anatomical regions, achieving state-of-the-art results in the process.
CFGPNet achieves up to 97.8% mAP in multispectral object detection, setting a new standard for accuracy and efficiency in challenging conditions.
FlaRe combines the power of neural radiance fields with explicit geometry, enabling interactive rendering and advanced scene manipulation in a single framework.
Multi-view geometric priors can drastically enhance 3D reconstruction quality, especially for complex scenes with specular surfaces.
Event-driven techniques reveal hidden dynamics, enabling EvReflection to outperform traditional methods in reflection removal by a significant margin.
Dense-Cast achieves a remarkable MAE of 0.235 mm for half-hourly precipitation nowcasting, setting a new benchmark for accuracy in this challenging domain.
Curia-MAE achieves superior performance in 3D medical image segmentation with a frozen encoder, challenging the need for extensive fine-tuning even in data-scarce environments.
Shape-aware conversion of OBB to HBB reduces background noise while preserving essential detection data, enhancing ship detection accuracy in aerial imagery.
VLMs can transform under-resourced historical languages by automating data extraction at unprecedented scales, as demonstrated by the mapping of Armenian commercial advertisements in Paris.
Achieving high-quality long video generation without any model retraining, Diff-VF redefines the capabilities of existing short-video diffusion frameworks.
HALO's innovative dual-prior approach eliminates attention drift, achieving unprecedented clarity and color accuracy in low-light remote sensing imagery.
Retailers can now generate highly accurate virtual try-ons that reflect true garment fit, reducing misleading representations in online shopping.
Achieving seamless identity replacement in videos, Vorch-IR can handle multiple subjects and backgrounds without requiring precise pose matching.
ODin redefines 3D human registration by transforming it into a generative diffusion process, achieving unprecedented accuracy and efficiency.
Decoupling coordinate frame selection from box regression leads to a remarkable 11% accuracy boost in 3D visual grounding tasks.
Models that generate convincing anomaly descriptions often fail to accurately track the corresponding instances, exposing a critical gap in video anomaly understanding.
Bayesian evidence acquisition can dramatically enhance diagnostic accuracy in whole-slide image reasoning by focusing on information gain rather than mere relevance.
G$^2$ARD-GS achieves up to 30x compression while enhancing image quality and preserving geometric fidelity, setting a new standard for 3D scene representation.
Tile-based background refinement in 360-degree telepresence can dramatically enhance perceived detail and interactivity, surpassing traditional video resolution techniques.
Coordinating global and local reasoning in long-video understanding leads to a significant 2.9 point improvement over traditional frame selection methods.
Early trajectory decoding from video diffusion models can cut planning latency by nearly 50% without sacrificing decision quality.
Aligning robot scene geometry with ARGUS enables manipulation policies to learn 4-6 times faster from diverse viewpoints, transforming their generalization capabilities.
Real-time character animation is now feasible with Wan-Animate-2, which achieves high-fidelity results without the pitfalls of traditional motion representation methods.
Achieving a balance between local detail and global coverage, this method significantly enhances UAV photogrammetry accuracy and completeness.
EviSelect achieves a 3.9x speedup in long video understanding by dynamically selecting relevant frames based on the MLLM's internal attention evidence, cutting visual token selection by half.
HOPE achieves accurate pressure estimation from monocular videos, enabling robust predictions of hand-object interactions without the need for specialized sensors or extensive labeled data.
Faked character detection in handwritten Chinese text recognition can be achieved without sacrificing recognition performance, thanks to DTRNet's innovative dual decoding approach.
By adapting model contributions to specific anatomical targets and institutional contexts, this framework significantly enhances segmentation performance in environments with scarce labeled data.
Self-supervised learning can unlock high-quality data extraction from bar charts without the need for extensive labeled datasets.
By combining multi-step latent self-prediction with observation-level dynamics, OG-SPR achieves superior performance in visual control tasks, revealing the limitations of traditional predictive methods.
OTLesMix significantly enhances synthetic lesion diversity, boosting segmentation performance by up to 6.6 points on critical medical imaging tasks.
Current XAI evaluation methods fall short, risking the effectiveness of bias detection and concept unlearning in evolving data environments.
Video language models falter dramatically in counting transient events, with less than 0.2% accuracy in high-frequency scenarios.
BioKD achieves a remarkable 68.01% accuracy in trial-wise arousal recognition, showcasing the power of reliability-aware knowledge distillation in emotion recognition tasks.
Expert-validated data and a compact model make Bangla Sign Language recognition feasible on personal devices, enhancing accessibility for the deaf community.
Structured video-grounded semantics can boost synthetic IMU generation, leading to a 19.86% improvement in tail-class recognition over traditional methods.
CogVis redefines change detection by decoupling temporal perception from semantic categorization, achieving unprecedented efficiency and accuracy across multiple benchmarks.
Grounding generative editing in physics dramatically enhances shadow removal quality, revealing that classic vision principles are still vital in AI-driven image editing.
Target-background relation shifts can cripple detection performance, but HyTBE's innovative approach expands training patterns to ensure robust generalization across unseen domains.
Implicit semantic guidance in UniVVT outperforms traditional geometric preprocessing, setting a new standard for high-fidelity video virtual try-on.
Localization errors for vehicles can be reduced by over 50% by accurately projecting their footprints onto the road plane, transforming traffic monitoring capabilities.
Prior-SG enables robots to redefine their spatial understanding in real-time, achieving zero-shot flexibility in semantic region segmentation even in the absence of physical boundaries.
Automated tooth-level mapping from smartphone images could revolutionize access to dental care in resource-limited settings.
Pretraining on synthetic data derived from CT scans can boost pose assessment accuracy in radiography by over 11%, tackling a critical barrier in patient positioning.
Achieving an FID of 1.45 in just 600 epochs, Energy-Guided Flow Matching redefines efficiency in high-quality image generation without the need for extensive model adaptations.
Fine-tuned foundation models can significantly outperform human-inspired methods in compositional analysis, but at the expense of interpretability and generalization.
Achieving robust hyperspectral image classification without access to source data could revolutionize remote sensing applications constrained by privacy regulations.
ConceptADapt achieves significant performance gains in few-shot anomaly detection by recalibrating feature statistics through a novel dynamic attention mechanism, even with minimal training data.
PaCoNet revolutionizes data extraction from parallel coordinate plots, achieving unprecedented accuracy and enabling deeper insights from complex visualizations.
UCD reveals that even state-of-the-art segmentation models like SAM3 can be severely compromised by a single, cleverly crafted adversarial perturbation.
Flow-Map Distillation achieves superior image restoration by transforming static knowledge transfer into a dynamic flow mapping process, cutting training variance in half.
Iterative refinement in LiDAR scene completion can yield significant improvements in specific contexts, but it’s not a universal solution—geometry matters.
MAVISEG reveals that diffusion transformers can retain and utilize more structured visual information than traditional methods, leading to significant improvements in segmentation accuracy.
STAIL achieves significant performance gains in medical imaging tasks while drastically minimizing storage requirements and privacy concerns associated with traditional rehearsal methods.
TTA can boost accuracy but often at the cost of calibration, and ZAEC is the key to restoring reliable confidence without labeled data.
Open-vocabulary segmentation can achieve spatially coherent and context-aware predictions without any training, thanks to SCI-CLIP's innovative use of a segment-centric inference framework.
LiteKD-Net achieves superior image denoising performance on mobile devices while slashing runtime costs, setting a new standard for efficiency in the field.
Achieving state-of-the-art HDR reconstruction quality, DOME-HDR reveals how dual-output synthesis can enhance both SDR and HDR imaging from bracketed inputs.
A single ranking can adapt to any frame budget, improving accuracy and reducing latency without retraining the model.
Pretraining on relevant scientific images can significantly boost quality assessment performance, outperforming larger datasets.
Uncertainty-aware segmentation can drastically improve performance in medical imaging, with DistMedVL achieving superior results using only 6.3M parameters.
Ambient temperature fluctuations can be harnessed to create robust adversarial attacks that consistently deceive multimodal perception systems.
Near-sensor computing slashes tactile response times from 170 ms to just 28 ms, revolutionizing robotic reflexes.
VLMs are failing to achieve human-level global spatial awareness, scoring only 42.68 on a new benchmark compared to 79.08 for humans.
U-Nets unexpectedly show greater robustness to resolution changes than anticipated, challenging assumptions about neural operator architectures in inverse imaging.
Target frame reconstruction from sparse event data achieves up to 3.29 dB improvement in PSNR, showcasing a breakthrough in video fidelity.
URNet achieves state-of-the-art RGB-D semantic segmentation performance while significantly reducing computational overhead by unifying feature extraction and fusion in a single encoder.
StreamMind's innovative architecture not only enhances long-horizon video understanding but also slashes query-to-answer latency, setting a new standard for multimodal agent performance.
Trajectory scoring in aerial navigation can be revolutionized by focusing on unexplainable prediction discrepancies, leading to more robust and efficient UAV navigation.
MT-GNN achieves a groundbreaking 2.29% improvement in predicting brain morphology, setting a new standard for accuracy in clinical imaging.
Explicit geometric priors in image restoration can outperform traditional methods while using fewer parameters, challenging the reliance on implicit regularization.
High reconstruction errors in VQ-VAD reveal anomalies in human motion, achieving 81.83% accuracy in detecting video anomalies from normal behavior patterns.
Outperforming larger models at one-tenth the computational cost, this framework redefines efficiency in classroom incident recognition.
The Predictor-Corrector sampler fails to outperform Stochastic Gradient Langevin Dynamics in practical scenarios, challenging its theoretical advantages in Joint Energy-Based Models.
ILDM outperforms traditional diffusion models by leveraging Riemannian geometry, achieving unprecedented generative quality in data-sparse environments.
Pre-computed embeddings can outperform traditional supervised models in estimating global biomass, reshaping our approach to carbon stock monitoring.
Recurrent Vision Transformers can outperform standard models in accuracy-to-parameter trade-offs when memory constraints are prioritized, challenging conventional wisdom about architectural efficiency.