Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Achieving double the speed of the leading five-point pose estimation method without sacrificing precision could revolutionize real-time camera tracking applications.
Achieving real-time, high-accuracy point correspondences without 3D scene knowledge could revolutionize image processing in dynamic environments.
Efficiently reducing metal artifacts in CBCT without manual intervention, our method sets a new standard for reconstruction accuracy and practicality in medical imaging.
Current AI-generated videos that mislead viewers are also the most challenging for existing detection systems to identify, revealing a critical vulnerability in misinformation defenses.
CPI-Bench reveals significant performance gaps among image editing models, offering a more nuanced evaluation that aligns with real-world user experiences.
Current AI-generated videos that mislead viewers are also the most challenging for existing detection systems to identify, revealing a critical vulnerability in misinformation defenses.
CPI-Bench reveals significant performance gaps among image editing models, offering a more nuanced evaluation that aligns with real-world user experiences.
SPARGen achieves competitive performance across diverse spatial tasks by unifying perception and reasoning into a single generative framework, challenging the need for separate architectures.
Achieving high-fidelity 3D object generation with up to 300 parts, MegaParts redefines the limits of part-aware modeling through token-efficient autoregressive techniques.
Unlocking the potential of multimodal AI to interpret complex scientific images could revolutionize how we access and understand experimental data.
GeoCache achieves over 2x speedup in multi-view texture diffusion without sacrificing fidelity by exploiting cross-view geometric redundancy.
Certified caching can boost edge image classification speed by 1.65x without compromising reliability.
Uniform Herding not only boosts accuracy but also reduces forgetting, outperforming traditional methods by effectively managing exemplar representation across tasks.
UniTraffic-Agent achieves top-tier performance in complex traffic video reasoning tasks, showcasing the potential of structured workflows in multimodal AI applications.
SSIM-based anchoring offers a more reliable fidelity control for denoising outputs, outperforming traditional PSNR methods across diverse noise conditions.
Achieving superior 3D scene reconstruction from a single measurement, this method combines Gaussian Splatting with advanced vision model priors to tackle the complexities of Snapshot Compressive Imaging.
Achieving a 12.1% boost in robust concept removal, MapRoute++ redefines the boundaries of visual concept unlearning.
Achieving state-of-the-art hand gesture recognition with a unified model that combines the strengths of CNNs and transformers while using fewer resources is a game changer for real-time applications.
SketchSense achieves remarkable gains in image inpainting by intelligently interpreting imperfect sketch guidance, outperforming traditional methods in both quality and structural accuracy.
RbFT-Net reveals that correcting radar measurements before fusion can significantly enhance depth completion accuracy in autonomous systems.
Contrast-based measures can significantly enhance the objectivity of evaluating historical manuscript restorations, outperforming traditional image quality metrics.
Non-ML cultural heritage experts can now independently analyze artifacts and validate findings using an intuitive, open-source computer vision platform.
Bridging the gaps in plant growth modeling, this framework enables accurate tracking of organ development over time, even with sparse data.
Training on the PassGen framework allows robots to anticipate human intentions earlier and more accurately, transforming human-robot collaboration.
Identifying intra-image predictive subsets can significantly boost visual classification performance, especially in challenging data shift scenarios.
TennisVAR redefines sports video analysis by grounding tactical reasoning in stroke-level evidence, enabling deeper insights into match strategies.
Smartphone photogrammetry can match commercial body scanners in accuracy, making 3D body scanning more accessible for healthcare applications.
Misclassification between known and unknown objects can be effectively mitigated using a unified framework that leverages dual perspectives on object discovery.
VOS-Agent outperforms existing methods by leveraging specialized agents for different target types, achieving state-of-the-art results in video object segmentation.
ASPIRE-VINS reduces trajectory estimation errors by adapting knot placement and refining splines based on local motion variations.
P2Fusion achieves state-of-the-art fusion quality by adaptively regulating modal competition through learnable dynamic regulators, outperforming existing methods in 14 out of 20 evaluation metrics.
GATO-Vid achieves superior spatial localization in text-to-video generation without the computational costs of traditional gradient-based methods.
Moving beyond anchor boxes, this method leverages signed distance functions to achieve superior instance segmentation performance, particularly for irregularly shaped objects.
A hybrid concept bottleneck model boosts interpretability in cancer imaging diagnostics with only 10% annotation, achieving a mean concept AUC increase from 0.619 to 0.741 for mammographic masses.
Generated retinal images can inherit critical clinical information, but they may not perform well against real-world classifiers, exposing a crucial representation gap.
DSCC is the first method to effectively ground long-form captions in multimodal models by integrating visual anchors during training, achieving unprecedented precision and length in caption generation.
Trusting the wrong pseudo-labels can lead to confirmation bias, but CW-BASS v2 intelligently adapts to the confidence landscape of strong foundation models, yielding significant performance improvements.
Every exactly calibrated convex surrogate for the Jaccard measure demands exponentially more prediction dimensions than previously understood.
TabSOM not only outperforms existing tabular-to-image methods but also enhances interpretability by revealing feature relationships and interactions.
ARMDIL outperforms traditional routers by dynamically selecting the best-suited model for each image, significantly enhancing cross-domain generalization.
Foundation models outperform traditional supervised methods in fall and stress detection, challenging the notion that bigger always means better in health monitoring tasks.
Complex models may shine in simulations, but they struggle in real-world fall detection, revealing the critical role of representation choice.
MergeOver reduces peak activation memory by over 37% while maintaining competitive accuracy, making it a game-changer for deploying Vision Transformers on edge devices.
GCache achieves a remarkable 2.17x speedup in video diffusion while enhancing visual quality, challenging the effectiveness of traditional caching heuristics.
CaC2-treated fruits exhibit distinct spectral signatures that can be detected with 95% accuracy, revealing a critical tool for food safety in the fruit industry.
Achieving over 30 PSNR in sign language video synthesis with a GAN framework that balances stability and detail could revolutionize communication for the hearing impaired.
IMU-based sensing outperforms egocentric vision in detecting freezing of gait, but the latter reveals crucial contextual insights that could transform clinical assessments.
Current MLLMs struggle with long-term memory, achieving only 71.8% accuracy on a benchmark where humans score 94.2%.
Vision-language models struggle to leverage visual evidence in medical VQA, with only one model surpassing human performance on a subset of questions.
SNM-VFI achieves superior perceptual quality and temporal coherence in video frame interpolation by effectively integrating flow-guided frames with diffusion-generated details.
Detecting partially forged videos is now feasible with a novel framework that leverages static images for enhanced supervision and accuracy.
SCULPT enables simultaneous part extraction and object reconstruction, overcoming the limitations of traditional segmentation and additive methods in 3D generation.
FineX boosts fine-grained action recognition accuracy by over 7% on challenging datasets, leveraging a unique fusion of visual and skeletal representations.
HPSD enables TI2V models to internalize high-quality visual cues, resulting in a remarkable boost in text-to-video performance while simultaneously enhancing image-to-video generation.
Achieving state-of-the-art character erasure without sacrificing image fidelity, this method transforms how we handle copyright in AI-generated content.
Inter-member disagreement in deep ensembles serves as a more sensitive indicator of model uncertainty than single-model confidence, especially under data shifts.
ProPose not only bridges the gap in pose estimation for diverse limb types but also enhances accuracy for long-tail prosthetic joints through innovative structure-aware loss functions.
Efficiently reducing metal artifacts in CBCT without manual intervention, our method sets a new standard for reconstruction accuracy and practicality in medical imaging.
Achieving state-of-the-art performance in 3D perception tasks, GeoUP reveals that integrating geometry-grounded representations can significantly enhance autonomous driving systems.
Achieving real-time, high-accuracy point correspondences without 3D scene knowledge could revolutionize image processing in dynamic environments.
Achieving double the speed of the leading five-point pose estimation method without sacrificing precision could revolutionize real-time camera tracking applications.
A unified framework that effectively fuses RGB and event data can drastically improve person re-identification accuracy across different camera views.
Active LED markers enable AUVs to maintain precise relative localization even in murky underwater conditions, overcoming traditional vision-based limitations.
Spatially grounded tokens in LocusGS lead to coherent Gaussian distributions that dramatically improve rendering quality in 3D scene synthesis.
RippleNet reveals that focusing on low-SNR forgery traces can significantly enhance the detection of AI-generated images, outperforming traditional methods.
GDI transforms defect classification by generating single-defect samples, leading to a remarkable 63.6% boost in F1-Score for rare defects.
Class-geometry supervision transforms how we approach sample-efficient open-world detection, yielding significant gains in novel-class insertion and unknown recall.
Achieving over 98% accuracy in classifying infected mosquitoes from video frames showcases the power of combining vision and language models for biological analysis.
DeSCon not only balances training data but also addresses bias in the critical tail of the non-match score distribution, leading to fairer face recognition outcomes.
Achieving a 5.28% improvement in segmentation accuracy while running 4.7% faster than existing methods, DiCoR redefines efficiency in referring remote sensing image segmentation.
EgoPHI transforms egocentric vision by enabling precise 3D force estimation from a single image, bridging the gap between contact localization and physical interaction reasoning.
Multi-view MRI inputs can enhance spatial localization but may compromise temporal reasoning, revealing critical limitations in current foundation models for clinical use.
HumanScore reveals that traditional kinematic metrics overlook critical failures in humanoid motion tracking, such as unstable support and incorrect contacts.
Latent SDS can produce noisy artifacts, but PixSDS offers a novel solution that preserves semantic content while reducing these distortions.
Achieving state-of-the-art performance in document parsing, NaviDC-OCR tackles geometric distortions and structural reasoning challenges that plague existing models.
V-RAE achieves a remarkable 2.13 rFVD on K600, outperforming traditional video VAEs by retaining significantly more semantic information in its latent representations.
Evaluations reveal that current models struggle with long-range narrative integration and cultural reasoning, highlighting a critical gap in video understanding capabilities.
Earth observation embeddings can boost weather downscaling accuracy by over 11% by effectively encoding persistent surface properties.
Achieving over 96% accuracy in exercise quality assessment could revolutionize remote rehabilitation by enabling effective, therapist-free patient monitoring.
SPAR-derived pulse wave images can accurately classify age groups with over 70% F1 score, revealing hidden patterns in cardiovascular health.
By reversing the conventional stroke generation process, this method allows for more intuitive and context-aware sketching directly from text prompts.
Urban expansion in Dhaka has surged by nearly 60%, with alarming declines in vegetation and water bodies, highlighting the urgent need for sustainable urban planning.
Achieving 94% accuracy in nationwide cashew orchard detection using satellite imagery could revolutionize environmental monitoring and agricultural management in Guinea-Bissau.
Radar imagery can outperform traditional methods in estimating air traffic complexity, achieving over 96% accuracy in modeling critical operational metrics.
VGG16 outshines other CNN architectures in Alzheimer's detection, but classifying early-stage dementia remains a significant hurdle.
Expert reassessment of ambiguous samples boosted model performance by up to 8.1 percentage points, revealing the hidden potential in small, imbalanced datasets.
AI-art detectors misclassify up to 40% of images from new generative models, revealing a dangerous vulnerability in copyright and authenticity verification.
Saliency-guided cutouts can improve image classification for natural images but fail to enhance malware detection, revealing a critical domain dependency in training strategies.
The evolution of Class Activation Mapping reveals a shift from simplistic CNN explanations to sophisticated, multi-layered insights that challenge traditional evaluation metrics.
NA-UNETR achieves unprecedented accuracy in segmenting the elusive LAD artery, outperforming traditional models and setting a new standard for cardiac imaging in radiotherapy.
Achieving 72.3% accuracy in freshness prediction with just three labeled days per fillet reveals a breakthrough in few-shot learning for food quality assessment.
SGNet achieves near-perfect freshness classification with a fraction of the parameters of conventional models, paving the way for real-time applications in food safety.
M-Net achieves up to a 12.37% improvement in segmentation accuracy by integrating mathematical features into deep learning architectures, challenging the notion that data alone suffices for optimal performance.
By reformulating spatial-temporal reasoning into localized coupled graph aggregation, HSTGFormer significantly enhances 3D human pose estimation accuracy while maintaining computational efficiency.
Identifying the right model for clinical deployment without target labels is fraught with challenges, leaving a significant performance gap even with advanced selection methods.
Achieving a 3.2x speedup in video diffusion without sacrificing fidelity could redefine efficiency benchmarks in the field.
Achieving a p-value of 0.0078 in the RxRx1 study highlights a powerful new approach to disentangling treatment effects from complex structured outcomes.
CoQui achieves superior image generation quality with fewer qubits by decoupling pixel control from quantum resource constraints.
Deformable convolution outperforms traditional methods, achieving a remarkable 20.79 dB PSNR and 0.9623 Dice score in nanophotonic absorber design.
A training-free multi-agent system for UAV image understanding not only surpasses leading models in accuracy but also addresses fundamental reasoning failures in MLLM applications.
Monocular 3D object detection can achieve robust performance without the pitfalls of 2D-to-3D lifting, thanks to a novel integration of metric reconstruction and detection.
FS-JEPA redefines how we optimize KANs by predicting structured edge function signatures, resulting in a significant boost in medical image segmentation performance.
SST-WSVADL reveals how targeted spatio-temporal analysis can mitigate background bias in anomaly detection, paving the way for more ethical AI systems.
Pooled AUC can mislead researchers by masking significant localization differences in anomaly detection models, highlighting the need for more granular evaluation methods.
Uncalibrated uncertainty in diffusion model-derived qMRI can mislead interpretations, but effective calibration transforms it into a powerful tool for reliability assessment.