Search papers, labs, and topics across Lattice.
Image recognition, object detection, segmentation, video understanding, and visual generation.
#2 of 24
1
EditaLive achieves real-time character video editing with state-of-the-art performance, preserving facial expressions while eliminating latency issues in live streaming.
TADP achieves an impressive 80.91% mAP on the KITTI dataset, setting a new benchmark for single-stage 3D object detection.
Editing long videos with multiple instructions can be done without hallucinations or loss of temporal continuity, thanks to a novel agentic framework that combines LLMs and VLMs.
QuantumBoostNet outperforms traditional models in cardiac ultrasound view identification, showcasing the potential of hybrid classical-quantum approaches in medical imaging.
A single destructive image can now yield comprehensive population-level statistics on crack degradation in lithium-ion cathodes, transforming how we assess battery aging.
SAGE achieves state-of-the-art forecasting accuracy by integrating multimodal semantic knowledge without the computational burden of large language models.
PETs that achieve similar classification accuracy can perform drastically differently across various vision tasks, revealing hidden vulnerabilities in their effectiveness.
Iterative uncertainty correction allows for accurate parameter estimation in inverse problems, even with incomplete prior knowledge.
Physics-informed deep learning can automate complex radiotherapy planning while achieving superior organ-at-risk sparing.
Achieving a record 38.8% mIoU on the SemanticKITTI hidden test with a single-sweep, single-sample approach, this work redefines the limits of LiDAR scene completion.
Coupling diffusion modeling with equivariant priors boosts inpainting quality, achieving remarkable results even with limited data.
LeVJEPA achieves up to 20.8x less pretraining compute while surpassing the performance of leading video representation methods, reshaping the landscape of video-based learning.
Generating physically accurate simulations just got faster—this method cuts deviations while sidestepping costly numerical simulations.
HCS not only detects AI-generated images with high accuracy but also reveals how different generation methods influence detection decisions through a structured analysis of CNN activations.
Property-specific discrepancies in PPG-to-rPPG recoverability reveal that not all physiological signals are equally preserved under varying observation conditions.
Four new event-based datasets could redefine the landscape of spiking neural network research by providing the high-quality data needed for robust object classification.
Transient objects in 3D reconstructions can be filtered out without retraining, significantly enhancing novel-view quality.
T2S transforms open-vocabulary semantic segmentation by generating precise seed points from text, leading to superior segmentation without the need for training.
Pixel-level table compression can dramatically reduce token usage while enhancing accuracy in document question answering, challenging conventional methods.
OCSD reduces path error by over 30% and generates more realistic long-term human motion forecasts by effectively integrating object cues and social interactions.
MedFG-VQA achieves superior performance in medical VQA tasks while being lightweight enough for practical clinical use, thanks to innovative memory and graph attention techniques.
Contact-induced constraints can dramatically enhance localization accuracy in underwater environments plagued by visual ambiguity and drift.
MedREAL achieves a remarkable 68.49% gIoU and 70.47% cIoU, setting a new benchmark for interpretable medical image analysis that aligns reasoning with pixel-level accuracy.
Achieving high-fidelity video virtual try-on in real time, LiveVVT reduces latency by 26x while enhancing generation quality.
Generative image retrieval just got a major upgrade—PailitaoGR boosts performance by 13.8% by mastering target focus and auxiliary evidence utilization.
CoGeo-GS achieves superior multi-object removal in 3D scenes by integrating concept-driven tagging with geometry-aware completion, outperforming traditional methods in both quality and stability.
VIG-Sampler boosts multimodal generation quality by leveraging visual attention, outperforming traditional methods with fewer decoding steps.
FIDA's innovative approach allows backdoor attacks to evade traditional defenses while preserving the functionality of facial recognition systems.
SSMB sets a new standard in keypoint detection under motion blur, achieving superior performance without relying on deblurring or handcrafted features.
Reducing model parameters by over 70% while boosting tumor detection rates reveals a critical trade-off in medical imaging segmentation tasks.
KDG-SemNOMA transforms 6G robotic vehicle communications by significantly enhancing visual perception while minimizing bandwidth and energy usage.
G2D boosts zero-shot image classification accuracy by up to 27.42 percentage points by effectively combining generative verification with discriminative retrieval.
GeoMAD achieves superior anomaly detection by seamlessly integrating geometric correspondence with distributional consistency, all while maintaining efficiency in 2D feature-space learning.
Malicious semantics can be stealthily injected into videos over time, revealing a critical vulnerability in I2V generation models that traditional single-frame attacks overlook.
Masking just 20 Visual Retrieval Heads in VLMs can lead to an 80-point drop in grounding accuracy, revealing their critical role in visual information extraction.
MILO redefines 3D human-object interaction reconstruction by leveraging Large Reconstruction Models, achieving unprecedented accuracy from just a single image.
Hard Negative Mining boosts precision-recall performance from 0.204 to 0.913, transforming the detection of Christmas tree plantations in complex landscapes.
Achieving superior accuracy in fetal limb assessment, UniFLM bridges critical gaps in ultrasound image analysis that have long hindered the detection of skeletal dysplasias.
Temporal coverage in Earth Observation embeddings is a tunable cost, enabling near-real-time mapping without sacrificing accuracy for certain land-use classes.
NeRF-based methods outperform traditional techniques in generating high-fidelity holograms of complex laboratory objects, transforming educational visualization.
Integrating depth information with visual data leads to a dramatic boost in 3D awareness, outperforming traditional methods on key benchmarks.
Unsupervised domain adaptation can dramatically enhance 3D CBCT segmentation accuracy without requiring any target-domain annotations.
Sidecar boosts character consistency in visual storytelling by seamlessly infusing missing identity semantics into prompts, all without the need for retraining.
CODE outperforms prior methods in Open World Object Detection by effectively balancing known and unknown object detection through innovative calibration and suppression techniques.
VLMs exhibit a significant performance gap in visual text understanding, with even the best models falling short of human accuracy in error correction tasks.
TetherMem enables video generation models to dynamically adapt scenes while keeping subjects consistent, achieving a significant leap in overall video quality and scene progression.
Storage limitations, not compute power, dictate the scalability of whole-slide image embedding extraction, reshaping our approach to handling large-scale medical data.
Ancient-Bench reveals that even advanced models fail to solve the challenges of recognizing ancient Chinese texts, underscoring a critical gap in AI capabilities.
SpatialCrafter achieves unprecedented 3D consistency in image-to-scene generation, effectively eliminating long-term drift and enhancing detail fidelity.
Aligning object poses with their geometric axes can dramatically enhance estimation accuracy while simplifying model architecture requirements.