Search papers, labs, and topics across Lattice.
100 papers published across 1 lab.
OCR technologies are evolving rapidly, yet significant challenges remain in recognizing diverse scripts and handwritten text, demanding innovative solutions for real-time applications.
MedFG-VQA achieves superior performance in medical VQA tasks while being lightweight enough for practical clinical use, thanks to innovative memory and graph attention techniques.
TAU-Agent leverages a novel retrieval mechanism to enhance traffic anomaly detection, achieving impressive benchmark results that challenge existing paradigms in video analysis.
EditaLive achieves real-time character video editing with state-of-the-art performance, preserving facial expressions while eliminating latency issues in live streaming.
TADP achieves an impressive 80.91% mAP on the KITTI dataset, setting a new benchmark for single-stage 3D object detection.
EditaLive achieves real-time character video editing with state-of-the-art performance, preserving facial expressions while eliminating latency issues in live streaming.
TADP achieves an impressive 80.91% mAP on the KITTI dataset, setting a new benchmark for single-stage 3D object detection.
Editing long videos with multiple instructions can be done without hallucinations or loss of temporal continuity, thanks to a novel agentic framework that combines LLMs and VLMs.
QuantumBoostNet outperforms traditional models in cardiac ultrasound view identification, showcasing the potential of hybrid classical-quantum approaches in medical imaging.
A single destructive image can now yield comprehensive population-level statistics on crack degradation in lithium-ion cathodes, transforming how we assess battery aging.
SAGE achieves state-of-the-art forecasting accuracy by integrating multimodal semantic knowledge without the computational burden of large language models.
PETs that achieve similar classification accuracy can perform drastically differently across various vision tasks, revealing hidden vulnerabilities in their effectiveness.
Iterative uncertainty correction allows for accurate parameter estimation in inverse problems, even with incomplete prior knowledge.
Physics-informed deep learning can automate complex radiotherapy planning while achieving superior organ-at-risk sparing.
Achieving a record 38.8% mIoU on the SemanticKITTI hidden test with a single-sweep, single-sample approach, this work redefines the limits of LiDAR scene completion.
Coupling diffusion modeling with equivariant priors boosts inpainting quality, achieving remarkable results even with limited data.
LeVJEPA achieves up to 20.8x less pretraining compute while surpassing the performance of leading video representation methods, reshaping the landscape of video-based learning.
Generating physically accurate simulations just got faster—this method cuts deviations while sidestepping costly numerical simulations.
HCS not only detects AI-generated images with high accuracy but also reveals how different generation methods influence detection decisions through a structured analysis of CNN activations.
Property-specific discrepancies in PPG-to-rPPG recoverability reveal that not all physiological signals are equally preserved under varying observation conditions.
Four new event-based datasets could redefine the landscape of spiking neural network research by providing the high-quality data needed for robust object classification.
Transient objects in 3D reconstructions can be filtered out without retraining, significantly enhancing novel-view quality.
T2S transforms open-vocabulary semantic segmentation by generating precise seed points from text, leading to superior segmentation without the need for training.
Pixel-level table compression can dramatically reduce token usage while enhancing accuracy in document question answering, challenging conventional methods.
OCSD reduces path error by over 30% and generates more realistic long-term human motion forecasts by effectively integrating object cues and social interactions.
MedFG-VQA achieves superior performance in medical VQA tasks while being lightweight enough for practical clinical use, thanks to innovative memory and graph attention techniques.
Contact-induced constraints can dramatically enhance localization accuracy in underwater environments plagued by visual ambiguity and drift.
MedREAL achieves a remarkable 68.49% gIoU and 70.47% cIoU, setting a new benchmark for interpretable medical image analysis that aligns reasoning with pixel-level accuracy.
Achieving high-fidelity video virtual try-on in real time, LiveVVT reduces latency by 26x while enhancing generation quality.
Generative image retrieval just got a major upgrade—PailitaoGR boosts performance by 13.8% by mastering target focus and auxiliary evidence utilization.
CoGeo-GS achieves superior multi-object removal in 3D scenes by integrating concept-driven tagging with geometry-aware completion, outperforming traditional methods in both quality and stability.
VIG-Sampler boosts multimodal generation quality by leveraging visual attention, outperforming traditional methods with fewer decoding steps.
FIDA's innovative approach allows backdoor attacks to evade traditional defenses while preserving the functionality of facial recognition systems.
SSMB sets a new standard in keypoint detection under motion blur, achieving superior performance without relying on deblurring or handcrafted features.
Reducing model parameters by over 70% while boosting tumor detection rates reveals a critical trade-off in medical imaging segmentation tasks.
KDG-SemNOMA transforms 6G robotic vehicle communications by significantly enhancing visual perception while minimizing bandwidth and energy usage.
G2D boosts zero-shot image classification accuracy by up to 27.42 percentage points by effectively combining generative verification with discriminative retrieval.
GeoMAD achieves superior anomaly detection by seamlessly integrating geometric correspondence with distributional consistency, all while maintaining efficiency in 2D feature-space learning.
Malicious semantics can be stealthily injected into videos over time, revealing a critical vulnerability in I2V generation models that traditional single-frame attacks overlook.
Masking just 20 Visual Retrieval Heads in VLMs can lead to an 80-point drop in grounding accuracy, revealing their critical role in visual information extraction.
MILO redefines 3D human-object interaction reconstruction by leveraging Large Reconstruction Models, achieving unprecedented accuracy from just a single image.
Hard Negative Mining boosts precision-recall performance from 0.204 to 0.913, transforming the detection of Christmas tree plantations in complex landscapes.
Achieving superior accuracy in fetal limb assessment, UniFLM bridges critical gaps in ultrasound image analysis that have long hindered the detection of skeletal dysplasias.
Temporal coverage in Earth Observation embeddings is a tunable cost, enabling near-real-time mapping without sacrificing accuracy for certain land-use classes.
NeRF-based methods outperform traditional techniques in generating high-fidelity holograms of complex laboratory objects, transforming educational visualization.
Integrating depth information with visual data leads to a dramatic boost in 3D awareness, outperforming traditional methods on key benchmarks.
Unsupervised domain adaptation can dramatically enhance 3D CBCT segmentation accuracy without requiring any target-domain annotations.
Sidecar boosts character consistency in visual storytelling by seamlessly infusing missing identity semantics into prompts, all without the need for retraining.
CODE outperforms prior methods in Open World Object Detection by effectively balancing known and unknown object detection through innovative calibration and suppression techniques.
VLMs exhibit a significant performance gap in visual text understanding, with even the best models falling short of human accuracy in error correction tasks.
TetherMem enables video generation models to dynamically adapt scenes while keeping subjects consistent, achieving a significant leap in overall video quality and scene progression.
Storage limitations, not compute power, dictate the scalability of whole-slide image embedding extraction, reshaping our approach to handling large-scale medical data.
Ancient-Bench reveals that even advanced models fail to solve the challenges of recognizing ancient Chinese texts, underscoring a critical gap in AI capabilities.
SpatialCrafter achieves unprecedented 3D consistency in image-to-scene generation, effectively eliminating long-term drift and enhancing detail fidelity.
Aligning object poses with their geometric axes can dramatically enhance estimation accuracy while simplifying model architecture requirements.
Vision generative AI models could revolutionize edge applications, but only if we rethink how they are designed in relation to hardware constraints from the outset.
Achieving camera calibration accuracy that meets the Cramer-Rao Lower Bound using flawed and asynchronous GPS data could revolutionize drone-based imaging systems.
OCR technologies are evolving rapidly, yet significant challenges remain in recognizing diverse scripts and handwritten text, demanding innovative solutions for real-time applications.
Automated segmentation pipelines can effectively track AMD and DME lesions in clinical settings, but their performance drops when generalizing beyond training data.
AnatoProto not only surpasses state-of-the-art models in fetal ultrasound detection but also reveals that combining anatomy-weighted pooling with prototype loss can dramatically enhance recall.
A deep learning approach can correct jitter in phase-contrast micro-CT images without needing a motion-free reference, significantly enhancing image quality.
Calibration in medical vision-language models is crucial, and MVC-Bench reveals that a simple train-time calibration method can outperform existing approaches in most scenarios.
Video-OPSD reveals that focusing on privileged visual evidence can drastically improve the efficiency and effectiveness of self-distillation in Video-LLMs.
Emotion recognition in visual intelligence is heavily skewed towards linguistic cues, revealing a critical gap in purely visual affect recognition capabilities.
Real-time emotion predictions can be made more reliable by selectively invoking multimodal reasoning only when necessary, improving accuracy without sacrificing speed.
RubricRM adapts evaluation criteria dynamically, leading to significant performance gains in visual generative tasks compared to static reward models.
Achieving a 3.09 percentage-point accuracy gain while reducing model size by 80% demonstrates the power of cross-architecture knowledge distillation in precision agriculture.
UniGeo achieves a remarkable 13.59-point improvement in retrieval accuracy for text-guided drone geo-localization by leveraging a unified multimodal framework.
Multi-agent SLAM can now achieve high-quality 3D reconstruction using only RGB and inertial data, making it accessible for consumer-grade devices.
FAN-LoRA outperforms existing methods by explicitly decoupling frequency components, leading to significant enhancements in medical image segmentation accuracy.
By integrating uncertainty modeling into SDF learning, NeuDonatello significantly boosts the accuracy of 3D surface reconstruction from RGB images, even in challenging scenarios.
AGW-PBR achieves superior brain MRI super-resolution by effectively addressing the partial-volume effect, enhancing fidelity in tissue transitions that traditional methods overlook.
Benchmark-driven results can obscure the true complexities of real-world imaging problems, leading to misguided research priorities.
Achieving minutes-long coherence in video generation, Ring Forcing reconciles the trade-off between historical fidelity and generative diversity.
Grounding glass surface detection in 3D geometry leads to state-of-the-art performance and improved scene reconstruction, challenging the limitations of traditional 2D methods.
sLoTh enables continual learning in sparse event-based transformers with less than 1% parameter updates, achieving competitive performance while slashing energy consumption by 6.5x.
Prioritizing content novelty over recency, RECAP-Forcing revolutionizes memory management in long video generation, leading to enhanced visual coherence.
Achieving real-time object discovery without any training, this framework outperforms traditional methods while remaining computationally efficient.
Real-world testing reveals that Assisted Lane Change systems may permit dangerous maneuvers that violate safety distance regulations, challenging current approval processes.
HullWake achieves a remarkable reduction in false positives and improved detection rates for slow or stationary vessels, redefining robustness in maritime detection systems.
DPA-I2P achieves a remarkable 45% reduction in pose estimation errors, setting a new standard for image-to-point cloud registration in autonomous driving.
Current MLLMs falter in following complex video instructions, revealing a significant oversight in their evaluation metrics.
Humans can pinpoint systematic flaws in AI-generated images, revealing critical insights into the limitations of current text-to-image models.
Pointwise convolutions, which dominate parameter volume in large-kernel CNNs, can be drastically reduced through a novel group-sharing strategy, enabling efficient deployment on edge devices.
FRAME reveals that up to 41% of racial performance differences in medical imaging may stem from sampling variation rather than actual bias, challenging conventional fairness assessments.
Cross-attention conditioning dramatically improves the realism of downscaled precipitation forecasts, especially for extreme events, outperforming traditional concatenation methods.
Spatially aware deep learning can extract accurate tropospheric profiles from satellite observations without relying on traditional weather forecasts.
Unsupervised diffusion pretraining can elevate medical image segmentation performance, achieving up to a 74% improvement in boundary precision without relying on extensive labeled datasets.
Achieving over 98% accuracy with a compact model, CropCop sets a new standard for plant-health recognition while ensuring data integrity through rigorous auditing.
Token-level visual dependence is the key to preventing catastrophic forgetting in multimodal continual learning, enabling models to adapt without losing prior knowledge.
TAU-Agent leverages a novel retrieval mechanism to enhance traffic anomaly detection, achieving impressive benchmark results that challenge existing paradigms in video analysis.
Multi-factor difficulty estimation boosts segmentation performance, achieving state-of-the-art results across multiple architectures.
With 36,000 new human annotations, VGA-BenchV2 not only enhances evaluation but also transforms how video generators can be optimized for aesthetic quality and realism.
Achieving state-of-the-art performance with 72% fewer parameters, CrossMambaTuning redefines efficiency in adapting image compression models for machine vision.
PoseOFF captures critical motion cues around human joints, enabling robots to anticipate actions with less data and faster response times.
Fine-grained dataset distillation can achieve superior performance by focusing on localized evidence rather than just global statistics.
Achieving real-time, interactive 4D video generation with precise control over both camera and object movements could revolutionize applications in virtual reality and gaming.
Polarimetric imaging achieves up to 93% accuracy in weld seam segmentation without the need for controlled acquisition, challenging the reliance on traditional RGB methods.
RefineCut's innovative verifier-grounded approach elevates video-editing performance to new heights, achieving a score of 0.924 while eliminating the need for teacher calls at inference.
Identity leakage in face-swapping anonymization is not just a flaw—it's a predictable phenomenon that can be systematically analyzed and improved.
ACT uncovers that only 97 observations drive the predictions for 221 clinical phenotypes, revealing potential shortcuts in medical imaging interpretations.
A groundbreaking dataset and framework that significantly enhance ulcer tissue segmentation accuracy, even with limited labeled data.
Lesion-guided ROI deep learning achieves up to 97.56% accuracy in ovarian ultrasound classification while significantly reducing the annotation workload.
Multi-token supervision can cut training time by 39% while improving image generation quality—an essential leap for scalable synthesis.
RefVideo-6M revolutionizes video editing datasets by providing 5 million high-quality editing samples that prioritize visual references over flawed automatic edits.