Search papers, labs, and topics across Lattice.
100 papers published across 2 labs.
Achieving real-time 3D hand pose estimation without the need for camera parameters could revolutionize applications in AR and robotics.
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
TLive-Omni achieves superior live-commerce understanding by seamlessly integrating multi-modal inputs and optimizing for real-time response quality.
EviRank redefines multimodal image re-ranking by transforming complex queries into structured evidence, achieving unprecedented accuracy and efficiency.
InfinityEdit allows for real-time, unbounded video editing that adapts to live streams, maintaining quality and coherence despite multiple edits.
TLive-Omni achieves superior live-commerce understanding by seamlessly integrating multi-modal inputs and optimizing for real-time response quality.
EviRank redefines multimodal image re-ranking by transforming complex queries into structured evidence, achieving unprecedented accuracy and efficiency.
InfinityEdit allows for real-time, unbounded video editing that adapts to live streams, maintaining quality and coherence despite multiple edits.
Explicit modular spatial verification can dramatically enhance diagnostic accuracy in medical imaging, outperforming traditional end-to-end models by a staggering margin.
A single architecture can achieve state-of-the-art results in pelvic imaging across multiple modalities, addressing critical data integrity issues in women's health research.
Flow matching can revolutionize PET image reconstruction by achieving better bias-variance trade-offs with fewer sampling steps than traditional methods.
Achieving real-time 3D hand pose estimation without the need for camera parameters could revolutionize applications in AR and robotics.
Learning assembly dependencies can drastically improve robotic manipulation, outperforming traditional object-centric approaches in complex tasks.
Existing methods falter on a new benchmark that captures the full complexity of basketball game events, revealing critical gaps in current visual understanding approaches.
Mix&Fix-Net bridges the monitoring gap for small vessels by integrating AIS and vision data, achieving superior trajectory prediction accuracy.
Freezing the backbone and optimizing inference strategies led to a surprising 0.041 APD improvement without any training, challenging conventional wisdom about fine-tuning.
Structured affinity enables Deep AINs to learn visual memory without replay, achieving impressive accuracy while retaining earlier class information.
Efficiently segmenting complex 3D curvilinear structures could revolutionize medical imaging by enhancing the accuracy of anatomical representations.
Optimizing feature engineering can double model accuracy and dramatically reduce error rates in ocean colour machine learning, but a one-size-fits-all approach won't work.
Training on geographically isolated samples can cut costs by over 140x while boosting model performance across multiple Earth observation tasks.
A novel reinforcement learning reward derived from positive pairs alone boosts keypoint learning stability and performance, even in low-texture environments.
Achieving unprecedented accuracy in 3D object detection by refining both object discovery and model training through innovative dual-guidance techniques.
Achieving unprecedented accuracy in biventricular motion synthesis from a single ED mesh could transform cardiac modeling and analysis.
Subtle timing in subtitle presentation can amplify video jailbreak effectiveness against LVLMs, achieving unprecedented attack success rates.
Core-KAN achieves superior performance in vision tasks by synthesizing spatial filters at arbitrary resolutions without the computational burden of per-location kernel generation.
Achieving high-accuracy disease detection without ever exposing sensitive patient data could revolutionize collaborative medical research.
Achieving a 4% accuracy boost in CNNs without changing datapath width while enhancing energy efficiency by 2.5x could revolutionize industrial visual inspection.
Combining gradient obfuscation with image encryption can drastically reduce data leakage risks in federated learning without sacrificing model accuracy.
Achieving a 0.499 face similarity score, WithEveryone revolutionizes group image generation by ensuring identity preservation for up to ten individuals without direct face copying.
IR-UWB outperforms other radar technologies in activity recognition, but FMCW shines in adapting to new environments—revealing a crucial trade-off for healthcare applications.
Streetscape qualities crucial for pedestrian-friendly urban design are found to be alarmingly scarce in suburban areas, revealing a significant gap in urban planning.
Integrating prompt-conditioned channel attention can boost segmentation accuracy by over 23% in challenging medical imaging tasks.
Real-world tennis serving by humanoid robots is now possible without motion capture, thanks to a novel adaptive framework that learns directly from video.
DARS achieves superior performance in instruction-based image editing by transforming outcome-level feedback into actionable, localized supervision for both planning and rendering stages.
A closed-form conditional-Gaussian estimator outperforms deep learning approaches in reconstructing complex cardiac shapes, achieving unprecedented accuracy in shape completion.
SEFS achieves superior artistic stylization by leveraging low-resolution image crops, enhancing content consistency while avoiding unwanted style transfer artifacts.
Under extreme visual degradation, adaptive fusion of sonar and visual data boosts underwater detection accuracy by over 33%, revealing the critical role of modality reliability.
Event memory enables real-time soccer commentary that adapts seamlessly to the evolving context of live matches, outperforming traditional methods.
Jointly leveraging selection-based and reinforcement-learning-based alignment can transform medical image captioning into a more trustworthy diagnostic tool.
CalcSeg achieves remarkable improvements in myocardial scar segmentation by intelligently prioritizing difficult cases, outperforming existing methods even in low-contrast environments.
DPC-Net achieves unprecedented image restoration quality by seamlessly integrating semantic understanding with low-level visual cues, outperforming existing methods on multiple benchmarks.
Eyeglasses removal from video can now be achieved with unprecedented fidelity and temporal stability, thanks to a physics-grounded approach that preserves identity and expression.
Achieving 28.5% better accuracy with 161 fewer primitives, this method revolutionizes sparse view 3D reconstruction efficiency.
Achieving leading performance in image generation with only 6 billion parameters, Swift-Image redefines the efficiency frontier for compact models.
DreamHand achieves a groundbreaking 40% reduction in error for 3D hand trajectory recovery in occluded environments, setting a new standard for egocentric video analysis.
G3Ego reveals that integrating gaze into graph construction can significantly enhance the efficiency and accuracy of egocentric action recognition.
A single AI model can outperform specialized counterparts in recognizing surgical phases across multiple centers, challenging the need for tailored approaches.
UPAL achieves a remarkable 4x speedup and 10x smaller memory footprint while maintaining state-of-the-art performance in multi-view feature extraction.
Micro-drones can autonomously navigate hazardous environments without GPS, preserving vital sensor data even when communication is lost.
MUST-PET achieves superior lesion segmentation and reconstruction accuracy, even with limited labeled data, by harnessing the power of multimodal self-supervised learning across diverse PET-CT scans.
By combining images with text queries, ID-VTG significantly improves the accuracy of video grounding in scenarios with visually similar entities.
Achieving accurate pose estimation with just one or two feature correspondences could revolutionize localization in consumer devices.
STEP transforms pose anomaly detection by ensuring that noise injection leads to physically plausible variations, enabling robust performance on longer video sequences.
AutoLumNet achieves state-of-the-art exposure correction by uniting monotonicity, optimal transport, and local adaptivity in a single trainable model.
State-of-the-art video object removal methods achieve high visual fidelity but systematically fail to maintain causal consistency in real scenes.
Injecting diffusion representations into CLIP pipelines boosts CZSL performance, revealing the untapped potential of generative models in zero-shot learning tasks.
S$^2$GS slashes per-frame optimization time and storage costs while delivering high-quality Free-Viewpoint Videos, making immersive IoT applications feasible on edge devices.
Transforming static avatars into dynamic, realistic representations could redefine the standards for avatar realism in virtual environments.
Achieving near fully supervised accuracy in tumor segmentation using only image-level labels, even in the presence of client-specific missing modalities, is a game changer for federated learning in healthcare.
Achieving a Dice score of 0.6796, AsymFeX outperforms existing methods by effectively utilizing brain symmetry for accurate stroke lesion segmentation across imaging modalities.
TextRefine achieves superior text editing in product posters by ensuring high fidelity and optimal placement, overcoming common pitfalls of existing models.
CVSD-Reg achieves a remarkable 97.7% success rate on challenging LiDAR datasets, outperforming existing methods by up to 44% without relying on camera data.
LFPR boosts bounding box accuracy by over 3% on average without any target annotations, revealing a complex interplay between referent selection and boundary precision.
Exciting specific molecular vibrations can double the yield of proton transfers in single-benzene fluorophores, revealing a new dimension of control in ultrafast photochemical processes.
Achieving high-quality 4D human reconstruction from casual monocular videos could revolutionize applications in virtual reality and gaming.
Transferable presentation attack representations can be learned without relying on facial content, achieving high performance across standard benchmarks.
Achieving a $5.15\times$ speedup in text-to-3D generation without sacrificing quality could revolutionize real-time 3D content creation.
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
Stream4D transforms video generation by enabling coherent motion and dynamic scene representation, outperforming static critics that hinder realism.
Achieving a 25.3% relative error reduction in character error rates, Phoenix redefines manuscript transcription as an auditable evidence management task rather than mere text replacement.
SCORE achieves a remarkable 53.23% Top-1 accuracy in EEG-to-image retrieval, outperforming existing methods by over 17 percentage points, even without target labels.
MDTIM not only separates missing from observed values but also directly predicts original time series values, leading to superior imputation performance.
Super-resolution GANs can accelerate EBSD analysis by 25x without sacrificing critical microstructural accuracy, revolutionizing battery material characterization.
Denoising in diffusion models acts as a dynamical Bayesian classifier, revealing that posterior probabilities can focus on a single cluster under specific conditions.
Fuzzy accuracy reveals that skin tone classification from PPG signals can achieve up to 96% accuracy, challenging traditional metrics that overlook label subjectivity.
Colorist outperforms complex generative models by safely generating clinically relevant domain variations while preserving anatomical structures.
Achieving competitive skin disease classification while ensuring fairness across skin tones, MIFR aligns clinical and dermoscopic data in a shared embedding space.
SelF-Rocket outperforms existing methods in fault classification while optimizing for both accuracy and computational efficiency.
CutMix boosts the reliability of semantic segmentation models under distribution shifts, even if it doesn't significantly improve accuracy.
A modular risk modeling framework can adapt to new environmental data, enabling utilities to prioritize inspections and interventions effectively.
Backdoor detection just got a major upgrade—DistScan identifies attacks by analyzing shifts in prediction distributions, achieving a 27.32% accuracy boost over existing methods.
One-stage object detectors can achieve remarkable speed and efficiency, but a critical gap remains between their benchmark performance and real-world reliability in autonomous driving.
Counterfactual images generated without reliance on specific classifiers can significantly reduce bias and improve interpretability in medical imaging tasks.
Tailoring data augmentation to individual learning states boosts performance by an average of 4.5% on natural images, revealing a new frontier in generative data strategies.
Iterative fine-tuning of OCR can drastically cut down the time and expertise needed for transcribing complex historical manuscripts.
Achieving competitive medical image segmentation with just 10 annotated cases could revolutionize the efficiency of clinical workflows.
Training prior and feature representation can significantly enhance the efficacy of AI-generated image detectors, often outperforming established methods.
Achieving over 2.5% better segmentation performance than existing methods, OptiModNet does so with a fraction of the computational cost.
Interactive correction of fish tracking predictions via natural language guidance shows promise, yet highlights the need for better integration of user input.
AgriNav achieves over 90% confidence in crop row detection even during GNSS outages, revolutionizing precision agriculture with minimal herbicide reliance.
Unlocking 22.6 million visual elements from historical books could revolutionize how we engage with digitized library collections and their applications in AI and research.
EgoHRV reveals that gaze video can be a powerful tool for continuous heart rate variability estimation, unlocking new dimensions in behavioral analysis.
Teeth2Point achieves a 1.44 DSC point improvement in segmentation accuracy for challenging dental cases, showcasing a breakthrough in handling missing or misaligned teeth.
Achieving state-of-the-art performance in lightweight semantic segmentation, SiConMo reveals that simplicity in design can outperform complex architectures.
RVLoss achieves a 20% performance boost in self-supervised LiDAR scene flow estimation by enforcing motion rigidity through a novel runoff voting mechanism.
Scattering-aware feature decomposition boosts few-shot SAR object detection performance by effectively leveraging sensor-specific characteristics.
CL4D redefines vision-language interaction by achieving state-of-the-art results in dynamic scene understanding without relying on traditional 2D representations.
Existing video quality metrics fall short for camera-controlled generation, but CWQA sets a new standard by accurately predicting perceptual quality with a tailored approach.
Achieving a tenfold reduction in reconstruction time for fetal cardiac MRI could revolutionize clinical workflows and diagnostic capabilities.
Current vision-language models excel at recognizing objects but falter in capturing dynamic interactions and user intent over time, revealing critical gaps in embodied AI.
CDGP achieves state-of-the-art anomaly localization without any human pixel annotations, revolutionizing the efficiency of industrial visual inspections.
Long-term video segmentation just got a boost—SAM2Dual raises performance by leveraging a novel Dual Memory approach that adapts to both immediate and historical context.
SIFT features can outperform some of the latest deep-learning image matching methods in UAV visual odometry, challenging the assumption that newer always means better.
PALATE enables personalized portrait retouching at a fraction of the cost, requiring only 512 bytes of user-specific data while achieving superior preference prediction accuracy.
ReX-Shot achieves seamless control over viewpoint, focal length, and photographic effects from a single image, outperforming existing methods in both quality and speed.