Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
Fusing lower-face video with upper-face EMG boosts emotion recognition accuracy in VR, achieving a 51% macro-F1 score despite facial occlusion.
Human trajectory logs are no longer the performance ceiling for autonomous driving: closed-loop reinforcement learning paired with distilled foundation models outperforms human demonstration baselines across major open and closed-loop benchmarks.
Frontier multimodal models that master complex single-image visual reasoning collapse to sub-35% accuracy when asked to detect basic low-level noise and texture differences between two images.
Srijika is presented, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia, and a negative-results catalogue covering failed conditioning, objective choices, and data-hull limits of reference-guided restyling.
Monolithic 3D scenes with hundreds of heavily occluded objects can be cleanly parsed into individual editable meshes without ever training on multi-object data.
Human trajectory logs are no longer the performance ceiling for autonomous driving: closed-loop reinforcement learning paired with distilled foundation models outperforms human demonstration baselines across major open and closed-loop benchmarks.
Frontier multimodal models that master complex single-image visual reasoning collapse to sub-35% accuracy when asked to detect basic low-level noise and texture differences between two images.
Srijika is presented, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia, and a negative-results catalogue covering failed conditioning, objective choices, and data-hull limits of reference-guided restyling.
Monolithic 3D scenes with hundreds of heavily occluded objects can be cleanly parsed into individual editable meshes without ever training on multi-object data.
Motion generation is no longer bound to fixed human templates: a single diffusion model can now animate arbitrary skeletal topologies—from serpents and insects to bipeds—directly from text in zero-shot fashion.
Agentic 3D scene generation can match state-of-the-art layout fidelity with a 24x speedup by replacing global iterative placement with localized VLM evolution over parametric image priors.
Neural rendering no longer requires per-scene fitting or specialized shaders to capture complex transport effects like caustics and participating media across arbitrary geometries.
High-fidelity, feed-forward 3D editing no longer requires paired 3D training data or agonizing test-time optimization: regularizing 2D visual and VLM distillation with a 3D latent distribution matching objective is enough to prevent geometric collapse.
Video diffusion models fail at few-step camera control not because of capacity limits, but because discretization errors bend the sampling trajectory—a failure mode resolved here to unlock 25× faster novel-view synthesis.
RARF achieves competitive 3D brain MRI inpainting by intelligently focusing on pathological regions while maintaining anatomical integrity, setting a new standard for medical image reconstruction.
A simple post-processing technique boosts SSIM in brain-MRI inpainting, enhancing image quality without retraining the ensemble model.
Causal optimal transport reveals a new pathway for guiding degenerate diffusion models, even when traditional score functions fail.
Natural language descriptions of dataset differences can reveal critical insights into the safety and robustness of autonomous driving systems.
TESSERA embeddings outperform traditional spectral-temporal features in tree species classification, especially when training data is scarce.
Visually grounded linguistic guidance transforms dense video captioning by dynamically adapting to semantic changes, leading to superior event localization.
A novel auditing framework reveals that image editors can satisfy regional plausibility constraints while still failing to provide a coherent global explanation.
Catalogue photography can significantly aid in the cold-start problem for carbide burr recognition, but its effectiveness in real-world applications is limited without targeted adjustments.
GazeFS reduces gaze trajectory errors by up to 0.400 degrees while maintaining temporal smoothness, revolutionizing target-centered gaze interaction.
Leveraging phase information can dramatically improve few-shot image classification, outperforming traditional methods by integrating frequency insights.
BioCLIP2 outperforms generic CLIP models by over 45% in recognizing Bangladeshi freshwater fish, revealing the critical role of nomenclature and context sensitivity in zero-shot learning.
Achieving rich 3D structures without sacrificing runtime, this framework redefines the limits of Bundle Adjustment in computer vision.
DropClick can maintain high segmentation performance with minimal user input, saving over 46% of annotation effort while still achieving competitive results.
Achieving mask-free video virtual try-on is now possible with BooM-VVT, which significantly reduces reliance on costly video-level data while enhancing garment realism.
SV-WAM achieves high-performance autonomous driving planning without the computational burden of generating future videos during inference.
Overcoming the 2D-to-4D spatial bottleneck in VLMs does not require native 3D architectures; factorizing visual projections into verifiable planar, depth, and temporal RL objectives delivers immediate 4-6% benchmark gains.
Fusing lower-face video with upper-face EMG boosts emotion recognition accuracy in VR, achieving a 51% macro-F1 score despite facial occlusion.
OCR-EDR transforms OCR error analysis into actionable repairs, achieving a remarkable 94.78% diagnostic accuracy and a 30.99-point boost in formula performance.
ToPO achieves superior image generation performance by optimizing token-conditioned preferences, outperforming existing methods across multiple evaluation metrics.
DeepSSIM++ reveals that you can achieve a 46-point boost in memorization detection accuracy without sacrificing computational efficiency in medical generative models.
Back mark-based tracking boosts individual pig monitoring accuracy by over 9% in challenging environments where traditional methods fail.
Text2Thermal synthesizes thermal images directly from text, achieving state-of-the-art results while eliminating the need for RGB image registration.
STARS-GS boosts aerial surface reconstruction accuracy by over 9% through innovative structure-aware techniques that tackle common pitfalls in existing methods.
Point-based neural editing can now adapt to large deformations without ground truth images, achieving superior consistency and fewer artifacts.
FoRIS redefines in-context segmentation by transforming it into a progressive refinement process, achieving state-of-the-art results without the need for training.
Counting unique animals in camera trap sequences is now possible without any count labels, thanks to a novel heuristic that leverages existing detection models.
Automated weld seam recognition could drastically cut down scanning time and data overload in robotic post-processing.
Achieving higher novel-view quality in 3D scene generation without requiring dense volumetric representations or ground-truth supervision is a game changer for practical applications.
CloudCast v2 not only extends cloud-cover forecasting from hours to 12 hours but also enhances spatial accuracy, outperforming previous models in real-time applications.
LeanGRPO achieves up to 1.83x speedup in diffusion RL without sacrificing optimization quality by eliminating redundant computations.
SurgeGen can generate diverse and realistic storm surge scenarios, offering a computationally efficient alternative to traditional physics-based models.
Beyond Intuition outperformed other methods in explaining model behavior, yet its correlation with heart rate accuracy was surprisingly weak, challenging conventional assumptions about explainability in AI.
REMIND can identify and correct noisy annotations in infant pose estimation, achieving up to 93% accuracy in real clinical settings.
Unseen sign classes can be recognized with over 21% accuracy using a reverse sign language dictionary that requires no gloss supervision, challenging the limitations of traditional classification methods.
Preprocessing defenses fail to recover adversarially perturbed inputs in depthwise-separable CNNs, but they reveal a new detection opportunity through measurable output divergence.
VideoLMs can significantly enhance their temporal understanding by leveraging a counterfactual approach that reveals hidden dynamics in video data.
OctWorld achieves unprecedented long-range video generation with spatial consistency, outperforming traditional methods by leveraging an innovative 3D memory architecture.
Achieving nearly full semantic accuracy with just 4.7% of the computational resources challenges the norms of dense segmentation in AI vision systems.
A groundbreaking dataset and benchmark that bridges the gap between anatomical and metabolic analysis in whole-body PET/CT imaging, enabling advanced multimodal reasoning.
By harnessing scene geometry, this method achieves significantly improved reflection consistency in mirror inpainting, outperforming conventional generative techniques.
FlexibleFusion's innovative approach allows for seamless object detection even when critical sensor data is missing, transforming how we handle multimodal perception in real-world applications.
Achieving up to 96% accuracy in parking space classification with minimal annotated data could revolutionize intelligent transportation systems.
Achieving 95.1% boundary localization accuracy, ARCOS revolutionizes corneal layer segmentation across diverse OCT devices without requiring retraining.
Small autoregressive camera motions can drastically enhance the stability of novel view synthesis, preventing geometric distortion and error accumulation.
Deepfake video detection can be revolutionized by a framework that preserves spatial and temporal cues independently, leading to superior adaptation to new forgery patterns.
Geometry-driven modeling in NC-TFAD enables robust anomaly detection in dynamic industrial environments, outperforming conventional methods.
OmniRSCLIP achieves strong performance across heterogeneous remote sensing modalities while preserving the advantages of RGB-based models.
PointGT allows for seamless editing of 3D geometry and texture, overcoming the limitations of traditional volumetric methods.
Restoration processes can inadvertently suppress defect detection, but SafeRestore offers a robust framework to audibly assess when to trust automated image transformations.
MudraGen generates culturally rich and anatomically precise two-hand dance gestures, setting a new standard in the preservation of Indian classical dance.
RGB-only salient object detection can outperform traditional RGB-D methods by leveraging reliable geometry distillation without using depth data during inference.
Lightweight conditional energies can dramatically enhance the observation-awareness of pretrained INRs without the need for retraining.
Achieving superior video compression performance by addressing misalignment issues in complex motion scenarios could redefine standards in neural video coding.
TruncGradGS tackles the gradient vanishing problem in 3D Gaussian Splatting, leading to superior scene reconstructions across various initialization methods.
Tree-VQ enables progressive image decoding where each prefix of the compressed bitstream is immediately usable, significantly enhancing efficiency and user experience.
CoFiE achieves a remarkable 2.54 times improvement in end-to-end inference latency while maintaining high accuracy in streaming video understanding.
Identity preservation in generative models isn't just a byproduct of model complexity; it can be systematically enhanced through a dedicated persistent identity layer.
Reconfiguring attention in frozen Vision Foundation Models can dramatically enhance their ability to detect anomalies in industrial settings, achieving superior localization without retraining.
Eliminating the need for per-scene optimization, this method boosts 3D model performance while enhancing visual fidelity and spatial consistency.
Semantic-Aware Subgraph State Space Model achieves unprecedented accuracy in histopathology classification by preserving the spatial organization of tissue structures.
VI3 achieves accurate metric scale recovery for 3D models using only inertial data, eliminating the need for ground-truth supervision.
Targeted modulation of just eight dominant channels can dramatically enhance image super-resolution quality without fine-tuning the entire model.
DSAQuant reveals that aligning quantization training with the stage-wise nature of video diffusion can drastically enhance visual fidelity in text-to-video generation.
Z3D reveals that 3D Foundation Models can effectively decode hidden surfaces to generate accurate depth maps for unseen views, transforming our approach to 3D reconstruction.
Dusty conditions can severely degrade sensor performance, but this innovative testing environment allows for controlled, reproducible evaluations that could transform agricultural automation.
Self-supervised learning can achieve state-of-the-art video tracking accuracy without any labeled data or additional inference costs.
TokenMatch achieves state-of-the-art performance in 3D shape correspondence estimation by leveraging curvature-guided tokenization, enabling robust matching under challenging conditions.
Online 3D reconstruction failure on long videos is not a representation collapse but an artifact of single-anchor pose extrapolation—and querying relative poses across multiple keyframes fixes it with just 1% parameter overhead.
Halving frame resolution costs virtually nothing in long-video MLLMs, but reinvesting those saved visual tokens to double temporal frame counts yields an immediate 2–3 point accuracy gain—all while a 30-year-old sparse approximation algorithm matches state-of-the-art custom selectors.
Monolithic video evaluation dilutes critical action cues, but chunking interactive rollouts into action-aligned visual evidence allows targeted reward models to outperform GPT-5.5 at scoring world model dynamics.
Visual design no longer requires choosing between diffusion fidelity and code editability: orchestrating modular asset diffusion through a VLM-driven HTML/CSS coding loop delivers fully interactive, layer-decoupled graphic layouts.
Streaming video MLLMs perform far better when historical context is actively internalized into evolving latent tokens rather than queried as passive external visual buffers.
High-fidelity image synthesis does not require paired text from day one: pre-training visual priors on uncaptioned images before multimodal alignment beats conventional joint training pipelines to establish a new open-source DiT benchmark.
World models no longer require fragile offline reconstruction pipelines—baking native physics, depth, and camera pose directly into a unified multimodal generative process unlocks self-calibrating, closed-loop 3D spatial simulation at scale.
Today's top video generators score near 0.8 on standard benchmarks like VBench, but none surpass 0.42 when tested on whether two co-occurring objects obey the same basic laws of physics.
Training-free video editing no longer requires fragmented, task-specific pipelines: unified attention-space control shatters prior zero-shot accuracy records on FiVE by nearly 20 points across both instruction- and subject-driven edits.
Visual place recognition models often crumble under weather and lighting shifts, but injecting 160K geometrically verified, route-aware synthetic hard positives boosts R@1 retrieval by up to 9.2% across foundation backbones.
Foundation models like SAM consistently mistake statues, paintings, and reflections for real-world entities—a fatal flaw for 3D reconstruction that is resolved here by routing ambiguous visual matches through a conditional VLM verification layer.
It is found that synthetic-only training can be competitive, and the 0.9B-parameter PaddleOCR-VL-1.6 is adapted into Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages, showing that synthetic-only training can be competitive.
SBOCF slashes the number of required simulator evaluations while achieving unprecedented accuracy in estimating complex materials parameters from experimental images.
Genesis achieves seamless multi-resolution satellite imagery synthesis, unifying super-resolution and outpainting for globally consistent quadtree generation.
TrajMind achieves a remarkable 41.1% reduction in latency while delivering over 15 percentage points improvement in anomaly diagnosis accuracy.
Visual Hallucinations in large vision-language models can be significantly reduced without the overhead of additional training or curated datasets, thanks to RVSD's innovative approach.
AVIS recovers nearly 70% of accuracy loss from quantization while ensuring real-time performance on lunar rovers, even under radiation constraints.
Disease distribution shifts, not skin tone, are the primary culprits behind the poor performance of dermatology AI models in unfamiliar clinical settings.
Trajectory geometry can be harnessed to dramatically enhance diffusion model sampling efficiency without the need for retraining, yielding significant improvements in output quality.
UGC-enhanced images can harbor subtle anomalies that existing quality assessment methods completely overlook, but our new framework identifies these issues with unprecedented precision.
VoRTeC slashes bit consumption by 58% while boosting decoding speeds up to 197 times, revolutionizing real-time video compression.
RouteGraph-Mona achieves superior mineral classification accuracy by dynamically adapting routing based on sample-specific scale preferences, addressing a critical challenge in geological imaging.
Saliency maps can mislead when images are rotated, but EquiGrad-CAM reveals how to achieve consistent explanations without retraining models.
Unstable appearance variations can lead to unreliable pseudo labels, but SAUF-Net's innovative structure-aware approach dramatically enhances segmentation accuracy in low-label settings.