Search papers, labs, and topics across Lattice.
Image recognition, object detection, segmentation, video understanding, and visual generation.
#3 of 24
3
UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity and Heterogeneity for all-in-one medical image restoration, is proposed and extensive experiments on two large-scale benchmarks demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration.
This work proposes RDDMPI, a conditional residual diffusion framework that operates directly in residual space, and reformulates probabilistic imputation as a baseline-residual decomposition, where a pretrained model captures the dominant signal and a diffusion process models the residual uncertainty.
Modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask, suggests modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.
This work presents FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder--decoder, achieving state-of-the-art results on major benchmarks, including Sintel, KITTI-2015, and Spring.
Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
The Logit Refiner is introduced, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features and generalizes to text-to-image generation, confirming that the mean-field bottleneck persists across VAR variants and is effectively alleviated by the method.
3D Point Splatting (3DPS), the first differentiable point renderer for radar, derived directly from the standard solid-angle form of the radar equation, is proposed.
This work generates synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer.
Equal parameter count evaluations show that modulating the primitive's wavefront is an effective and lightweight enhancement for hologram representations, and frequency domain analysis illustrates that CVQPG has successfully preserved the mid-to-high frequency band of natural images.
Inspired by Mixture-of-Experts (MoE) principles, Evolvable-Substrate HyperNEAT is partitioned into non-overlapping spatial segments, each assigned to a separately evolved specialist network, which forces evolution to discover features across the entire image.
This paper proposes a meta-learning framework that leverages a comprehensive set of meta-features capturing dataset complexity to predict classifier performance without exhaustive training, and achieves an average ranking prediction accuracy exceeding 86%, demonstrating its effectiveness in guiding model selection.
SolCloudLLM aligns sky-image patches with time-series patches and fuses their corresponding representations through bidirectional multimodal fusion, yielding a unified representation that is subsequently mapped into the embedding space of an LLM.
Control ablations show that gains in cross-regime adaptive propagation are not explained by single-basis propagation, an additional same-family branch, or within-family adaptive order alone, supporting complementary two-basis propagation as an effective design principle for visual representation learning.
Evaluating pre-trained convolutional neural networks for skin lesion classification using dermatoscopic and histopathological image datasets demonstrates that model performance differs between dermatoscopic and histopathological image modalities, showing that architectures exhibiting similar performance on dermatoscopic images exhibit different performance on histopathological data.
A model, InterIL, is proposed, which jointly generates the two modalities, background image and layout, in a single generative process and has no design-specific inductive bias, which allows it to better preserve the original characteristics of realistic designs.
MC-DeTra improves dynamic, socially-situated trajectory forecasting while preserving or improving detection accuracy; a gradient-based loss-calibration analysis exposes how the auxiliary objectives compete at the shared backbone, and the ablation identifies which signals contribute most.
BruNet is a segmentation framework that combines a ViT-based visual encoder with a SAM-based mask decoder that outperforms CNN-based models, state-of-the-art segmentation models, ChatGPT-4o/5-assisted SAM2 zero-shot baselines, and the medical-oriented MedSAM model, demonstrating strong cross-domain generalisation to bruise segmentation.
This paper proposes MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs, including speech audio, text descriptions, and trajectory data, to generate coherent and lifelike motions without requiring aligned multimodal data.
A deep-learning pipeline for enhancing the detection of faint moving objects in optical space situational awareness (SSA) imagery through automated star removal and background reconstruction is presented, thereby providing an effective data-driven preprocessing strategy for faint moving-object detection in optical SSA scenarios.
Mi-Ripple separates periodic lattice artifacts from content-entangled granular texture, then combines selective spectral notching, structure-aware smoothing, and cleaned-reference regeneration to link measurable artifact reduction to visibly cleaner generated images, rather than optimizing a spectral score alone.
Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model, is presented and the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing is explored.
A novel semi-tensor product for third-order tensors under the t-product framework induced by arbitrary invertible linear transforms is introduced, yielding an accelerated multi-term randomized semi-tensor product SVD algorithm that achieves a balance between reconstruction accuracy and computational efficiency.
It is tested whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs against the domain-specific specialist BioCLIP on a 96-species task, and comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.
Calibration-Aware Uncertainty Cascades is proposed, a simple post-hoc framework that independently calibrates each model's confidence and selects deployment policies using validation data and shows theoretically that calibration gives confidence thresholds an explicit selective-risk interpretation, whereas uncalibrated scores offer no comparable reliability guarantee.
This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.
Single-model 3D Gaussian Splatting can match the calibrated uncertainty bounds of a 10-model ensemble at 280 FPS simply by factorizing spatial rendering geometry from view-dependent difficulty under a conformal framework.
Novel classes are rarely entirely novel—treating them as unseen compositions of known visual primitives enables robust open-world category discovery without modifying downstream heads or objective functions.
Autonomous vehicles can maintain millimeter-wave beam alignment even when critical sensors drop out by leveraging generative cross-modal reconstruction across camera, LiDAR, and radar feeds.
Standard open-vocabulary detectors fail on nearly 70% of ambiguous, task-driven queries, but layering structured semantic retrieval and dynamic LLM concept expansion over frozen detectors pushes grounding success to 85%.
Clinical 2D X-rays can now yield high-fidelity 3D volumetric CTs without paired real-world training data, beating prior 2D-to-3D synthesis baselines by up to 14% in PSNR through multi-view slice refinement and progressive transfer learning.
Scaling generalized zero-shot recognition to massive vocabularies does not require re-architecting backbones: treating seen-versus-unseen gating as a downstream Monte Carlo calibration problem boosts unseen accuracy by over 20% in just 15 dimensions.
Unsupervised signal-processing pseudo-labels can outperform ground-truth contact sensors when training deep physiological vision models on imperfectly synchronized video.
Contrastive flow matching can upper-bound forward-KL divergence on fixed preference pairs, resolving the training instability of unbounded DPO surrogates without requiring online rollouts.
High-dimensional small-sample learning can achieve robust classification without parameter-heavy deep networks simply by systematically balancing kernel variance, margin, and compactness.
Unstructured conversational AI fails complex creative handoffs—anchoring revision tasks to a structured intent-evidence-action record prevents criteria drift while keeping human aesthetic authority intact.
Fragmented satellite data no longer means blind spots: learning over incomplete multi-sensor constellations nearly triples usable methane plume events while slashing false positives by 8.19 points.
Unfolding convolutional kernels breaks Muon's optimization geometry; aligning polar updates with the true convolution operator in the frequency domain cuts flow-matching compute by nearly 40% while radically outperforming Adam.
Shrinking a 102M-parameter nnU-Net by 81× costs barely 2.5% in segmentation Dice while actually boosting lesion-level detection F1 by over 5 points.
Neither NeRF nor 3D Gaussian Splatting is globally superior across a complex scene: dynamically swapping teacher-student roles ray-by-ray boosts 3DGS fidelity by 1.56 dB without requiring shared feature spaces or point correspondences.
Relying on arbitrary 2D slice selection for joint pathology is obsolete when implicit neural representations can reconstruct continuous, multi-planar anatomical profiles directly from standard clinical MRIs.
Inverse rendering no longer requires trading physical validity for generative expressiveness: reciprocally conditioned latent bridge matching achieves state-of-the-art albedo estimation across five benchmarks while drastically cutting iterative inference costs.
Re-encoding a compact set of distilled synthetic images during test-time adaptation dynamically realigns source anchors with shifting model weights, eliminating both catastrophic forgetting and pseudo-label drift under severe distribution shifts.
The radius of a hyperbolic embedding acts as a natural uncertainty signal, allowing open-world detectors to reliably isolate unknown object classes without relying on brittle Euclidean decision boundaries.
Instead of forcing entire videos into a single, lossy feed-forward context, training an agent to actively read and write temporal memory yields 78.8% accuracy on MINERVA with virtually zero performance drop on long-duration clips.
Directly conditioning video diffusion on audio wastes massive capacity on static background and identity pixels—routing control transitively through a causal motion latent distilled under a single frozen video teacher achieves real-time streaming at 15.4 FPS with zero fidelity loss.