Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
Inspired by Mixture-of-Experts (MoE) principles, Evolvable-Substrate HyperNEAT is partitioned into non-overlapping spatial segments, each assigned to a separately evolved specialist network, which forces evolution to discover features across the entire image.
High-dimensional small-sample learning can achieve robust classification without parameter-heavy deep networks simply by systematically balancing kernel variance, margin, and compactness.
UFO is proposed, the first unified framework for omni-condition alignment simultaneous evaluation, and UFO-Bench is presented, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.
UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity and Heterogeneity for all-in-one medical image restoration, is proposed and extensive experiments on two large-scale benchmarks demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration.
This work proposes RDDMPI, a conditional residual diffusion framework that operates directly in residual space, and reformulates probabilistic imputation as a baseline-residual decomposition, where a pretrained model captures the dominant signal and a diffusion process models the residual uncertainty.
UFO is proposed, the first unified framework for omni-condition alignment simultaneous evaluation, and UFO-Bench is presented, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.
UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity and Heterogeneity for all-in-one medical image restoration, is proposed and extensive experiments on two large-scale benchmarks demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration.
This work proposes RDDMPI, a conditional residual diffusion framework that operates directly in residual space, and reformulates probabilistic imputation as a baseline-residual decomposition, where a pretrained model captures the dominant signal and a diffusion process models the residual uncertainty.
Modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask, suggests modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.
This work presents FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder--decoder, achieving state-of-the-art results on major benchmarks, including Sintel, KITTI-2015, and Spring.
Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
The Logit Refiner is introduced, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features and generalizes to text-to-image generation, confirming that the mean-field bottleneck persists across VAR variants and is effectively alleviated by the method.
3D Point Splatting (3DPS), the first differentiable point renderer for radar, derived directly from the standard solid-angle form of the radar equation, is proposed.
This work generates synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer.
Equal parameter count evaluations show that modulating the primitive's wavefront is an effective and lightweight enhancement for hologram representations, and frequency domain analysis illustrates that CVQPG has successfully preserved the mid-to-high frequency band of natural images.
Inspired by Mixture-of-Experts (MoE) principles, Evolvable-Substrate HyperNEAT is partitioned into non-overlapping spatial segments, each assigned to a separately evolved specialist network, which forces evolution to discover features across the entire image.
This paper proposes a meta-learning framework that leverages a comprehensive set of meta-features capturing dataset complexity to predict classifier performance without exhaustive training, and achieves an average ranking prediction accuracy exceeding 86%, demonstrating its effectiveness in guiding model selection.
SolCloudLLM aligns sky-image patches with time-series patches and fuses their corresponding representations through bidirectional multimodal fusion, yielding a unified representation that is subsequently mapped into the embedding space of an LLM.
Control ablations show that gains in cross-regime adaptive propagation are not explained by single-basis propagation, an additional same-family branch, or within-family adaptive order alone, supporting complementary two-basis propagation as an effective design principle for visual representation learning.
Evaluating pre-trained convolutional neural networks for skin lesion classification using dermatoscopic and histopathological image datasets demonstrates that model performance differs between dermatoscopic and histopathological image modalities, showing that architectures exhibiting similar performance on dermatoscopic images exhibit different performance on histopathological data.
A model, InterIL, is proposed, which jointly generates the two modalities, background image and layout, in a single generative process and has no design-specific inductive bias, which allows it to better preserve the original characteristics of realistic designs.
MC-DeTra improves dynamic, socially-situated trajectory forecasting while preserving or improving detection accuracy; a gradient-based loss-calibration analysis exposes how the auxiliary objectives compete at the shared backbone, and the ablation identifies which signals contribute most.
BruNet is a segmentation framework that combines a ViT-based visual encoder with a SAM-based mask decoder that outperforms CNN-based models, state-of-the-art segmentation models, ChatGPT-4o/5-assisted SAM2 zero-shot baselines, and the medical-oriented MedSAM model, demonstrating strong cross-domain generalisation to bruise segmentation.
This paper proposes MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs, including speech audio, text descriptions, and trajectory data, to generate coherent and lifelike motions without requiring aligned multimodal data.
A deep-learning pipeline for enhancing the detection of faint moving objects in optical space situational awareness (SSA) imagery through automated star removal and background reconstruction is presented, thereby providing an effective data-driven preprocessing strategy for faint moving-object detection in optical SSA scenarios.
Mi-Ripple separates periodic lattice artifacts from content-entangled granular texture, then combines selective spectral notching, structure-aware smoothing, and cleaned-reference regeneration to link measurable artifact reduction to visibly cleaner generated images, rather than optimizing a spectral score alone.
A novel semi-tensor product for third-order tensors under the t-product framework induced by arbitrary invertible linear transforms is introduced, yielding an accelerated multi-term randomized semi-tensor product SVD algorithm that achieves a balance between reconstruction accuracy and computational efficiency.
It is tested whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs against the domain-specific specialist BioCLIP on a 96-species task, and comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.
Calibration-Aware Uncertainty Cascades is proposed, a simple post-hoc framework that independently calibrates each model's confidence and selects deployment policies using validation data and shows theoretically that calibration gives confidence thresholds an explicit selective-risk interpretation, whereas uncalibrated scores offer no comparable reliability guarantee.
Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model, is presented and the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing is explored.
This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.
Single-model 3D Gaussian Splatting can match the calibrated uncertainty bounds of a 10-model ensemble at 280 FPS simply by factorizing spatial rendering geometry from view-dependent difficulty under a conformal framework.
Novel classes are rarely entirely novel—treating them as unseen compositions of known visual primitives enables robust open-world category discovery without modifying downstream heads or objective functions.
Autonomous vehicles can maintain millimeter-wave beam alignment even when critical sensors drop out by leveraging generative cross-modal reconstruction across camera, LiDAR, and radar feeds.
Standard open-vocabulary detectors fail on nearly 70% of ambiguous, task-driven queries, but layering structured semantic retrieval and dynamic LLM concept expansion over frozen detectors pushes grounding success to 85%.
Clinical 2D X-rays can now yield high-fidelity 3D volumetric CTs without paired real-world training data, beating prior 2D-to-3D synthesis baselines by up to 14% in PSNR through multi-view slice refinement and progressive transfer learning.
Scaling generalized zero-shot recognition to massive vocabularies does not require re-architecting backbones: treating seen-versus-unseen gating as a downstream Monte Carlo calibration problem boosts unseen accuracy by over 20% in just 15 dimensions.
Unsupervised signal-processing pseudo-labels can outperform ground-truth contact sensors when training deep physiological vision models on imperfectly synchronized video.
Contrastive flow matching can upper-bound forward-KL divergence on fixed preference pairs, resolving the training instability of unbounded DPO surrogates without requiring online rollouts.
High-dimensional small-sample learning can achieve robust classification without parameter-heavy deep networks simply by systematically balancing kernel variance, margin, and compactness.
Unstructured conversational AI fails complex creative handoffs—anchoring revision tasks to a structured intent-evidence-action record prevents criteria drift while keeping human aesthetic authority intact.
Fragmented satellite data no longer means blind spots: learning over incomplete multi-sensor constellations nearly triples usable methane plume events while slashing false positives by 8.19 points.
Unfolding convolutional kernels breaks Muon's optimization geometry; aligning polar updates with the true convolution operator in the frequency domain cuts flow-matching compute by nearly 40% while radically outperforming Adam.
Shrinking a 102M-parameter nnU-Net by 81× costs barely 2.5% in segmentation Dice while actually boosting lesion-level detection F1 by over 5 points.
Neither NeRF nor 3D Gaussian Splatting is globally superior across a complex scene: dynamically swapping teacher-student roles ray-by-ray boosts 3DGS fidelity by 1.56 dB without requiring shared feature spaces or point correspondences.
Relying on arbitrary 2D slice selection for joint pathology is obsolete when implicit neural representations can reconstruct continuous, multi-planar anatomical profiles directly from standard clinical MRIs.
Inverse rendering no longer requires trading physical validity for generative expressiveness: reciprocally conditioned latent bridge matching achieves state-of-the-art albedo estimation across five benchmarks while drastically cutting iterative inference costs.
Re-encoding a compact set of distilled synthetic images during test-time adaptation dynamically realigns source anchors with shifting model weights, eliminating both catastrophic forgetting and pseudo-label drift under severe distribution shifts.
The radius of a hyperbolic embedding acts as a natural uncertainty signal, allowing open-world detectors to reliably isolate unknown object classes without relying on brittle Euclidean decision boundaries.
Instead of forcing entire videos into a single, lossy feed-forward context, training an agent to actively read and write temporal memory yields 78.8% accuracy on MINERVA with virtually zero performance drop on long-duration clips.
Directly conditioning video diffusion on audio wastes massive capacity on static background and identity pixels—routing control transitively through a causal motion latent distilled under a single frozen video teacher achieves real-time streaming at 15.4 FPS with zero fidelity loss.
Decoupling spatial features into high and low frequency priors allows a single sparse MoE model to generalize across heterogeneous facial landmark benchmarks without dataset-specific tuning.
High-resolution, multi-view consistent 3D texturing no longer requires costly per-scene optimization or custom fine-tuning: frozen 2D diffusion models can directly synthesize production-ready texture atlases with baked shadows at an 80% speedup.
Forcing state-space models into the primary convolutional path actively harms tiny object detection, but routing selective scanning off-path enables an efficient detector that beats YOLOv8s by 10.8 pp on VisDrone with two-thirds fewer parameters.
Unconstrained sampling offsets routinely degrade dynamic spatial modeling—freezing the kernel's center and filtering deformations with high-frequency edge priors consistently outperforms DCNv4 without adding architectural bloat.
Standard pixel augmentations frequently corrupt delicate vision-language alignment, but injecting diffusion-style isotropic noise directly into embedding spaces breaks through the longstanding performance ceiling of stacked CutMix, Mixup, and RandAug recipes.
Unsupervised PCA paired with Random Forest beats supervised LDA-based pipelines and SVMs in hyperspectral classification, challenging the assumption that class-aware dimensionality reduction is strictly necessary for high-dimensional remote sensing.
3D Gaussian Splatting typically collapses without dozens of views, but anchoring explicit primitives to statistical shape models unlocks robust volumetric reconstruction from just five X-ray projections.
Supervised few-shot adaptation no longer has to break conformal prediction guarantees: aligning 1D nonconformity score distributions restores distribution-free calibration on unlabeled query sets without forcing practitioners to freeze model weights.
Leading MLLMs still struggle to articulate and localize physical changes across revisited scenes, exposing a fundamental blind spot in multi-image spatial grounding that targeted synthetic data can overcome.
Diffusion-generated thermal faces can effectively break the multi-modal data bottleneck in biometrics, outperforming models trained on scarce real pairs without requiring costly image translation at inference time.
Low-quality modalities can quietly sabotage multi-modal networks because they receive minimal training gradients yet dominate test-time skip connections—a failure mode solved by explicitly filtering and rebalancing degraded feature streams.
2D foundation model priors can be transferred directly into coordinate-free 3D point representations via conditional diffusion, forming an emergent multimodal bottleneck that achieves non-trivial 3D semantic understanding without native 3D labels.
Tiny models can conquer severe clinical domain shifts: grounding 3D kinematic sequences in LLM-generated semantic descriptions and merging source-domain weights delivers top-tier Parkinsonian gait severity prediction with just 637K active parameters.
Multi-view diffusion models routinely hallucinate duplicate objects and warp geometry under camera rotation; enforcing spherical-projection attention constraints directly during denoising solves this across complex streetscapes without modifying backbone weights.
Wavelet transforms in deep learning no longer require trading off learnability for mathematical stability: lattice-structured lifting steps guarantee perfect invertibility and stability under arbitrary network parameter updates.
Wearable-free video models can match the precision of physical IMU sensors during severe turning-induced self-occlusion by distilling kinematic topologies directly into visual latent spaces.
Standard Gumbel-Sigmoid pruning fails in 3DGS because it forces binary decisions before importance ranks can stabilize—swapping it for a simple linear activation cuts primitive count by up to 3.6x while actually increasing rendering quality.
Zero-shot composed video retrieval can jump by over 35% absolute R@1 without any parameter updates—provided you route frozen foundation models as a dynamic compute ladder rather than a uniform scoring function.
This work evaluates knee-angle estimation under a missing-ankle-keypoint condition and test a first-order temporal interpolation scheme as a recovery mechanism to recover a critical missing joint, without resorting to learned reconstruction models.
Directly mapping driver facial cues to traffic objects via cross-attention outperforms intermediate point-of-gaze pipelines, slashing background-object misclassification errors by 49.7%.
Human-in-the-loop segmentation breaks down when expert attention is treated as an infinite resource—allocating clinician queries strictly by expected value of information transforms fragile few-shot models into risk-bounded decision systems.
Dense pixel-level segmentation can be performed directly inside an autoregressive token space without external decoders by serializing binary change masks into grammar-constrained quadtrees.
Petabyte-scale brain reconstruction no longer requires massive compute clusters—prioritizing sparse topological skeletons over dense voxel sweeps enables complete neuronal recovery on a single GPU in just one week.
The longstanding performance gap between CNNs and transformers in cell microscopy collapses to under 0.5 macro-F1 once pretraining is controlled, paving the way for distilled tiny students that outperform massive frontier backbones.
Today's specialized video hallucination detectors peak at an abysmal 34.63% accuracy under adversarial conditions, revealing that the verifiers meant to guardrail multimodal foundation models are fundamentally untrustworthy.
TeethGNN is developed, a novel graph-based framework designed to combine CBCT image features with morphological information for accurate and efficient malocclusion grading, providing an accurate and reliable solution for vision-based clinical measurement and diagnosis.
Animal visual re-identification no longer requires shared anatomical priors: decoupling environmental style from dynamic cross-species relational neighborhoods enables unified representation learning across radically different morphologies.
Commodity laptop webcams can infer accurate, distance-invariant 3D gaze by extracting passive depth-from-defocus cues without dedicated RGB-D hardware or external markers.
Accurate void fraction estimation no longer requires flow-disrupting physical probes: multi-view video coupled with spatio-temporal modeling recovers complex multiphase fluid parameters and directly transfers from synthetic CFD to real experimental flows.
Neural motion trackers routinely accumulate drift over repetitive biological cycles; pairing persistent cross-window memory tokens with cyclic teacher-student supervision forces trajectories to close naturally without sacrificing frame-to-frame precision.
Coarse, out-of-date GIS vectors can match dense human annotations when intermediate masks are repurposed as component-wise spatial prompts for targeted local refinement.
Most test-time adaptations in vision-language models are either redundant or actively destructive—and skipping up to 85% of them via simple consistency checks actually boosts final accuracy.
Single-step flow-matching DiTs can surpass multi-step generative models in geometric accuracy, slashing depth error by up to 26% while recovering fine-grained details down to individual strands of hair.
Multimodal models master chart semantics with near-perfect accuracy, yet fundamentally collapse when reasoning about 2.5D visual occlusions—dropping to 61% on layer ordering while generative editors fail under visibility constraints with mIoUs as low as 0.37%.
Real-time object detection and tracking in autonomous racing can be achieved with a multi-modal fusion approach that significantly improves performance under challenging conditions.
Strong adversarial robustness can emerge without generating a single adversarial example during training, reaching 76.6% AutoAttack accuracy on CIFAR-10 purely through oscillatory neural dynamics and predictive self-supervision.
Instead of relying on crude, handcrafted perturbations for radiation therapy safety margins, generative latent velocity fields can now simulate realistic, continuous 3D patient anatomical deformations at full clinical CT scale.