Search papers, labs, and topics across Lattice.
Training data quality, synthetic data generation, data filtering, deduplication, and dataset construction.
#14 of 24
3
The transition from file storage to database management systems transformed stored data into managed resources into AI model management systems, and AI now faces an analogous transition from AI model storage to AI model management.
This work addresses the adverse interactions between sparsity and data repetition, and presents evidence for the core mechanisms of overfitting and its potential remediation, and suggests promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
This work generates synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer.
LoaDiff is introduced, a diffusion-based generative model for year-long, sub-hourly smart-meter load curves that generates realistic and diverse load profiles, achieves a favorable trade-off between generation quality and limited evidence of memorization, and preserves information useful for downstream energy applications.
This study examines how synthetic and real historical training data should be combined for low-resource OCR, and finds that joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines.
Within the tested range, concentration sets neither the destination nor the pace of collapse; the pace follows whose text fills the pool; what emerges instead is an invariance.
This work constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks and designed an evaluation framework featuring diverse tasks and tailored metrics.
This work systematically reviews FSL approaches for NIDS published from 2022 to 2026 with PRISMA 2020-like reporting to search ACM Digital Library, IEEE Xplore, and Scopus, and compares reported performance.
LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories, enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development.
This work introduces AmazonSWE, a dataset for training and evaluating large-scale spatiotemporal graph imputation methods that integrates processed satellite altimetry measurements from a range of sources, including the recent wide-swath SWOT sensor.
The results show that differentially private perturbation can be integrated into EEG processing workflows, but the selected mechanism, privacy parameters, and sensitivity calibration strongly influence data utility.
Using one instrument and period, a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom is held, governed by true-positive prevalence in each pool via Bayes, not solely by detector quality.
Tightening top-p, which cuts the low-probability tail at generation time, nearly stops collapse within three generations and stabilizes six checkpoints spanning the whole spectrum together, while data-side filtering slows collapse without stopping it.
The Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages, is detailed, which involves multilingual Multiple-Choice Question Answering across 21 languages.
It is shown the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, indicating that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing.
Post-cutoff test sets only prevent verbatim window memorization, revealing that time-series foundation model outperformance is driven by pretraining corpus familiarity rather than generalized temporal reasoning.
Winning lottery tickets at up to 95% sparsity can be harvested essentially for free simply by piggybacking iterative magnitude pruning onto the natural retraining loop of active learning.
Even "leak-free" protein interaction benchmarks are riddled with non-biological shortcuts that models readily exploit—and standard negative-sampling heuristics often make the bias worse.
Treating automated feature transformations as permutation-invariant hierarchies rather than ordered sequences eliminates a fundamental representation bias, enabling policy-guided RL to efficiently navigate non-convex search spaces.
Compact transformers trained purely on synthetic clinical text can match or exceed real-data de-identification performance while proving more robust to format perturbations across out-of-distribution primary care domains.
Full rationale supervision burns up to 254× more tokens without guaranteeing reliability, whereas isolating representations with high rationale-boundary sensitivity selectively immunizes medical LLMs against choice-order brittleness.
Open-weight embedding models fine-tuned via contrastive ranking beat frontier LLMs like GPT-5.2 and Claude-Sonnet-4.6 by over 10 percentage points on real-world entity disambiguation.
Industrial warehouse safety exposes a massive capability gap in frontier LMMs, where even top proprietary models consistently fail to jointly ground visual hazards and reason over operational risks.
Romanized and regional dialects represent a massive evaluation blind spot for frontier LLMs—even across globally dominant low-resource languages like Bangla.
Standard SFT degrades financial reasoning benchmarks by up to 4 percentage points, but pairing self-distillation with rule-verified GRPO transforms domain post-training from a capability tax into a 3-point gain.
Discarding "easy" examples where the base model already follows glossary constraints yields an 11-point accuracy leap at fixed data volume, outperforming RL-based alignment methods with standard supervised fine-tuning alone.
Standard pixel augmentations frequently corrupt delicate vision-language alignment, but injecting diffusion-style isotropic noise directly into embedding spaces breaks through the longstanding performance ceiling of stacked CutMix, Mixup, and RandAug recipes.
Diffusion-generated thermal faces can effectively break the multi-modal data bottleneck in biometrics, outperforming models trained on scarce real pairs without requiring costly image translation at inference time.
Curriculum learning's edge on hard examples is not just an artifact of data exposure: modeling curricula as Wasserstein transport paths reveals that ordering genuinely drives capability shifts, though no universal pacing strategy dominates across tasks.
Counterpart data sharing silently distorts production A/B tests, but selectively filtering bid and ranking divergences outperforms both naive log-sharing and data-starved log-splitting.
Differentiable cardiovascular physics can filter out ICU sensor artifacts that fool standard deep nets, yielding a 5-point SOTA leap on VT alarm reduction alongside a 2x boost in label efficiency.
Open-set GNNs silently break when graphs violate homophily, but synthesizing pseudo-unknown proxies along cross-class displacement vectors restores robust out-of-distribution detection on heterophilic topologies.