Search papers, labs, and topics across Lattice.
91 papers published across 5 labs.
The transition from file storage to database management systems transformed stored data into managed resources into AI model management systems, and AI now faces an analogous transition from AI model storage to AI model management.
This work addresses the adverse interactions between sparsity and data repetition, and presents evidence for the core mechanisms of overfitting and its potential remediation, and suggests promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
This work generates synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer.
LoaDiff is introduced, a diffusion-based generative model for year-long, sub-hourly smart-meter load curves that generates realistic and diverse load profiles, achieves a favorable trade-off between generation quality and limited evidence of memorization, and preserves information useful for downstream energy applications.
This study examines how synthetic and real historical training data should be combined for low-resource OCR, and finds that joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines.
The transition from file storage to database management systems transformed stored data into managed resources into AI model management systems, and AI now faces an analogous transition from AI model storage to AI model management.
This work addresses the adverse interactions between sparsity and data repetition, and presents evidence for the core mechanisms of overfitting and its potential remediation, and suggests promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
This work generates synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer.
LoaDiff is introduced, a diffusion-based generative model for year-long, sub-hourly smart-meter load curves that generates realistic and diverse load profiles, achieves a favorable trade-off between generation quality and limited evidence of memorization, and preserves information useful for downstream energy applications.
This study examines how synthetic and real historical training data should be combined for low-resource OCR, and finds that joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines.
Within the tested range, concentration sets neither the destination nor the pace of collapse; the pace follows whose text fills the pool; what emerges instead is an invariance.
This work constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks and designed an evaluation framework featuring diverse tasks and tailored metrics.
This work systematically reviews FSL approaches for NIDS published from 2022 to 2026 with PRISMA 2020-like reporting to search ACM Digital Library, IEEE Xplore, and Scopus, and compares reported performance.
LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories, enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development.
This work introduces AmazonSWE, a dataset for training and evaluating large-scale spatiotemporal graph imputation methods that integrates processed satellite altimetry measurements from a range of sources, including the recent wide-swath SWOT sensor.
The results show that differentially private perturbation can be integrated into EEG processing workflows, but the selected mechanism, privacy parameters, and sensitivity calibration strongly influence data utility.
Using one instrument and period, a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom is held, governed by true-positive prevalence in each pool via Bayes, not solely by detector quality.
Tightening top-p, which cuts the low-probability tail at generation time, nearly stops collapse within three generations and stabilizes six checkpoints spanning the whole spectrum together, while data-side filtering slows collapse without stopping it.
The Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages, is detailed, which involves multilingual Multiple-Choice Question Answering across 21 languages.
It is shown the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, indicating that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing.
Post-cutoff test sets only prevent verbatim window memorization, revealing that time-series foundation model outperformance is driven by pretraining corpus familiarity rather than generalized temporal reasoning.
Winning lottery tickets at up to 95% sparsity can be harvested essentially for free simply by piggybacking iterative magnitude pruning onto the natural retraining loop of active learning.
Even "leak-free" protein interaction benchmarks are riddled with non-biological shortcuts that models readily exploit—and standard negative-sampling heuristics often make the bias worse.
Treating automated feature transformations as permutation-invariant hierarchies rather than ordered sequences eliminates a fundamental representation bias, enabling policy-guided RL to efficiently navigate non-convex search spaces.
Compact transformers trained purely on synthetic clinical text can match or exceed real-data de-identification performance while proving more robust to format perturbations across out-of-distribution primary care domains.
Full rationale supervision burns up to 254× more tokens without guaranteeing reliability, whereas isolating representations with high rationale-boundary sensitivity selectively immunizes medical LLMs against choice-order brittleness.
Open-weight embedding models fine-tuned via contrastive ranking beat frontier LLMs like GPT-5.2 and Claude-Sonnet-4.6 by over 10 percentage points on real-world entity disambiguation.
Industrial warehouse safety exposes a massive capability gap in frontier LMMs, where even top proprietary models consistently fail to jointly ground visual hazards and reason over operational risks.
Romanized and regional dialects represent a massive evaluation blind spot for frontier LLMs—even across globally dominant low-resource languages like Bangla.
Standard SFT degrades financial reasoning benchmarks by up to 4 percentage points, but pairing self-distillation with rule-verified GRPO transforms domain post-training from a capability tax into a 3-point gain.
Discarding "easy" examples where the base model already follows glossary constraints yields an 11-point accuracy leap at fixed data volume, outperforming RL-based alignment methods with standard supervised fine-tuning alone.
Standard pixel augmentations frequently corrupt delicate vision-language alignment, but injecting diffusion-style isotropic noise directly into embedding spaces breaks through the longstanding performance ceiling of stacked CutMix, Mixup, and RandAug recipes.
Diffusion-generated thermal faces can effectively break the multi-modal data bottleneck in biometrics, outperforming models trained on scarce real pairs without requiring costly image translation at inference time.
Curriculum learning's edge on hard examples is not just an artifact of data exposure: modeling curricula as Wasserstein transport paths reveals that ordering genuinely drives capability shifts, though no universal pacing strategy dominates across tasks.
Counterpart data sharing silently distorts production A/B tests, but selectively filtering bid and ranking divergences outperforms both naive log-sharing and data-starved log-splitting.
Differentiable cardiovascular physics can filter out ICU sensor artifacts that fool standard deep nets, yielding a 5-point SOTA leap on VT alarm reduction alongside a 2x boost in label efficiency.
Open-set GNNs silently break when graphs violate homophily, but synthesizing pseudo-unknown proxies along cross-class displacement vectors restores robust out-of-distribution detection on heterophilic topologies.
Instead of relying on crude, handcrafted perturbations for radiation therapy safety margins, generative latent velocity fields can now simulate realistic, continuous 3D patient anatomical deformations at full clinical CT scale.
Masking text-level script and font details during English-only training yields a massive 20.8% absolute F1 boost across 15 unseen multi-script languages without requiring a single multilingual training sample.
Scaling and post-hoc interpretability hit a hard wall in agentic AI because models cannot recover causal invariances that observational interaction data never contained in the first place.
Domain gaps baked in during mid-training are virtually immutable: compensatory SFT failed to bridge a single domain performance disparity at a 5% threshold despite lifting overall baseline accuracy.
Clip-level captions compress away the continuous visual dynamics needed for true long-horizon reasoning; Kairos restores this signal with dense, time-resolved annotations tracking actions, entities, and attributes across 10-to-30-minute videos.
This is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability and all three checkpoints rank 1st among models at comparable parameter scales.
Statistical physics meets LLM evaluation: mapping multi-category rater agreement onto a Potts model preserves true rubric ordinality without requiring rigid threshold or distance assumptions.
Offloading causal DAG discovery to LLM world knowledge enables Bayesian networks to generate high-fidelity synthetic tabular data from just a 2% sample while strictly retaining empirical grounding for numerical parameters.
Arbitrarily splitting identical retrieved records into separate chunks swings an LLM's evidence weighting by up to 32 percentage points without altering a single word of context.
A 4B parameter model hits 86.4% on the Berkeley Function Calling Leaderboard using just 11K synthetic examples, proving that decomposed generate-verify-refine loops can outperform massive, brute-force filtered datasets.
Hardware interpolation artifacts in standard 4-analyzer polarization sensors quietly corrupt polarimetric ground truth, but dense 180-angle temporal sampling slashes AoLP error from $13.36^\circ$ to $2.21^\circ$ to unlock clean RGB-to-polarization synthesis.
Angular-to-Cartesian projection distortions have long bottlenecked spherical 3D perception, but explicitly modulating voxel features with spherical range-azimuth geometry outperforms top surround-view models like SurroundOcc and TPVFormer across varied wild environments.
Millimeter-wave person re-identification does not need continuous walking: conditioning point-cloud experts on everyday indoor activities catapults Rank-1 accuracy from 59.1% to 82.1%.
Synthetic domain shifts cannot replace authentic historical variance, but using synthetic degradation to complete missing paired identities yields up to a 3.69 percentage point R@1 boost under severe data scarcity.
DYAD (DYadic Assistance Dataset), a synchronized multimodal record of human-human assistance during gearbox assembly, is introduced, a linked interaction structure spanning help seeking, intervention choice, execution, and outcome under egocentric and workspace sensing.
Better human mesh recovery does not guarantee more accurate gait parameter estimation, exposing a critical disconnect between standard visual 3D pose metrics and downstream biomechanical fidelity.
Reducing educational data quality to a single scalar leaves benchmark gains on the table: decomposing text value into six distilled rubric classifiers outperforms FineWeb-Edu mixtures and doubles as an effective reward signal for GRPO post-training.
This work introduces ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans, and introduces ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness.
Even frontier LLMs collapse to sub-10% accuracy when facts in a long context must be dynamically updated or revoked, but training on executable state-machine simulations reliably repairs this state-tracking failure across diverse out-of-distribution tasks.
Hardware dataset curation no longer requires manual code refactoring, thanks to an agentic pipeline that autonomously trawls academic repositories to carve out synthesizable HLS designs for LLM training.
Transforming archival collections into engaging conversational experiences can revolutionize how we interact with history in public spaces.
RoboCousin can generate over a million expert trajectories for bimanual manipulation, making it easier than ever to scale training data without extensive manual effort.
These results establish technical consistency for the evaluated FINALLY workflow and show that the implemented strategies follow their intended optimization directions within the investigated configuration space, but do not establish the scientific suitability, global optimality, or practical superiority of the generated selections.
A data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations is proposed.
Existing cattle data contains rich semantic information that can vastly improve data interoperability without the burden of creating new metadata.
Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts, is presented, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts.
A methodological framework for autonomous, AI-based harvesting of cultural knowledge from the open web, designed to expand corpus coverage while intensifying per-entry data depth is presented, including the architectural guarantee that the machine never overwrites human contributions.
Key to the work is the use of bottom-up clustering to preserve local structure in the embedding space, improving the clustering quality of the resulting Semantic IDs and their utility for downstream generative retrieval.
This work introduces SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination and proposes a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks.
Human-to-robot demonstration transfer notoriously collapses when tasks leave the 2D plane, but decoupling boundary keyframes from dense trajectory generation unlocks precise 3D flips and rotations across more than 1,200 objects.
Stripping boilerplate and line duplicates before running heuristic filters rescues 19% of valid web text that standard data pipelines prematurely throw away.
Swapping network traffic features in the frequency domain preserves critical structural signatures while generating diverse, out-of-distribution variations that standard time-domain augmentations fail to synthesize.
Multimodal RGB-pose fusion maximizes in-domain fall detection accuracy, but skeletal representations alone are what survive out-of-distribution shifts to unseen environments.
Data providers no longer have to feed continuous ML training streams on blind trust: dormant consumers are automatically cut off, and any data leak can be cryptographically traced back to the exact recipient using as few as 40 tabular rows.
Even after targeted fine-tuning, LLMs peak at an F1 score of just 0.58 on cyber threat level determination, revealing that current models remain far too brittle for operational SecOps workflows.
Merging independently specialized checkpoints in parameter space beats joint multi-source continued pretraining for domain adaptation, retaining complementary domain signals that joint training tends to wash out.
Pretraining loss actively misleads video data curation: broad data coverage systematically beats aggressive quality filtering across downstream benchmarks in both Wan 2.1 and V-JEPA 2.1.
D DianShi-RxnDB is presented, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes integrating patent text, images, and reaction schemes.