Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
A single random feature space can achieve high spectral accuracy for multidimensional targets while exposing a critical trade-off with ill-conditioning.
Prematurely abandoned in modern scaling recipes, layer dropout can actually slash LLM pre-training compute by 25% while natively unlocking 1.5x faster inference through zero-shot layer skipping.
LLMs can match or outperform traditional CKD screening methods with minimal examples, but their stability diminishes with increased input complexity.
Achieving an 11% boost in energy efficiency for edge AI applications by smartly balancing throughput and latency through hierarchical operator parallelism.
Prematurely abandoned in modern scaling recipes, layer dropout can actually slash LLM pre-training compute by 25% while natively unlocking 1.5x faster inference through zero-shot layer skipping.
LLMs can match or outperform traditional CKD screening methods with minimal examples, but their stability diminishes with increased input complexity.
Achieving an 11% boost in energy efficiency for edge AI applications by smartly balancing throughput and latency through hierarchical operator parallelism.
Bypassing retraining, OSR achieves competitive classification performance while significantly improving computational efficiency and privacy preservation.
Causal optimal transport reveals a new pathway for guiding degenerate diffusion models, even when traditional score functions fail.
GBTs can outperform neural networks in game-playing scenarios, revealing a potential misalignment in the community's reliance on deep learning methods.
Query-independent eviction signals can maintain 100% retrieval accuracy in KV caches while compressing data without any measurable cost.
Principled replay selection can significantly cut down wall-clock time in RL training while maintaining or even improving performance metrics.
Attention parameterization can trap models in uninformative states, but symmetry-breaking mechanisms can dramatically enhance recovery efficiency.
Sequential refinements using MPI can enhance offline RL performance beyond traditional behavior constraints, outperforming established baselines.
Correlated weight initialization can fundamentally alter the asymptotic behavior of deep residual networks, revealing critical hyperparameters that influence performance.
Compressed models may achieve high accuracy but suffer a dramatic drop in adaptability during test-time, revealing a critical trade-off in model deployment.
Achieving up to 9.9 times the throughput of traditional methods, this GPU-accelerated solver transforms the efficiency of constant optimization in symbolic regression.
ESPO not only boosts accuracy by nearly 4 percentage points but also slashes prompt length by nearly half, redefining efficiency in prompt optimization.
Auxiliary views can enhance LLM learning efficiency, revealing that reallocation of training tokens can lead to better factual recall even when using weaker teacher models.
Speed up 3D Gaussian Splatting training by over 70% without sacrificing reconstruction quality through a novel frequency-staged approach.
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
Achieving dimension-independent linear convergence at unit step size could revolutionize how we compute Bures-Wasserstein barycenters in high-dimensional spaces.
Current AI methods for data center optimization overlap significantly, making it impossible to rank their effectiveness—CLEAR-DC aims to change that.
WeatherNext 3 achieves state-of-the-art probabilistic medium-range forecasting by integrating real-time satellite data, outperforming traditional models in accuracy and resolution.
QAT-FM reduces coupling construction costs while achieving competitive generative performance, transforming how we approach high-dimensional generative tasks.
LeanGRPO achieves up to 1.83x speedup in diffusion RL without sacrificing optimization quality by eliminating redundant computations.
Task-adaptive pruning can cut model size by nearly 25% while boosting accuracy in histopathology tasks, challenging the notion that bigger always means better.
A single random feature space can achieve high spectral accuracy for multidimensional targets while exposing a critical trade-off with ill-conditioning.
Achieving high-quality model performance with just 10% of the required labels could revolutionize the scalability of RLVR in large language models.
Scheduling imagination in VLA models can cut GPU costs by 80% while boosting performance and robustness in real-world tasks.
sp-DBA accelerates transform-domain computations by up to 28.1 times, revolutionizing how we approach scientific simulations on parallel hardware.
BU-MBAR achieves faster and more stable convergence for MBAR equations by dynamically adjusting bin widths, revolutionizing statistical analysis in thermodynamics.
TruncGradGS tackles the gradient vanishing problem in 3D Gaussian Splatting, leading to superior scene reconstructions across various initialization methods.
DSAQuant reveals that aligning quantization training with the stage-wise nature of video diffusion can drastically enhance visual fidelity in text-to-video generation.
LLM agents trained without any programmatic verifiers can actually outperform models trained on ground-truth reward signals when trajectory-level rubric judgments are dynamically decomposed into step-level advantages.
TaRA achieves superior initialization for LoRA by ensuring low-rank gradients closely match full-rank gradients, leading to better performance in fine-tuning tasks.
LoRA-TSD achieves superior performance over all existing LoRA optimizers while offering the first global convergence guarantees for low-rank adaptations.
Polyak's momentum can significantly expand the critical batch size, enabling more parallelism without compromising data efficiency.
Achieving a new anytime lower bound of \(Ω(n^{-1.2408})\) reveals critical insights into the acceleration limits of gradient descent.
Switching to E5M3 block scaling not only simplifies FP4 pretraining but also yields significantly improved training and validation performance.
Fixing the teacher's weights during Test-Time Adaptation can lead to substantial performance gains and greater robustness against hyperparameter changes.
Verified reliability outperforms domain expertise in teacher selection, leading to substantial performance gains in multi-domain LLMs.
SAPE-FL achieves superior model performance in heterogeneous environments by intelligently balancing global and local learning through similarity-aware personalization.
Matching graph complexity to task granularity can dramatically enhance performance in power grid control, challenging the notion that more complex representations are always better.
IFW-BLS achieves superior robustness against noise and outliers, outperforming traditional models by intelligently down-weighting unreliable samples.
DKL boosts RAG accuracy by over 25 points in retrieval failure cases without the need for costly instruction fine-tuning.
HyperMC achieves superior hyperparameter tuning for SGMCMC, significantly improving Bayesian inference performance while ensuring reproducibility.
Efficiency in GUI agents is as crucial as task success, with recent advancements converging on innovative strategies like selective reading and hybrid execution models.
Training on the entire deployed matrix can unlock up to 81.4% of a model's capacity that was previously unreachable, leading to unprecedented performance gains in low-rank distillation.
GaLe achieves exact-inference performance with a staggering 65% speedup and 90% RAM reduction, revolutionizing how we deploy models on embedded devices.
Forget costly supervised fine-tuning—this new contrastive reward mechanism enables efficient zoom-in tool learning in MLLMs, outperforming traditional methods.
Gradient-based optimization can transform how data centers allocate load by leveraging differentiable electricity market clearing, achieving near-optimal solutions efficiently.
Privacy and robustness in federated learning are not just additive; they interact in ways that can significantly impact model performance in adversarial settings.
Tuning polyhedral optimizations with hill climbing can yield up to 28% faster execution compared to static optimizers, bridging the gap between fixed-cost compilation and full autotuning.
BASP slashes communication overhead in LLM training, boosting efficiency by over 30% while maintaining model performance.
BBYT slashes candidate-selection time by over 18% while maintaining the accuracy of timing-aware logic rewrites, revolutionizing how we approach timing evaluations in circuit design.
Neural networks prune themselves during optimization not through gradual simplification, but via abrupt percolation phase transitions where architectural symmetries force subnetworks to merge in discrete, cascading blocks.
Because mode-seeking reverse KL aggressively amplifies incorrect teacher signals, gating dense distillation on verifier-scored teacher probes systematically outperforms uniform distillation while reclaiming massive amounts of idle teacher compute.
Influence-guided response rewriting is introduced, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed, motivating intervention-aware evaluation of TDA methods.
Learnable tokenization can dramatically boost sample efficiency in language model pretraining, leading to superior performance in zero-shot tasks.
Conflict rates in modern optimizers can reach up to 86.3%, but a new alignment technique can ensure updates remain conflict-free, drastically improving training outcomes.
Achieving significant network compression without sacrificing performance, LRNBA allows for the creation of deeper and wider models under the same parameter budget.
Global allocation of precision budget in LLMs can improve accuracy by 21-52 points compared to local layer-specific repairs, defying conventional wisdom about quantization damage.
PLES can uncover optimal hyperparameter scaling laws for LLMs using a fraction of the resources typically required, revolutionizing hyperparameter tuning efficiency.
Algorithm-dependent learnability reveals that focusing on the optimizer's trajectory can significantly enhance offline optimization performance, leading to state-of-the-art results on complex tasks.
TRIAGE transforms LLM efficiency by turning historical execution into reusable skills, slashing token consumption to zero for many queries.
Optimal hyperparameter choices for supervised fine-tuning can vary dramatically with model scale and architecture, challenging conventional wisdom in AI training practices.
ModalShare reallocates bandwidth based on each modality's contribution, boosting accuracy by over 15 percentage points compared to traditional equal keep-ratios.
Warmup strategies can drastically alter the perceived benefits of online adaptation, with variations leading to performance shifts of nearly 19 percentage points.
Committed reveal sampling (CRS) significantly lowers generative perplexity by leveraging persistent context, outperforming traditional top-$p$ sampling methods.
A fixed span can recover most of the benefits of movable low-rank factors in LLM adaptation, achieving superior performance with drastically fewer trainable parameters.
The central flow of gradient descent reveals a surprising interplay of fast oscillations and slow dynamics that could redefine our understanding of stability in deep learning.
HarnessEvolve not only solves the credit assignment problem but also prevents agents from falling into the traps of shortcut learning and catastrophic forgetting, ensuring robust self-evolution.
Training agents on compressed contexts can lead to significant log-probability gaps, but innovative methods like LogitTree and SDCC offer a robust solution that maintains performance consistency.
Subspace Levenberg-Marquardt algorithms can outperform traditional methods in training larger neural networks without sacrificing performance.
Fine-tuning small-to-medium LLMs can be dramatically improved with just two online rollouts, reshaping efficiency in model training.
Self-Routing adapts optimization strategies on-the-fly, leading to significant improvements in mathematical reasoning tasks without relying on external supervision.
Outperforming prior open models, Instella-MoE sets a new standard for efficiency and performance in language modeling with its innovative architecture and training methods.
SFAD achieves a remarkable 2.48x speedup in inference while enhancing contextual faithfulness, addressing one of the most pressing challenges in large language models.
Understanding how concurrent AI training jobs disrupt each other could redefine resource allocation strategies in supercomputing environments.
Stochastic Riemannian optimizers for tree tensor networks achieve competitive performance while enabling stable compression, revolutionizing their utility in machine learning applications.
Pix2Rep-v2 achieves remarkable data efficiency in dense medical imaging, outperforming fully supervised methods even in few-shot settings.
Tailoring iteration counts for nonlinearity approximations can cut inference latency in encrypted language models by over 40%.
GlitchLab achieves a staggering 2-85x reduction in attempts and 26-1,237x decrease in time for fault discovery compared to traditional methods.
By eliminating bidirectional communication, L-shaped SFT allows edge devices to participate in fine-tuning without the need for constant connectivity, drastically cutting down on communication overhead.
A new protocol achieves a contraction rate of \( 1/\sqrt{2} \), pushing the boundaries of convergence in multidimensional approximate agreement closer to optimality.
Achieving a 15.1% accuracy boost in neural network training while enhancing energy efficiency by over 33 times could redefine on-chip training paradigms.
A hybrid CNN-BiLSTM model with intermediate layers achieves unprecedented accuracy in State of Health estimation for battery systems.
Structural and magnitude pruning can retain significant functional redundancy, revealing that model compression strategies may overlook critical parameter directions.
Silent data corruption can derail LLM training, but TrainSDC offers a targeted solution that maintains performance even under fault conditions.
T3S achieves unprecedented efficiency in multi-task reinforcement learning by tailoring features and task selection, outperforming existing methods in robotics tasks.
Achieving superior performance with one-third the resources, Qwen3.8-Flash-Next redefines efficiency in large-scale language models.
Precomputed memory in language models can degrade significantly unless rebuilt frequently and updated with specifically phrased corrections.
Speculative decoding can be significantly accelerated by training models to anticipate verification outcomes, leading to faster and more efficient inference.
OPD's effectiveness hinges less on teacher supervision than previously thought, with a new method achieving a staggering 263% relative gain without any teacher input.
A fixed-dimensional latent space can achieve vanishing approximation error in neural networks, challenging conventional beliefs about the necessity of high-dimensional representations.
TMB-based neural network mixed-effects models eliminate the burdensome manual derivations, streamlining complex modeling while enhancing statistical accuracy.
Achieving state-of-the-art performance in continual learning with a single shared adapter, FACET reduces parameter usage while enhancing feature discrimination across tasks.
TSPFN outperforms both traditional tabular models and specialized deep learning approaches in classifying physiological time series, showcasing its superior generalization capabilities.
Discrete differentiation of gradient descent reveals unexpected curvature behaviors that challenge conventional stability assumptions in ReLU training.
Fine-tuning low-bit models can be both efficient and deployment-faithful, thanks to a new optimization technique that leverages code surrogate gradients.
Self-play training for autonomous driving reveals critical failure modes that could hinder real-world deployment, including reward hacking and inadequate adherence to traffic rules.
Class-imbalance bias in long-tailed semi-supervised learning can be systematically mitigated with a novel dynamic pruning approach that reallocates gradient budgets for better generalization.