Search papers, labs, and topics across Lattice.
75 papers published across 4 labs.
ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.
This work derives the generalized score matching objective on a convex subset of $\mathbb{R}^{d}$ constructively starting from Minimum Probability Flow (MPF) learning, and shows how classical score matching as well as domain-adapted variants for non-negative data arise naturally within the proposed framework.
This work proposes MomentUm SpEctral Clipping (Musec), which replaces Muon's spectral flattening with spectral clipping: rather than setting all singular values of the momentum matrix to approximately one, Musec clips singular values that exceed a threshold while preserving the underlying spectral structure of the momentum.
Non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization, and transfer from non-language data is less efficient than additional language data.
Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop.
ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.
This work derives the generalized score matching objective on a convex subset of $\mathbb{R}^{d}$ constructively starting from Minimum Probability Flow (MPF) learning, and shows how classical score matching as well as domain-adapted variants for non-negative data arise naturally within the proposed framework.
This work proposes MomentUm SpEctral Clipping (Musec), which replaces Muon's spectral flattening with spectral clipping: rather than setting all singular values of the momentum matrix to approximately one, Musec clips singular values that exceed a threshold while preserving the underlying spectral structure of the momentum.
Non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization, and transfer from non-language data is less efficient than additional language data.
Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop.
This work addresses the adverse interactions between sparsity and data repetition, and presents evidence for the core mechanisms of overfitting and its potential remediation, and suggests promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
This paper extends CONES to allow for loss functions $f_t'$s to also change over time, and shows that any online algorithm with sublinear {\it anytime} regret has a movement cost of $\Omega\left(\log T\right)$.
Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.
A novel error analysis is provided that provides substantially sharper bounds for products of operators, thereby significantly relaxing existing restrictions on the maximum number of local machines while retaining optimal learning rates for the distributed kernel-based robust gradient descent algorithm.
This work introduces CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve refinement that combines cross-space ranking robustness with low-cost architecture selection.
A four-coefficient parameterization lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t) in which direction-aligned proxies of EOPD and ToDi appear as one-dimensional (1D) restrictions, and which adds multi-channel composition and an explicit bias as further degrees of freedom.
AdamX, a first-order optimizer that incorporates cosine similarity as an adaptive mechanism for controlling update magnitudes is introduced, which is scalable, model-agnostic, and straightforward to integrate into existing training pipelines.
Looped flows are proposed, an approach that allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples.
LOCUS, a method that selects a task-aware low-rank adaptation subspace to minimize output-token cost subject to a utility constraint, is presented, a method that selects a task-aware low-rank adaptation subspace to minimize output-token cost subject to a utility constraint.
This work describes an automatic criterion for full-state rejuvenation of the Gibbs sampler, derived from the Gelman-Rubin statistic, which plays a key role in speeding up learning convergence.
The model incorporates three core components: spatiotemporal embedding module, spatiotemporal fusion module, and LLM backbone, which adopts a differentiated parameter adaptation strategy to balance training efficiency and traffic data adaptability.
This paper model MTO as a sequence of sequential transfer optimization problems, concentrating evaluations on a single target per iteration, and proposes a likelihood-informed task prioritization mechanism to maximize transfer utility by identifying the task most likely ready for knowledge integration.
The XAI-assisted Recurrent neural network Attribution for Channel Estimation (X-RACE) framework is proposed, and novel temporal XAI metrics: Saturation Time, Importance Drift, and Relevance Contrast are proposed to characterize the LSTM's learning dynamics and memory convergence.
The proposed greedy algorithm with alternating updates guarantees the four-peak distribution required for stable 2-bit clustering of each factor and admits closed-form initialization of cluster centers, removing the multi-restart k-means bottleneck of prior work.
Whether large language models (LLMs) are narrowing the variety of methods archaeologists use is evaluated, with results consistent with LLMs pushing methodological choice towards convergence, although this study cannot establish a causal effect.
Intermediate activations in split LLM fine-tuning trivially leak raw training prompts past standard privacy defenses, but a learned obfuscation pipeline closes this leakage vector without destroying model utility.
Cross-silo federated learning achieves up to 40% faster time-to-accuracy when SDN-level network telemetry—rather than client-side heuristics—determines which institutions train synchronously versus asynchronously.
Rockafellar's classic interior-domain constraint qualification is insufficient to guarantee that the sum of two maximally monotone operators remains maximally monotone, overturning a foundational conjecture in convex analysis.
Winning lottery tickets at up to 95% sparsity can be harvested essentially for free simply by piggybacking iterative magnitude pruning onto the natural retraining loop of active learning.
Randomization slashes ERM-oracle query complexity from linear down to logarithmic for online threshold learning, revealing that an oracle's internal tie-breaking rule can single-handedly trigger an exponential computational gap.
Standard quadratic regularization fails to contain runaway client drift under extreme heterogeneity, but tuning the proximal exponent to $p \in [5, 7]$ slashes severe-stress federated loss by over 23%.
Worst-case adversarial bandits no longer have to learn from scratch: transferring predictable task-level priors across random action sets achieves provably sublinear transfer regret with intrinsic-dimension $\mathcal{O}(\sqrt{n})$ guarantees.
Skipping gradient updates on rounds without prediction errors completely removes the long-standing $\log T$ horizon dependence in online inverse optimization, securing constant $O(d^2)$ regret for forward integer programs without expensive geometric oracles.
Deterministic black-box optimization no longer hits a wall beyond low dimensions, unlocking 27% better solution quality and the fastest convergence on 40% of medium-scale benchmark problems.
Unfolding convolutional kernels breaks Muon's optimization geometry; aligning polar updates with the true convolution operator in the frequency domain cuts flow-matching compute by nearly 40% while radically outperforming Adam.
Overparameterized networks are thermodynamically driven toward smooth, data-adaptive solutions by a function-space fluctuation potential that penalizes misalignment between dynamic mobility and parameter-space curvature.
Standard SFT wastes gradient budget on tokens models either already know or cannot yet grasp; trimming supervision from both extremes yields up to a +26.9 point boost on MATH500 with zero reference-model overhead.
Standard Gaussian teacher assumptions hide severe optimization traps: hidden node alignment alone dictates whether gradient descent fails into boundary versus interior local minima.
Frontier-scale agentic RL on 700B+ MoEs is no longer locked inside hyperscaler proprietary stacks, achieving stable 263-second step times across 64 GB300 GPUs via a verified, open-source training infrastructure.
Constant learning rate combined with weight decay is mathematically incapable of maintaining a stable interior equilibrium in normalized networks, driving recurrent instabilities that can be precisely mapped and controlled via a single scalar law.
Curriculum learning's edge on hard examples is not just an artifact of data exposure: modeling curricula as Wasserstein transport paths reveals that ordering genuinely drives capability shifts, though no universal pacing strategy dominates across tasks.
Polynomial-time sampling in spin glasses and sparse Bayesian regression is tractable deep into low-temperature regimes, slashing the measurement barrier for spike-and-slab posteriors from $k^3$ to $k^{3/2}$ while approaching the Almeida–Thouless phase boundary.
Recurrent models can extrapolate up to 128× beyond their training horizon simply by stabilizing backward-pass state credit without altering the forward computation.
Vanilla gradient descent cannot beat the silver stepsize schedule, proving that predetermined learning rates hit a hard theoretical limit of $\mathcal{O}(n^{-\log_2(1+\sqrt{2})})$ without momentum.
Amari's long-overlooked Bayesian duality and standard convex duality are fundamentally two sides of the same coin, unlocking a unified geometric foundation for probabilistic update rules in AI.
Nearly 40% of GRPO rollouts yield zero gradient due to uninformative all-or-nothing groups, but an offline anchor pass can slash this cold-start compute waste before the target policy generates a single token.
Optimal learning rate and batch size in MoEs depend directly on the expert activation ratio rather than active or total parameter counts, unlocking predictable hyperparameter transfer even at 1/64 sparsity.
Analog over-the-air model aggregation can match the theoretical $\mathcal{O}(1/\sqrt{T})$ convergence of ideal FedAvg without requiring instantaneous channel state information or strict phase alignment.
Functional linear regression can achieve minimax-optimal prediction rates even when covariates are masked by discrete, noisy sampling, bridging the long-standing gap between idealized continuous theory and imperfect real-world sensor streams.
On-policy LLM distillation does not actually need precise advantage magnitudes: retaining merely the directional sign of token advantages matches standard distillation performance while Total Variation smoothing eliminates late-stage training instability.
Hidden-state representation alignment in diffusion models isn't just an empirical heuristic—it is mathematically equivalent to score distillation operating directly within the latent representation space.
Standard SGD can get arbitrarily close to an anytime convergence rate of $\sqrt{\log n / n}$, but hitting that benchmark exactly is mathematically impossible under any deterministic step-size schedule.
Adam silently breaks down on equivariant networks because irrep multiplicities distort spectral step sizes across channels—a flaw that simple blockwise update normalization fixes to match Muon's performance.
Neural network feature learning obeys a strict Fisher-information speed limit, governing the exact sequence and timescales with which latent representations can emerge from stochastic gradient dynamics.
Nearly a third of variable-wise gradients directly oppose each other during standard multivariate forecasting, but resolving these conflicts via single-pass gradient surgery reliably boosts accuracy without multiple backward passes.
Exploiting spatial proximity between hospitals enables direct AUC maximization under strict privacy constraints, preventing the discriminative performance collapses typical of standard federated learning on skewed clinical data.
Naive linear reward scalarization inevitably triggers destructive gradient tug-of-wars, but decoupling marginal objectives via orthogonal projections and localized constraints stabilizes training across both recommender systems and complex LLM alignment tasks.
Constrained min-max optimization just caught up to its unconstrained counterpart: the long-standing $\widetilde{O}(\varepsilon^{-4})$ stochastic complexity bound for natural residuals collapses down to optimal $\widetilde{O}(\varepsilon^{-2})$.
Orthogonally constrained low-rank adapters produce dangerously overconfident predictions, but transporting ensemble particles directly along the Stiefel manifold via Stein variational inference recovers reliable calibration without sacrificing accuracy.
Black-box optimization can now safely leverage noisy, low-cost gradient approximations to bridge the gap between $O(d/T)$ zeroth-order and $O(1/T)$ first-order rates without stalling when guidance fails.
Treating all deleted nodes uniformly causes catastrophic utility collapse in graph unlearning, but ordering removals by topological difficulty retains 74% performance at 20% mass deletion where prior baselines plummet to 26%.
Pretraining loss is a deceptive selection metric: at 30B MoE scale, downstream SFT performance is governed not by benchmark scores, but by the checkpoint's solution density under local weight perturbations.
Rigorous optimality certificates are no longer restricted to polynomial optimization: constructive Positivstellensätze now extend to broad classes of non-polynomial, definable learning objectives with bounded computational graph complexity.
Despite the allure of physics-inspired curves, BrachistoneLR reduces algebraically to cosine annealing terminating one epoch early, proving that adopting smooth decay matters significantly while the choice between specific smooth parameterizations does not.
Attempts to fix the vanishing Jacobian in SAC lead to worse performance, highlighting that saturating a bound is not a solution for optimal control problems.
Adjusting joint stiffness dynamically can transform the training landscape for legged robots, enabling successful execution of complex tasks that standard methods fail to solve.
DiffLUT-Net is presented, an FPGA-native network connected by six-input LUTs that are trained from scratch, demonstrating the effectiveness of jointly learning LUT functions and sparse connectivity for compact FPGA-native inference.
StochBench is introduced, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source that better represents domain-specific applied mathematics while remaining challenging for advanced provers.