Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
Achieving substantial speedups in kNN regression calculations on large datasets by leveraging CUDA GPUs could transform how we approach real-time data analysis.
YOLO-PEFT transforms the fine-tuning landscape for real-time detectors by replacing trial-and-error with structured, auditable planning, achieving notable performance gains.
Modular TTT reveals that the right combination of learning rate and weight decay can significantly boost performance, while deeper architectures may actually degrade results.
An exact closed-form update for orthogonality-constrained optimization could revolutionize efficiency in matrix-aware methods like Muon.
AutoThread reduces simulation execution time by up to 83.8% while boosting throughput to 1.8x that of existing RL inference methods.
Achieving substantial speedups in kNN regression calculations on large datasets by leveraging CUDA GPUs could transform how we approach real-time data analysis.
YOLO-PEFT transforms the fine-tuning landscape for real-time detectors by replacing trial-and-error with structured, auditable planning, achieving notable performance gains.
Modular TTT reveals that the right combination of learning rate and weight decay can significantly boost performance, while deeper architectures may actually degrade results.
An exact closed-form update for orthogonality-constrained optimization could revolutionize efficiency in matrix-aware methods like Muon.
AutoThread reduces simulation execution time by up to 83.8% while boosting throughput to 1.8x that of existing RL inference methods.
BaKron accelerates neural network quantization by harnessing richer curvature information, achieving significant computational efficiency without sacrificing performance.
Early stopping can transform gradient descent from a source of suboptimality into a strategy for achieving minimax-optimal classification performance in noisy settings.
Leveraging large language models to generate synthetic transitions, ProDVI boosts sample efficiency in deep reinforcement learning without the need for extensive pre-collected data or simulators.
SkillTFM achieves a remarkable AUC boost of up to 0.142 without any training, redefining how we adapt tabular models in the face of distribution shifts.
Operational failures in multi-node fine-tuning can be mitigated by prioritizing power monitoring over utilization metrics, revealing critical insights for practitioners.
Early stopping of accumulation in binary neural networks can slash computation by over 86% with minimal accuracy trade-offs.
Self-PreTraining boosts transformer accuracy in medical time series by up to 6 percentage points, even with limited data.
Co-optimizing network and ML parameters can accelerate training by up to 42%, unlocking new efficiencies in AI workloads.
Learning to Rank models can significantly enhance the efficiency of tensor-network contraction plans, outperforming traditional methods and adapting across different GPU architectures.
LC-Implicit-QAOA slashes the computational overhead of QAOA training, achieving remarkable accuracy with a fraction of the resource demands.
Zero-loss exactness in optimal transport dynamics could revolutionize how we approach matching problems in high-dimensional spaces.
Langevin correction can transform the way flow-based generative models optimize policies, leading to sharper and more accurate sample generation during reinforcement learning.
Late fusion in quantum machine learning achieves near-identical accuracy to full reconstruction while slashing costs and enhancing noise resilience.
Cheap models can recover early evolutionary progress, enabling a shift in budget allocation that dramatically enhances LLM-driven algorithm discovery.
Gated Hindsight Distillation allows GUI agents to learn from future observations, drastically improving their reasoning capabilities in complex environments.
Viveka achieves up to 75% energy savings in smart wearables by intelligently adapting sensing strategies based on context reliability.
Self-supervised learning can outperform fully supervised methods in machinery fault diagnosis, achieving high accuracy with significantly less labeled data.
GROM achieves rapid, effective unlearning in seconds while maintaining model performance, outpacing traditional methods that struggle with computational efficiency and content recovery.
Skill contamination in LLM agents can lead to irreversible performance degradation, but a structured filtering approach can prevent this and enhance overall capabilities.
LiteKD-Net achieves superior image denoising performance on mobile devices while slashing runtime costs, setting a new standard for efficiency in the field.
The shift from parameter-centric to system-level adaptation in continual learning could redefine how we build and interact with AI models.
Fine-tuning on planning-aware trajectories can enhance model performance across diverse CLI environments, mitigating the pitfalls of scaffold-specific training.
Achieving 90% sparsity, BnBERT-iPET rivals larger models while drastically reducing computational costs for Bengali NLP tasks.
MALT outperforms Muon in pretraining language models while keeping memory usage and computation time nearly identical.
Path-level pretraining in MultiPathFormer leads to a dramatic 59% improvement in wireless propagation estimations, reshaping the landscape of wireless foundation models.
Gradient descent can slash held-out error by up to 44% compared to Adam in low-rank matrix recovery, revealing the critical role of optimizer choice in performance outcomes.
Exploration bonuses can either amplify or neutralize memory architecture differences, depending on how memory content is acquired and supervised.
Optimal training time in gradual adaptation scales inversely with the number of tasks, revealing a critical balance for effective learning.
Robustness to adversarial perturbations can dramatically change the sample complexity landscape, shifting accuracy dependence from linear to polynomial rates.
Self-distillation conditioned on privileged information may lead to a model that is less capable of reasoning, as it optimizes for a misleading signal rather than task success.
CL-PINN outperforms traditional methods by balancing accuracy and efficiency, making it a game-changer for solving parameterized PDEs without observational data.
Strong $L^2$ convergence in functional flow matching reveals that learned flows can achieve robust performance even without uniqueness assumptions in their underlying dynamics.
Recurrent Vision Transformers can outperform standard models in accuracy-to-parameter trade-offs when memory constraints are prioritized, challenging conventional wisdom about architectural efficiency.
MUON's generalized variant reveals a surprising non-convergence behavior, challenging its effectiveness in deep learning applications.
Early-stage gradient descent can achieve weak alignment with the max-margin direction in far fewer iterations than previously thought, challenging conventional beliefs about convergence rates.
Counterfactual recoverability transforms how we approach error correction in on-policy distillation, leading to a staggering AUC of 1.000 compared to 0.392 with divergence alone.
MESH achieves a 62.5% reduction in optimizer-state memory for Mixture-of-Experts training while preserving performance, challenging assumptions about memory efficiency in deep learning.
GOAL not only optimizes advertising incentives but also adapts seamlessly to varying ROI constraints without the need for retraining, revolutionizing how we approach incentivized user engagement.
Achieving a 3.56× increase in decoding throughput without sacrificing accuracy, BinaryPC revolutionizes efficiency in long-context LLMs.
ORACLE achieves up to 104.4x faster circuit design while meeting nearly all target specifications, revolutionizing multi-objective optimization in analog circuit design.
Downscaled training can yield gradients nearly identical to native ones, but only within a specific noise window that defies spectral predictions.
LoRA+ outshines other PEFT methods, achieving the best energy efficiency in 19 out of 24 configurations, paving the way for practical on-device personalization of language models.
A novel IDS framework achieves near-perfect accuracy while revealing critical vulnerabilities in replay buffers that can be exploited by adversarial attacks.
Selecting the right VLM adaptation strategy can drastically improve performance in federated learning for remote sensing, balancing efficiency and generalization.
Rollout generation can be transformed from a static process into a dynamic, learning-driven strategy that adapts to policy changes in real-time.
SpecRoll achieves up to 2.15x faster generation in RL rollouts by cleverly balancing fast and slow adaptation strategies.
Achieving over 83% uplink data savings in federated learning while maintaining competitive accuracy highlights a new frontier in optimizing communication costs.
SPOT redefines on-policy distillation by ensuring that probing decisions directly enhance downstream reasoning performance, not just teacher alignment.
Muon optimizer's unique advantage lies in its ability to enhance token efficiency specifically when applied to the output projection, challenging assumptions about conditioning in state-space models.
Omega-S retains over 84% of a model's original capabilities during fine-tuning, significantly outperforming conventional regularization techniques.
Energy-aware DNN design can now be optimized offline, paving the way for truly autonomous AI on intermittent power sources.
Achieving high task performance with ultra-low-rank adapters, SALT recovers accuracy while slashing memory usage by up to 16x.
Achieving MCMC-level accuracy at a fraction of the computational cost, this method revolutionizes how we approach Bayesian inference in cognitive decision-making models.
Subword composition methods can drastically reduce continued pre-training steps while enhancing accuracy in LLM vocabulary extensions.
A unified criterion for Markov chain convergence reveals that traditional assumptions can be bypassed, streamlining the path to convergence analysis.
A smaller batch size and larger learning rate can lead to flatter minima in SAM, revealing a critical trade-off in hyperparameter tuning that impacts generalization.
Noise-aware shrinkage can enhance the utility of differentially private fine-tuning, outperforming traditional methods without compromising privacy or efficiency.
Bridging the gap between ANNs and SNNs could revolutionize federated learning on resource-constrained devices, achieving high accuracy without sacrificing efficiency.
Achieving near-uncompressed accuracy with significantly lower communication costs, FraQ revolutionizes federated LoRA efficiency.
Early training telemetry can predict deep learning outcomes with over 99% accuracy after just one epoch, potentially saving significant compute resources.
Pruning the vocabulary of multilingual models can lead to a 60% memory savings without sacrificing translation quality, challenging the need for large vocabularies in MNMT.
MoEGen achieves instance-specific adaptations without the storage burden of full LoRA experts, revolutionizing how we think about parameter-efficient fine-tuning.
Dynamic adaptation in vision-language models can significantly boost performance while cutting down computational costs.
ComFuse achieves a breakthrough in GPU compilation by enabling concurrent execution of memory-intensive and compute-intensive operations, leading to significantly improved performance in complex workloads.
Sparse rewards can be effectively optimized without losing the advantages of fine-grained feedback through a novel two-stage training approach.
Achieving exact acceleration of neural-network potentials without retraining could revolutionize molecular dynamics simulations by enhancing efficiency without sacrificing accuracy.
Achieving up to 1,002× faster power analysis, DiffPower revolutionizes design optimization and enables unprecedented efficiency in power management tasks.
SlimVLM achieves unprecedented efficiency in Vision-Language Models by intelligently pruning redundant visual tokens without sacrificing performance.
Targeted optimization of normalization affine parameters can dramatically enhance low-bit quantization performance, breaking the limits of conventional training methods.
ALiBi positional encoding can blind attention heads, drastically impairing token retrieval without impacting standard performance metrics.
MoE dLLMs can outperform leading models with significantly fewer training tokens, challenging assumptions about data efficiency in large-scale language models.
FedCARE achieves up to 12.5% better predictive accuracy by enabling personalised model adaptations in federated healthcare settings without compromising data privacy.
Incorporating incremental knowledge into hierarchical reinforcement learning can drastically enhance sample efficiency in challenging environments with sparse rewards.
MFU can predict GPU power consumption with remarkable accuracy, achieving a 1% error rate in compute-bound LLM training scenarios.
Achieving a 1.43x speedup in MoE training while slashing communication costs by up to 74% could redefine efficiency benchmarks in large-scale LLM training.
SpecDrop shows that fixed, category-conditioned routing can outperform learned routers, achieving up to 6.53% accuracy gains without additional parameters.
Offline KD can achieve the same training loss as online methods while being 29% faster, revolutionizing how we approach model distillation efficiency.
Simply adding more multimodal environments can hinder agent performance, but targeted diversity and structured difficulty can transform training outcomes.
Biased client selections in federated learning can severely degrade model accuracy, but a new scoring method offers a way to optimize client contributions while preserving privacy.
KANs outperform MLPs in FTN BPSK detection, achieving 18.6 times lower bit error rates with drastically fewer parameters.
A new Gaussian approximation reveals that ridge regression can be optimized for finite samples, leading to improved parameter selection strategies that significantly enhance prediction accuracy.
Achieving a Wasserstein mixing time that scales with $\sqrt{d}/\varepsilon$ could revolutionize the efficiency of sampling algorithms in high-dimensional settings.
CoPES recovers 92% of the performance gains of traditional methods while slashing GPU memory requirements to less than one-eighth.
BRiG-AFA achieves up to 10.20 percentage points improvement in accuracy over greedy methods by leveraging a novel risk-to-go learning approach for active feature acquisition.
CRIP selectively fuses features from clients based on representational similarity, leading to superior performance in one-shot federated learning under extreme domain heterogeneity.
CoRe-GNN achieves competitive accuracy on long-range graph tasks while dramatically improving memory efficiency through a novel dual-message passing approach.
A single shared intervention in the query-key channel can stabilize low-precision transformer training, eliminating the need for individual fault repairs.
Leveraging realized input information can accelerate evolutionary optimization processes, achieving faster convergence in complex tasks.
Open-DiffLoco achieves robust quadruped locomotion with minimal training complexity, outperforming traditional reinforcement learning methods.
AOS-R accelerates convergence by 43% while improving accuracy on 75% of tested model-dataset pairs, reshaping the optimizer landscape for deep learning.
Convex neural energy elements transform neural operators into reusable components that guarantee stability and accuracy, even in complex geometries.
Finite-time convergence rates for risk-sensitive reinforcement learning algorithms reveal that model-free approaches can achieve robust performance without complex parameter tuning.
A hands-on framework that transforms AI education in power systems, making complex concepts accessible to newcomers and practitioners alike.
Chunked Muon accelerates Diffusion Transformer training by over 2x while achieving state-of-the-art performance, solving critical convergence issues.
Anomaly detection just got faster—CARE achieves up to 4.8x inference speedup without sacrificing accuracy by intelligently filtering normal data.