Search papers, labs, and topics across Lattice.
28 papers published across 4 labs.
Discarding irrelevant reasoning tokens can make language models three times faster at test time without sacrificing performance.
The required rank for effective low-rank adaptation in Transformer attention can be much smaller than expected, especially under softmax saturation conditions.
Weight-decay pulses can dictate the timing of generalization in neural networks, revealing a surprising canalization effect that precedes visible performance improvements.
Increasing model size trumps inference compute for grammar-constrained text-to-SQL tasks, with beam search proving superior to sample+vote under matched budgets.
Romanization during pretraining can dramatically enhance multilingual model performance, outpacing traditional text-based approaches.
Discarding irrelevant reasoning tokens can make language models three times faster at test time without sacrificing performance.
The required rank for effective low-rank adaptation in Transformer attention can be much smaller than expected, especially under softmax saturation conditions.
Weight-decay pulses can dictate the timing of generalization in neural networks, revealing a surprising canalization effect that precedes visible performance improvements.
Increasing model size trumps inference compute for grammar-constrained text-to-SQL tasks, with beam search proving superior to sample+vote under matched budgets.
Romanization during pretraining can dramatically enhance multilingual model performance, outpacing traditional text-based approaches.
APT accelerates high-resolution diffusion models by up to 8.16× while enhancing energy efficiency, revolutionizing the feasibility of real-time generative AI applications.
TransRetrieval boosts recommendation recall by over 19 points while slashing computational costs by 85%, proving that smarter aggregation can outperform brute force scaling.
Visual intelligence could be the key to unlocking AGI, challenging the dominance of language-based models in the quest for general intelligence.
Recursive self-improvement in LLMs can lead to unprecedented performance, with Meta^n surpassing all prior agents on challenging benchmarks by leveraging fixed meta-operations.
Matching effective learning rates across different training setups leads to remarkably consistent loss trajectories, challenging conventional wisdom about learning rate variability.
Progressive growth strategies can significantly bias neural network training towards flatter loss landscapes, but flatter does not always mean better performance.
Generative models exhibit distinct phases of learning that reveal how they balance generalization and memorization in high-dimensional spaces.
ALPHABET achieves Bayes oracle performance with a mere 6,437 parameters, revolutionizing efficiency in sequence modeling.
Degrading LLM service to save costs can actually backfire, leading to increased churn and inflated demand during peak times.
RoI slashes the parameter overhead of semi-structured sparsity by up to 8.75 times, paving the way for more efficient large language model deployment.
RIBOSPAN achieves unprecedented RNA modeling capabilities by maintaining high-resolution context across full-length transcripts, outperforming existing models in both reconstruction and representation quality.
Amortizing distillation across model size and variant can produce a continuum of optimized LLMs with a single distillation run, revolutionizing model deployment efficiency.
CNNs outperform larger models in real-time plasma equilibrium prediction, achieving a remarkable balance of speed and accuracy essential for tokamak control.
Achieving up to a 72.81× speedup in long-sequence inference could redefine efficiency standards for Diffusion Language Models.
Tokenization can unlock sustained performance gains in user representation learning, even as raw data scaling hits diminishing returns.
The shift from ad-hoc to standardized carbon accounting in climate modeling reveals significant discrepancies in energy and emissions reporting that could reshape climate research practices.
Channel selection in convolutional networks can dramatically reduce computational costs while maintaining or exceeding performance benchmarks.
Intermediate knowledge distillation can dramatically improve performance in data-scarce environments, turning the conventional wisdom on its head.
Reclaiming memory from idle KV cache during decode phases yields negligible latency improvements, challenging assumptions about prefill chunk sizes in LLM serving.
Predictive delta prefetching can achieve up to 12x faster LLM inference on edge devices, even when models exceed memory limits.
Deeper partitions in federated fine-tuning may boost throughput and privacy, but they can also cause LLM performance to collapse catastrophically.
The synchronization tax can consume over 50% of communication time in GPU scale-up domains, fundamentally challenging our understanding of bandwidth scaling.
NOVA's innovative architecture delivers 4.5x higher throughput and 69.8% lower latency for hybrid LLMs, pushing the boundaries of memory processing efficiency.