Search papers, labs, and topics across Lattice.
Power-law relationships in model scaling, emergent capabilities at scale, and compute-optimal training.
#23 of 24
5
Adapting LLM inference scheduling to bursty traffic can boost throughput by leveraging real-time request intensity estimation.
Self-PreTraining boosts transformer accuracy in medical time series by up to 6 percentage points, even with limited data.
Energy-Guided Flow Matching achieves a remarkable FID of 1.45 on ImageNet with fewer epochs, redefining efficiency in high-dimensional generative modeling.
U-Nets unexpectedly show greater robustness to resolution changes than anticipated, challenging assumptions about neural operator architectures in inverse imaging.
Optimal training time in gradual adaptation scales inversely with the number of tasks, revealing a critical balance for effective learning.
Tiny transformers can achieve impressive reasoning capabilities through protoreasoning, challenging assumptions about the necessity of larger models for effective step-by-step reasoning.
Attention-free models can outperform transformers in language generation, especially at smaller dataset scales, challenging the dominance of attention mechanisms in NLP.
Pretraining on behavioral data can boost neural decoding performance by over 11%, making it a game-changer for brain-computer interface development.
Elbow-based routing can cut inference latency by over 5% in MoE models while preserving accuracy, revolutionizing expert selection efficiency.
CARVE slashes visual token usage by 80% while maintaining nearly all performance, reshaping how we approach token allocation in 3D medical imaging.
K-EXAONE 2.0 achieves over three times the capacity of its predecessor while enhancing multilingual capabilities and context understanding.
Smaller models can rival larger ones in familiar tasks, but only larger models excel at generalizing to new challenges.
Replacing just 12% of traditional training data with OctoLong's curated code contexts leads to substantial improvements in long-range retrieval and state tracking for language models.
Test-time scaling can significantly enhance LLM reasoning capabilities, but without clear protocols, results are often incomparable and misleading.
Logic pre-pretraining accelerates language model skill acquisition by 36B tokens while enhancing compressibility, revealing a new path for efficient model training.
Reusing KV caches during model swaps can retain up to 98% of prefill accuracy while running 2.7-25x faster than traditional methods.
Standard music transformers lose equivariance as they scale, but the Equivariant Music Transformer recovers this crucial property, enhancing generative performance.
The smallest eigenvalue of ReLU NTK Gram matrices is tightly bounded, revealing critical geometric insights into neural network behavior.
Test error in gradient boosting decision trees can peak and then drop as the split-candidate budget increases, revealing a surprising double descent effect tied to the geometry of candidate paths.
Quantum circuits can sample distributions that classical language models cannot, revealing a profound computational divide that challenges the supremacy of classical architectures.
MoE dLLMs can outperform leading models with significantly fewer training tokens, challenging assumptions about data efficiency in large-scale language models.
Task-specific optimization in multimodal models can significantly boost performance by dynamically reallocating shared computation resources between vision and language components.
AdaMX slashes accuracy loss in low-bit LLM inference by 83% while keeping energy costs minimal.
Geometry-guided width allocation in Transformers can lead to substantial improvements in model performance, reducing validation loss more effectively than traditional methods.
Hierarchical local attention in TextNCA reveals that the arrangement of attention windows can dramatically influence language modeling performance, even more than the model's iterative nature.
LLMs' predictive accuracy plummets as data dimensionality increases, unlike classical models that maintain or improve performance.
DART achieves a remarkable 75% reduction in inference cache size while enhancing associative recall in long-context sequence modeling.
Achieving 94.5% attribution accuracy while processing novels 20 times faster than traditional methods reveals a breakthrough in quotation attribution efficiency.
REFLEX redefines MoE inference by aligning expert computation with the distinct refinement needs of tokens, achieving efficiency gains without compromising quality.
SmartGR achieves an 8.6% boost in recommendation performance while slashing inference time by over 2.3 times, tackling unique challenges in generative recommendation systems.
HiResNets achieve human-like foveation in video recognition, enabling efficient processing of high-resolution content without the typical quadratic memory costs.
Converged diffusion loss in visual generation improves linearly with structured language, leading to a new training paradigm that outperforms both open-weight and closed-weight models.
Repeated sampling outperforms self-reflection methods in language models, revealing a critical flaw in approaches that rely on self-critique and rewriting.
Diminishing returns in inference-time scaling reveal that more computation doesn't always equate to better performance in local computer-use agents.
ROCS can triple the throughput of recommendation systems without sacrificing prediction accuracy, revolutionizing how we handle user requests in large-scale settings.
CCFormer delivers a 3.57% increase in click-through rates and a 1.71% boost in advertising revenue, all while cutting model training time by over 2x.
Chimera achieves 7.3x compute efficiency over traditional models while enabling zero-shot extrapolation from short video clips to significantly longer sequences.
Sequence models outperform LLMs in predictive process monitoring, challenging the assumption that larger models always yield better results.
Agents can now autonomously teach themselves creative skills using high-quality human texts, bypassing the need for expensive human feedback.
CoMem achieves a 7.83x prefill speedup and drastically reduces memory usage while maintaining high performance on long-context tasks, challenging conventional memory management in LLMs.
SCSE transforms recurrent computation in Looped Transformers by ensuring that deviations from a learned anchor enhance performance without compromising stability.
Achieving high-fidelity 3D generation with just 1.5% of the training data could revolutionize resource allocation in 3D modeling.
Expert subspaces in MoE models overlap significantly, yet selected routes yield better token representations than their strongest unselected counterparts, challenging traditional views on redundancy.
Independently scaling long-term memory in language models can yield better performance with fewer parameters than simply increasing model size.
Shifting the burden of intelligence from runtime search to a belief-based framework allows for professional-level Go play on consumer hardware without the need for extensive MCTS.
Predicting the next token's KV entries can boost long-context LLM throughput by over 2.5 times without sacrificing latency or quality.
Knowledge transfer in MLLM fusion is not a blanket inheritance but a selective process favoring high-level reasoning over perception.
A photonic memory architecture can slash LLM inference latency by over 50%, paving the way for scalable, high-performance KV cache management.
BM25 outperforms all competitors at scale, revealing that lexical retrieval is the most efficient choice for large corpus sizes.
Exploration as a pretraining axis can boost generative model performance by up to 36%, revolutionizing how we approach end-to-end generation.