Search papers, labs, and topics across Lattice.
35 papers published across 4 labs.
Gambit achieves up to 6.7% higher accuracy and over 2x throughput by intelligently reallocating compute resources during reasoning.
The Skaling law slashes loss estimation errors by up to 3x and reduces compute needs by 90% for reliable model performance predictions.
Scaling language models with interpretability as a core constraint reveals that they can become more capable and understandable simultaneously, defying traditional trade-off assumptions.
Adapting LLM inference scheduling to bursty traffic can boost throughput by leveraging real-time request intensity estimation.
Self-PreTraining boosts transformer accuracy in medical time series by up to 6 percentage points, even with limited data.
Gambit achieves up to 6.7% higher accuracy and over 2x throughput by intelligently reallocating compute resources during reasoning.
The Skaling law slashes loss estimation errors by up to 3x and reduces compute needs by 90% for reliable model performance predictions.
Scaling language models with interpretability as a core constraint reveals that they can become more capable and understandable simultaneously, defying traditional trade-off assumptions.
Adapting LLM inference scheduling to bursty traffic can boost throughput by leveraging real-time request intensity estimation.
Self-PreTraining boosts transformer accuracy in medical time series by up to 6 percentage points, even with limited data.
Achieving an FID of 1.45 in just 600 epochs, Energy-Guided Flow Matching redefines efficiency in high-quality image generation without the need for extensive model adaptations.
U-Nets unexpectedly show greater robustness to resolution changes than anticipated, challenging assumptions about neural operator architectures in inverse imaging.
Optimal training time in gradual adaptation scales inversely with the number of tasks, revealing a critical balance for effective learning.
Tiny transformers can achieve impressive reasoning capabilities through protoreasoning, challenging assumptions about the necessity of larger models for effective step-by-step reasoning.
Attention-free models can outperform transformers in language generation, especially at smaller dataset scales, challenging the dominance of attention mechanisms in NLP.
Pretraining on behavioral data can boost neural decoding performance by over 11%, making it a game-changer for brain-computer interface development.
Elbow-based routing can cut inference latency by over 5% in MoE models while preserving accuracy, revolutionizing expert selection efficiency.
CARVE slashes visual token usage by 80% while maintaining nearly all performance, reshaping how we approach token allocation in 3D medical imaging.
Smaller foundation models can rival larger counterparts in familiar tasks but struggle with generalization, revealing a critical trade-off in cognitive modeling.
Replacing just 12% of traditional training data with OctoLong's curated code contexts leads to substantial improvements in long-range retrieval and state tracking for language models.
K-EXAONE 2.0 achieves over three times the capacity of its predecessor while enhancing multilingual capabilities and long-context reasoning.
Test-time scaling can significantly enhance LLM reasoning capabilities, but without clear protocols, results are often incomparable and misleading.
Logic pre-pretraining accelerates language model skill acquisition by 36B tokens while enhancing compressibility, revealing a new path for efficient model training.
Reusing KV caches during model swaps can retain up to 98% of prefill accuracy while running 2.7-25x faster than traditional methods.
Standard music transformers lose equivariance as they scale, but the Equivariant Music Transformer recovers this crucial property, enhancing generative performance.
The smallest eigenvalue of ReLU NTK Gram matrices is tightly bounded, revealing critical geometric insights into neural network behavior.
Test error in gradient boosting decision trees can peak and then drop as the split-candidate budget increases, revealing a surprising double descent effect tied to the geometry of candidate paths.
Quantum circuits can sample distributions that classical language models cannot, revealing a profound computational divide that challenges the supremacy of classical architectures.
MoE dLLMs can outperform leading models with significantly fewer training tokens, challenging assumptions about data efficiency in large-scale language models.
Task-specific optimization in multimodal models can significantly boost performance by dynamically reallocating shared computation resources between vision and language components.
AdaMX slashes accuracy loss in low-bit LLM inference by 83% while keeping energy costs minimal.
Maglev achieves superior validation loss and performance benchmarks by cleverly combining full attention with sliding-window mechanisms, all while conserving memory.
Geometry-guided width allocation in Transformers can lead to substantial improvements in model performance, reducing validation loss more effectively than traditional methods.
Hierarchical local attention in TextNCA reveals that the arrangement of attention windows can dramatically influence language modeling performance, even more than the model's iterative nature.
LLMs' predictive accuracy plummets as data dimensionality increases, unlike classical models that maintain or improve performance.
DART achieves a remarkable 75% reduction in inference cache size while enhancing associative recall in long-context sequence modeling.
Achieving 94.5% attribution accuracy while processing novels 20 times faster than traditional methods reveals a breakthrough in quotation attribution efficiency.
REFLEX redefines MoE inference by aligning expert computation with the distinct refinement needs of tokens, achieving efficiency gains without compromising quality.
SmartGR achieves an 8.6% boost in recommendation performance while slashing inference time by over 2.3 times, tackling unique challenges in generative recommendation systems.
HiResNets achieve human-like foveation in video recognition, enabling efficient processing of high-resolution content without the typical quadratic memory costs.