Search papers, labs, and topics across Lattice.
39 papers published across 5 labs.
Minor architectural tweaks can lead to a staggering 47% drop in long context performance, challenging assumptions about model design.
Surprisingly, larger LLMs benefit from increased repetition of high-quality domain data, challenging conventional wisdom about data diversity in training.
Achieving high-fidelity 3D object generation with up to 300 parts, MegaParts redefines the limits of part-aware modeling through token-efficient autoregressive techniques.
Forecast collapse in time-series models reveals a critical calibration-ranking tradeoff that could mislead financial decision-making.
RMM reveals that optimizing attention-side computations can lead to substantial runtime gains in Transformer models without sacrificing accuracy.
Surprisingly, larger LLMs benefit from increased repetition of high-quality domain data, challenging conventional wisdom about data diversity in training.
Achieving high-fidelity 3D object generation with up to 300 parts, MegaParts redefines the limits of part-aware modeling through token-efficient autoregressive techniques.
Forecast collapse in time-series models reveals a critical calibration-ranking tradeoff that could mislead financial decision-making.
RMM reveals that optimizing attention-side computations can lead to substantial runtime gains in Transformer models without sacrificing accuracy.
A universal quadratic model reveals that diverse neural architectures share a common training dynamic, leading to predictable power law behaviors in learning.
Independently trained depth slices can be recombined to match the performance of monolithic models, revealing a new avenue for efficient language model training.
Post-norm normalization significantly enhances performance in LLMs when depth is introduced through a curriculum, outperforming pre-norm by an order of magnitude.
Traditional probing methods fail to reveal the true memorization capabilities of large code LLMs, leading to inflated performance scores that obscure their genuine understanding.
DINOv2-style pretraining outperforms other SSL methods in resource-limited settings, but combining it with video objectives reveals critical trade-offs in performance.
A single generalist model can achieve state-of-the-art performance in Referring Expression Comprehension while generalizing across diverse datasets without fine-tuning.
Operator-level scaling can reduce GPU usage by over a third while maintaining strict service level objectives for LLMs.
Reducing parameter redundancy in KANs, HYDRA achieves superior predictive performance while enhancing interpretability in hyperbolic spaces.
TREX enables compact models to match or exceed the performance of large foundation models while dramatically speeding up inference times.
Hyperparameter tuning is the hidden key that unlocks the scaling laws in small models, revealing insights previously thought to be exclusive to larger architectures.
OPD may improve sampling efficiency, but it risks making previously solvable problems unsolvable, challenging the notion of true capability expansion in LLMs.
Autonomous Semantic Solitons emerge from a new dynamical framework, enabling LLMs to generate diverse outputs while avoiding stagnation.
XYZFlow achieves up to 8.5X speed improvements in generative modeling without compromising image quality, redefining the efficiency landscape in high-fidelity image generation.
Training LLMs with long contexts can paradoxically weaken their ability to retain knowledge, leading to poorer performance when context is not available.
Pre-attention spikes and inter-spike plateaus reveal a surprising organization in hybrid linear attention models that could redefine our understanding of activation dynamics in LLMs.
The optimal vocabulary size for LLMs can shift dramatically based on serving conditions, with potential divergences from training norms by up to 16x.
Nonlinear $β$-VAEs reveal a surprising tradeoff where deeper models concentrate utility but compromise the fidelity of less significant dimensions.
Machine learning models can reduce path loss prediction errors in LPWAN by over 30% compared to traditional methods, revolutionizing network planning for smart cities.
Optimal sparsity in sparse MoE models emerges only when considering the cluster's systems constraints, challenging traditional compute-centric design paradigms.
Inverse-distance attention can achieve exact retrieval with constant resources, outperforming softmax's logarithmic scaling in complexity.
Sharing expert knowledge before routing leads to a substantial reduction in computational demand while boosting model performance.
Retrieval-augmented reasoning can boost LRM accuracy by up to 60% during test-time scaling, transforming how models handle complex problem-solving.
Minor architectural tweaks can lead to a staggering 47% drop in long context performance, challenging assumptions about model design.
Native encoding outperforms external-function serialization in thread scaling for TDVRPTW, yielding better solutions and stability across multiple threads.
Replacing fixed attention mechanisms with a learned, input-generated operator reveals surprising invariance properties that could redefine our understanding of attention in LLMs.
Motif 3 achieves unprecedented efficiency and performance in language modeling by leveraging a novel Mixture-of-Experts architecture that activates only a fraction of its parameters per token.
Task arithmetic reveals that not all fine-tuning combinations yield predictable model behaviors, challenging assumptions about parameter addition's reliability.
A random transformer can achieve universal approximation without any pretraining, challenging conventional beliefs about model training requirements.
ICBQ not only cuts perplexity in quantized models but also salvages performance where traditional methods falter, redefining efficiency in model compression.
Latent feedback in full-bandwidth transformers enables deeper contextual understanding without sacrificing efficiency, leading to improved performance across multiple tasks.
A single checkpoint can now adapt to any model size, streamlining the deployment of elastic retrieval systems and achieving faster performance without sacrificing quality.
The evolution of Mixture-of-Experts architectures reveals a critical shift towards decoupling semantic routing from computational budgets, reshaping our understanding of model efficiency.
DistillCache retains over 94% accuracy on long-context tasks while slashing memory usage by 75%, outperforming traditional methods and setting a new standard for efficient LLM inference.
Gambit achieves up to 6.7% higher accuracy and over 2x throughput by intelligently reallocating compute resources during reasoning.
The Skaling law slashes loss estimation errors by up to 3x and reduces compute needs by 90% for reliable model performance predictions.