Search papers, labs, and topics across Lattice.
41 papers published across 3 labs.
Scaling native multimodal pre-training reveals that text-heavy data mixtures require larger models for optimal efficiency, challenging conventional resource allocation strategies.
Larger parameterized quantum circuits can enhance generalization performance, defying traditional expectations of model degradation.
Fine-tuning on bad advice can trigger a pre-existing persona structure, leading to unexpected and broad misalignment in language models.
Expanding Flow Maps redefine generative modeling by allowing output size to be a learned, controllable parameter, enabling unprecedented flexibility in generation tasks.
By bounding the draft's KV working set to a constant, Windowed-MTP slashes decoding costs by up to 44% at million-token contexts, redefining efficiency in autoregressive models.
Scaling native multimodal pre-training reveals that text-heavy data mixtures require larger models for optimal efficiency, challenging conventional resource allocation strategies.
Larger parameterized quantum circuits can enhance generalization performance, defying traditional expectations of model degradation.
Fine-tuning on bad advice can trigger a pre-existing persona structure, leading to unexpected and broad misalignment in language models.
Expanding Flow Maps redefine generative modeling by allowing output size to be a learned, controllable parameter, enabling unprecedented flexibility in generation tasks.
By bounding the draft's KV working set to a constant, Windowed-MTP slashes decoding costs by up to 44% at million-token contexts, redefining efficiency in autoregressive models.
Incorporating local context into discrete flow matching can reduce generative perplexity by up to 63%, revolutionizing how we approach generative modeling.
TTEL not only boosts reasoning efficiency but also slashes token generation by nearly 50%, setting a new standard for inference-time computation in language models.
Self-Output Fine-Tuning (SOFT) effectively curbs the butterfly effect in weather forecasting, transforming initial prediction errors into a manageable calibration process.
Progressive cramming reveals that achieving perfect reconstruction is not enough, as it can lead to substantial performance degradation in downstream tasks.
Pruning less promising seed trajectories early can significantly boost image generation quality without increasing compute costs.
SLPO enables latent reasoning models to leverage outcome-reward learning, dramatically enhancing their performance and efficiency in complex reasoning tasks.
StatLoRA leverages statistical hypothesis testing to optimize rank allocation in LoRA, achieving superior performance while maintaining efficiency.
Treating AI models as a single, scalable architecture may be a fundamental mistake, as distinct cognitive tasks require qualitatively different structures for optimal performance.
Clustering LLM inputs can cut inference costs and latency by 50-fold while maintaining personalization, a game-changer for scaling AI applications.
Treating LLM-surprisal as interchangeable risks masking crucial representational choices that shape model behavior.
Solar Open 2 outperforms its predecessors and competitors with a groundbreaking 1M-token context window, redefining the capabilities of large language models in agentic tasks.
Collective electronic entanglement can be achieved without the typical O(1/N) dilution penalty, unlocking new possibilities for scalable quantum technologies.
Emergent task decomposition in robot policies reveals that learned experts can specialize and be reused, challenging traditional notions of task hierarchy in AI.
Coupled memory rewriting during model pondering leads to a surprising learning-speed penalty without sacrificing ultimate recall performance. WHY_IT MATTERS: This insight challenges existing assumptions about memory management in neural networks and could inform future designs for more efficient learning systems.
PortLLM patches retain their effectiveness over time, eliminating the need for constant fine-tuning even as base models evolve.
Achieving desired accuracy in optimization may require significantly less memory than previously thought, with a critical phase transition in online iteration budgets that can redefine efficiency benchmarks.
Multi-Head Attention Residuals achieve superior validation loss by allowing Transformers to leverage multiple attention heads, revealing that subspace disagreement is a key factor in model performance.
A minimalist generative model achieves state-of-the-art performance without the complexity of iterative denoising or advanced architectures.
Spectral Higher-Order Neural Networks can drastically reduce computational costs while improving performance on notoriously difficult tasks like N-bit parity.
Stochastic bandit convex optimization is fundamentally harder than linear bandits, with a new lower bound that reveals a surprising increase in regret complexity.
Relative positional encodings not only enable extrapolation in transformers but also reveal a profound connection between implicit bias and sequence length generalization.
Compact models like Athena-Brain-8B can outperform larger counterparts in embodied tasks while maintaining strong general intelligence.
AGI won't emerge from mere architectural tweaks; it demands a nuanced understanding of twenty-three interdependent constraints across multiple domains.
Achieving up to 2.64x speedup in Mixture-of-Experts execution by cleverly overlapping computation and communication could redefine efficiency benchmarks in large-scale AI models.
Structural generalization is mathematically unattainable for pure Transformers, revealing a critical limitation in their learning capabilities compared to neuro-symbolic systems.
Hypernetworks can achieve reliable out-of-distribution generalization and scaling benefits that traditional adaptation methods struggle to match.
Task-specific transformer architectures can dramatically boost learning efficiency but often at the cost of versatility, challenging the notion that one design fits all.
Relational hidden states in neural networks can unlock emergent planning capabilities in model-free reinforcement learning, challenging traditional distinctions between model-based and model-free methods.
Attention-only transformers can match standard architectures in performance when resource allocation favors attention depth, challenging assumptions about the necessity of feed-forward layers.
Achieving up to 17× faster inference for long-context LLMs without compromising output quality could redefine efficiency standards in AI applications.
L1 augmented attention achieves a remarkable 14.5% reduction in perplexity by integrating L1 geometry into Transformer models, challenging the dominance of traditional dot product methods.
Systemic AI risks are not just technical failures but emergent societal threats that can cascade through interconnected systems, challenging our governance frameworks.
ExpertPlex slashes over 95% of duplicate model weights and boosts goodput by up to 2.01 times, transforming how we serve large language models.
ThAME achieves a staggering 15.7x speedup in MoE inference, revolutionizing the efficiency of Large Language Models.
Analyzing Transformers through continuous stochastic differential geometry reveals surprising predictive insights into their stability limits and optimization dynamics.
Surprising insights emerge as low-parameter models accurately replicate the behavior of complex LLM-agent societies, challenging the notion that high computational power is essential for meaningful simulations. WHY_IT MATTERS: This approach could democratize access to agent-based modeling, allowing researchers with limited resources to explore complex interactions in LLM societies without sacrificing accuracy.