Search papers, labs, and topics across Lattice.
31 papers published across 9 labs.
Loopie outperforms traditional Transformers by leveraging a novel architecture that maximizes efficiency without sacrificing reasoning power, achieving gold-medal performance in competitive settings.
Heterogeneous AI agent networks can evolve to outperform homogeneous counterparts, revealing a new scaling law for collaboration that defies traditional strength metrics.
Expanding tokenizers in-place can lead to up to 4x fewer tokens for new languages, drastically improving decoding efficiency for multilingual models.
Coordinated scaling of Behavior Foundation Models can enhance humanoid robot control performance, achieving up to 82% error reduction in real-world tasks.
LongStraw enables RL post-training with over 2 million tokens on a fixed GPU budget, pushing the boundaries of context length in AI applications.
Loopie outperforms traditional Transformers by leveraging a novel architecture that maximizes efficiency without sacrificing reasoning power, achieving gold-medal performance in competitive settings.
Heterogeneous AI agent networks can evolve to outperform homogeneous counterparts, revealing a new scaling law for collaboration that defies traditional strength metrics.
Expanding tokenizers in-place can lead to up to 4x fewer tokens for new languages, drastically improving decoding efficiency for multilingual models.
Coordinated scaling of Behavior Foundation Models can enhance humanoid robot control performance, achieving up to 82% error reduction in real-world tasks.
LongStraw enables RL post-training with over 2 million tokens on a fixed GPU budget, pushing the boundaries of context length in AI applications.
Scaling visuomotor context to 8K timesteps enables robots to master complex tasks and adapt in real-time, outperforming previous models by a staggering margin.
Scaling Hyper-Connections beyond four streams is now feasible, yielding substantial performance gains without prohibitive costs.
A task's representability determines not just generalization but also memorization, leading to a binary success-failure outcome in training.
ExTernD achieves near-bf16 accuracy in ternary quantization, outperforming traditional methods by correcting quantization errors with expanded rank components.
Local redundancy predicts neural network adaptability more accurately than traditional metrics, revolutionizing how we assess model plasticity.
Large models can achieve sudden performance leaps without retraining, unlocking their hidden capabilities through innovative shortcut modifications.
CIMERA achieves up to 25x energy efficiency improvements for LLM inference, revolutionizing how we approach resource-constrained environments.
TmallGS redefines e-commerce search ranking by optimizing feature representation and interaction, resulting in substantial performance gains over traditional models.
GFlowRL achieves unprecedented stability and performance in large language models by eliminating the problematic learned partition function, setting a new standard for GFlowNet-style reinforcement learning.
NexForge transforms the landscape of LLM training by synthesizing 43.2K tasks, propelling model performance to unprecedented levels without the need for domain-specific infrastructure.
DRIFT slashes communication time from 97% to under 6% of the forward-pass time, enabling unprecedented speedups in distributed Fourier Neural Operators.
Looping Transformers can achieve better performance with shared parameter updates, revealing that scaling rules must adapt to parameter visits for stable recurrent depth.
Collective training of large language models can now be achieved with consumer GPUs, making frontier AI development accessible to a broader community.
Muon's touted superiority in large-scale training may not hold up under controlled conditions, challenging its status as a go-to optimizer.
Smaller multimodal emotion models can outperform their larger counterparts, achieving state-of-the-art results with significantly improved efficiency.
Achieving 2-3x near-lossless compression of KV caches could revolutionize the efficiency of transformer inference without sacrificing performance.
MESH achieves a 14x improvement in scaling for fresh content, transforming how retrieval systems handle diverse content tiers.
Scaling zero RL to a trillion parameters reveals that models can spontaneously develop advanced cognitive behaviors, making traditional heuristics obsolete.
Transforming our understanding of Transformers, this work reveals that learnability may be as crucial as expressivity in optimizing large language models.
ARMT-augmented models can handle inputs far beyond their original context limits while using 30% less compute, revolutionizing efficiency in long-context processing.
UMoE transforms underperforming expert pools into high-performing domain-specific models, achieving up to 6.0 points improvement on key benchmarks without increasing computational costs.
Training dynamics of Transformers can be reduced to a low-dimensional manifold, revealing how inductive reasoning emerges from data statistics and model initialization.
Dimensionality collapse is a key precursor to grokking, and manipulating representation geometry can dramatically speed up generalization in neural networks.
Data synergy can either amplify or diminish model performance, revealing that the right dataset combinations are crucial for optimal language model training.
Energy calculus reveals that energy optimization can be systematically composed, transforming how we approach energy efficiency in AI systems.
Achieving state-of-the-art navigation performance with only 0.58M trainable parameters challenges the notion that bigger models are always better.