Search papers, labs, and topics across Lattice.
13
0
12
18
Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.
Models can autonomously bootstrap their own dense token-level supervision simply by extrapolating the trajectory of their own RL updates away from a trailing checkpoint.
Expanding an agent's harness with new tools and skills often degrades its performance on tasks it previously solved, exposing a critical failure mode of "harness-induced forgetting."
Despite advances in AI, top-performing agents struggle with real-world data science, achieving only 56.70% success in automating complex workflows.
Visualization code editing is a tough nut to crack, with leading models only achieving a 74.46% pass rate on complex multimodal tasks.
Token likelihood changes can mislead researchers into overestimating the value of intermediate actions in self-distillation, with experiments showing near-chance performance in scoring effectiveness.
Achieving 4.7x to 8.2x higher throughput for trillion-parameter MoE models could redefine the limits of large-scale model training.
Automatically generated Multi-Agent Systems are not only outperformed by Single-Agent Systems but also exhibit architectural inefficiencies that challenge the very foundations of multi-agent design principles.
OrchRM slashes training costs while boosting orchestration accuracy, proving that self-supervised reward modeling can revolutionize multi-agent coordination.
LVLM judges, despite excelling in English, exhibit surprisingly inconsistent and unreliable behavior when evaluating content in other languages, revealing a critical blind spot in current alignment and evaluation pipelines.
LLMs can generate syntactically correct tests, but their ability to *reason* about code faults is surprisingly poor, hindering autonomous debugging.
SkillOrchestra slashes the learning costs of AI agent orchestration by up to 700x while improving performance by explicitly modeling agent skills and costs, offering a more scalable and interpretable alternative to RL-based methods.
Reference-guided LLM evaluators can boost alignment in non-verifiable domains, enabling self-improvement to rival reward model training.