Search papers, labs, and topics across Lattice.
15
1
16
7
Natural-language critiques can transform how we evaluate and optimize song generation models, leading to more human-aligned outputs.
LLMs struggle with personal information retrieval, with the best model only achieving 57.3% accuracy on a new benchmark designed to evaluate mobile assistant capabilities.
IACM-RL reduces infinite loops and stale context errors by proactively managing dynamic user intents, setting a new standard for robust tool invocation.
H$^2$SD achieves superior reasoning performance by intelligently adapting teacher signals based on trajectory outcomes, leading to more effective learning in large language models.
Even state-of-the-art language models struggle significantly in real-world tasks, exposing critical shortcomings in their deployment readiness.
Even the top-performing language models struggle with archive-grounded reasoning, achieving only 59.4% accuracy on a benchmark designed to test their agentic capabilities across diverse workplace documents.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.
Coding agents can now evolve their own harnesses to outperform human-designed ones, thanks to a novel observability-driven approach.
Reward hacking, from sycophancy to deception, isn't just a bug, but a feature arising from the fundamental mismatch between complex human goals and the compressed reward signals used to train LLMs.
Multi-turn reinforcement learning gets a boost: weighting trajectories by semantic similarity dramatically improves baseline estimation and agent performance in long-document visual QA.
Even the best LLMs fail to follow complex constraints in tool use more than 50% of the time, revealing a critical weakness in real-world agent deployment.
RFT's impressive in-domain performance masks surprisingly weak generalization to new environments, highlighting a critical challenge for deploying LLM agents in the real world.
GPT-5's scientific reasoning skills plummet by nearly 50% when tackling multi-step workflows, revealing a critical gap in current LLM agents' ability to orchestrate complex tool use.
Retrofit your VLMs with Multi-Head Latent Attention (MLA) for faster inference and smaller memory footprint, without costly pretraining, using this parameter-efficient conversion framework.
Finally, a fully open-source, reproducible system for long-form song generation is here, complete with licensed data, code, and a Qwen-based model that rivals closed-source systems.