Search papers, labs, and topics across Lattice.
22
1
17
33
Natural-language critiques can transform how we evaluate and optimize song generation models, leading to more human-aligned outputs.
LLMs struggle with personal information retrieval, with the best model only achieving 57.3% accuracy on a new benchmark designed to evaluate mobile assistant capabilities.
IACM-RL reduces infinite loops and stale context errors by proactively managing dynamic user intents, setting a new standard for robust tool invocation.
Even state-of-the-art language models struggle significantly in real-world tasks, exposing critical shortcomings in their deployment readiness.
ConsumerSim reveals that consumer confidence is driven more by individual interpretations of salient events than by aggregate trends, transforming our understanding of economic sentiment dynamics.
Legal agents exhibit starkly different performance across litigation stages, challenging the notion of a one-size-fits-all model in legal AI applications.
VLA models can lose up to 40% of their effectiveness in real-world scenarios due to occlusion, but viewpoint imagination offers a powerful solution to this pervasive challenge.
Gradual bridging with embodied trajectory-coupled data transforms VLMs into robust robot control policies, overcoming significant transfer challenges.
LLMs can drive agentic RAG to match the performance of complex hybrid retrieval systems, even with a simple inverted index, by expressing retrieval intents as logical expressions.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.
Coding agents can now evolve their own harnesses to outperform human-designed ones, thanks to a novel observability-driven approach.
Current benchmarks miss the point: the real value of AI peer review lies in the quality of its textual justification, not just predicting a rating.
Learned critics in RLHF can actually *increase* variance and hurt performance in sparse-reward settings, but a simple explained variance metric can tell you when to ditch the critic and get better results.
Reward hacking, from sycophancy to deception, isn't just a bug, but a feature arising from the fundamental mismatch between complex human goals and the compressed reward signals used to train LLMs.
Multi-turn reinforcement learning gets a boost: weighting trajectories by semantic similarity dramatically improves baseline estimation and agent performance in long-document visual QA.
You can dial up or down how obvious an AI's hallucinations are, giving you control over whether users catch the errors.
Even the best LLMs fail to follow complex constraints in tool use more than 50% of the time, revealing a critical weakness in real-world agent deployment.
Forget benchmarks: AI can now learn "scientific taste" and propose research ideas with higher potential impact than humans, thanks to a novel reinforcement learning approach using citation data.
RFT's impressive in-domain performance masks surprisingly weak generalization to new environments, highlighting a critical challenge for deploying LLM agents in the real world.
Current LLMs fall short in understanding implicit intentions and modeling long-term user preferences, as revealed by a new benchmark, LifeSim-Eval, designed to simulate real-world user-assistant interactions.
GPT-5's scientific reasoning skills plummet by nearly 50% when tackling multi-step workflows, revealing a critical gap in current LLM agents' ability to orchestrate complex tool use.
Finally, a fully open-source, reproducible system for long-form song generation is here, complete with licensed data, code, and a Qwen-based model that rivals closed-source systems.