Search papers, labs, and topics across Lattice.
Fudan University
10
1
14
26
LLMs struggle with personal information retrieval, with the best model only achieving 57.3% accuracy on a new benchmark designed to evaluate mobile assistant capabilities.
IACM-RL reduces infinite loops and stale context errors by proactively managing dynamic user intents, setting a new standard for robust tool invocation.
MemArbiter closes the Memory-Action Gap, boosting LLM decision-making success rates by over 20% in long-horizon tasks.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.
Coding agents can now evolve their own harnesses to outperform human-designed ones, thanks to a novel observability-driven approach.
Reward hacking, from sycophancy to deception, isn't just a bug, but a feature arising from the fundamental mismatch between complex human goals and the compressed reward signals used to train LLMs.
Multi-turn reinforcement learning gets a boost: weighting trajectories by semantic similarity dramatically improves baseline estimation and agent performance in long-document visual QA.
Forget end-to-end training: breaking down long-context reasoning into atomic skills and training on targeted pseudo-data unlocks a 7.7% performance boost.
GPT-5's scientific reasoning skills plummet by nearly 50% when tackling multi-step workflows, revealing a critical gap in current LLM agents' ability to orchestrate complex tool use.
Finally, a fully open-source, reproducible system for long-form song generation is here, complete with licensed data, code, and a Qwen-based model that rivals closed-source systems.