Search papers, labs, and topics across Lattice.
8
0
10
6
Interleaving reasoning and execution allows PACE to hide 66.8% of thinking time within action execution, drastically improving planning efficiency.
IACM-RL reduces infinite loops and stale context errors by proactively managing dynamic user intents, setting a new standard for robust tool invocation.
Even state-of-the-art language models struggle significantly in real-world tasks, exposing critical shortcomings in their deployment readiness.
Achieving an 85.71% repair success rate, VeriPilot transforms Verilog debugging by intelligently tracing dependencies and aligning code semantics.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.
Learned critics in RLHF can actually *increase* variance and hurt performance in sparse-reward settings, but a simple explained variance metric can tell you when to ditch the critic and get better results.
RFT's impressive in-domain performance masks surprisingly weak generalization to new environments, highlighting a critical challenge for deploying LLM agents in the real world.
GPT-5's scientific reasoning skills plummet by nearly 50% when tackling multi-step workflows, revealing a critical gap in current LLM agents' ability to orchestrate complex tool use.