Search papers, labs, and topics across Lattice.
14
0
12
6
Recursive belief updates in AgentOPSD reveal pivotal decision points, leading to a 89.1% success rate on complex RL tasks.
World rehearsal enables LLM agents to internalize environment dynamics, achieving superior performance without costly external interactions.
Selective distillation can unlock critical learning signals in reinforcement learning, leading to significant performance gains in complex tasks.
VAD reveals that isolating visual evidence can dramatically improve target reconstruction in multimodal learning, leading to more accurate student outputs.
SkillRise achieves up to 8.5 percentage points better performance than leading methods by effectively reusing transferable skills across related tasks.
SEED transforms past experiences into actionable skills, allowing reinforcement learning policies to evolve and improve in real-time.
OPID achieves a remarkable boost in agent performance by leveraging hierarchical skills extracted from on-policy trajectories, transforming sparse rewards into dense, actionable insights.
GUI agents can learn world knowledge more efficiently by internalizing causal relationships during mid-training, rather than relying on implicit learning through action annotations or reward signals in post-training.
Forget monolithic models: a lightweight RL policy can dynamically orchestrate ensembles of frozen experts to outperform GPT-5 and Gemini-2.5-Pro on multimodal tasks, even generalizing to unseen models and skills.
Uncertainty-driven zoom-in boosts GUI grounding accuracy by up to 13.4% without any retraining, showing that targeted attention to model uncertainty can significantly improve performance.
Offloading memory and computation to a copilot lets a 7B parameter GUI agent outperform larger models on long-horizon tasks, suggesting a path to more efficient and capable GUI automation.
Even frontier models like Claude Sonnet 4.6 stumble when asked to infer user preferences and proactively assist in mobile tasks, achieving less than 50% success despite excelling at explicit task execution.
LLM agents can internalize skills via in-context RL, achieving zero-shot autonomous behavior without the token overhead and retrieval noise of traditional methods.
By adversarially co-evolving code and test LLMs, Code-A1 achieves code generation performance on par with human-annotated training, while simultaneously boosting the LLM's ability to find bugs.