Search papers, labs, and topics across Lattice.
5
0
7
27
SEED transforms past experiences into actionable skills, allowing reinforcement learning policies to evolve and improve in real-time.
TACO reveals that agentic models can learn to optimize tool usage without external judges, achieving higher accuracy and efficiency in multimodal tasks.
OPID achieves a remarkable boost in agent performance by leveraging hierarchical skills extracted from on-policy trajectories, transforming sparse rewards into dense, actionable insights.
Forget monolithic models: a lightweight RL policy can dynamically orchestrate ensembles of frozen experts to outperform GPT-5 and Gemini-2.5-Pro on multimodal tasks, even generalizing to unseen models and skills.
Finally, a watermarking method exists that can be embedded directly into the weights of open-source neural speech generation models, enabling proactive copyright protection.