Search papers, labs, and topics across Lattice.
Beijing National Research Center for Information Science and Technology
6
0
9
27
SEED transforms past experiences into actionable skills, allowing reinforcement learning policies to evolve and improve in real-time.
Achieving competitive multimodal performance while training only a fraction of the parameters reveals a new pathway for efficient visual processing in large language models.
TACO reveals that agentic models can learn to optimize tool usage without external judges, achieving higher accuracy and efficiency in multimodal tasks.
OPID achieves a remarkable boost in agent performance by leveraging hierarchical skills extracted from on-policy trajectories, transforming sparse rewards into dense, actionable insights.
Forget monolithic models: a lightweight RL policy can dynamically orchestrate ensembles of frozen experts to outperform GPT-5 and Gemini-2.5-Pro on multimodal tasks, even generalizing to unseen models and skills.
Finally, a watermarking method exists that can be embedded directly into the weights of open-source neural speech generation models, enabling proactive copyright protection.