Search papers, labs, and topics across Lattice.
Tsinghua University, Beijing National Research Center for Information Science and Technology
Tsinghua AI4
0
5
7
SEED transforms past experiences into actionable skills, allowing reinforcement learning policies to evolve and improve in real-time.
Achieving competitive multimodal performance while training only a fraction of the parameters reveals a new pathway for efficient visual processing in large language models.
TACO reveals that agentic models can learn to optimize tool usage without external judges, achieving higher accuracy and efficiency in multimodal tasks.
OPID achieves a remarkable boost in agent performance by leveraging hierarchical skills extracted from on-policy trajectories, transforming sparse rewards into dense, actionable insights.