Search papers, labs, and topics across Lattice.
Shanghai Jiao Tong University
3
0
8
Allocating rollout budgets based on state informativeness allows LLM agents to achieve superior performance in complex decision-making tasks without increasing computational costs.
Forget hand-tuning: VisPCO automatically finds optimal visual token pruning configurations in VLMs, outperforming predefined strategies across diverse benchmarks.
Forget memorizing surface patterns: RADAR leverages reinforcement learning to teach LLMs genuine relational reasoning, boosting knowledge graph performance by 5-6%.