Search papers, labs, and topics across Lattice.
1
0
3
0
This work proposes an exploration-guided prompt scaffolding framework that adapts the training prompt distribution dynamically throughout RL post-training of multimodal large language models (MLLMs), and uses a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory.