Search papers, labs, and topics across Lattice.
2
0
6
4
This work proposes an exploration-guided prompt scaffolding framework that adapts the training prompt distribution dynamically throughout RL post-training of multimodal large language models (MLLMs), and uses a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory.
Forget scaling laws – AgenticQwen proves that clever training with dual data flywheels can enable small language models to rival giants in real-world agentic tasks.