Search papers, labs, and topics across Lattice.
Affiliation:
3
0
7
7
This work proposes an exploration-guided prompt scaffolding framework that adapts the training prompt distribution dynamically throughout RL post-training of multimodal large language models (MLLMs), and uses a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory.
VLM-hyster achieves unprecedented segmentation accuracy in hysteroscopic surgeries, outperforming all existing models and paving the way for more effective AI-assisted surgical interventions.
Forget scaling laws – AgenticQwen proves that clever training with dual data flywheels can enable small language models to rival giants in real-world agentic tasks.