Search papers, labs, and topics across Lattice.
This paper introduces PaperGym, a novel framework that transforms research papers into training environments for AI research planning by utilizing rubrics as critics. By synthesizing research questions and deriving evaluation criteria from the structure of scientific papers, PaperGym significantly reduces criterion leakage to 3.7% and improves model performance across multiple benchmarks. The Qwen3-8B model trained on PaperGym-20k outperforms larger models, achieving a score of 73.48 on ResearchQA, demonstrating the effectiveness of this approach in enhancing AI's research planning capabilities.
PaperGym achieves a remarkable 73.48 on ResearchQA, outperforming larger models and redefining how AI can generate and evaluate research plans.
Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks paired with a critic. Rubrics extracted from scientific papers can supply the critic. Existing pipelines, however, draw the question and the criteria from the same content, so the reward can be earned by paraphrase. The rubric is further compressed into a single scalar per rollout. We introduce PaperGym, a unified framework that turns each research paper into a complete training environment. PaperGym exploits the structure of a paper: the question is synthesized from the research goal and background, while the criteria are derived from the method and experiments. The criteria span methodological innovation and experimental design, and criterion leakage falls to 3.7%, versus 11.90% to 34.10% in existing datasets. Training uses the rubric twice: first as privileged context for OPSD's self-teacher, then as the reward for GRPO. Across Qwen3-1.7B/4B/8B, this schedule outperforms supervised fine-tuning, either stage alone, and the reverse ordering, improving five-benchmark averages by +5.6, +5.0, and +4.8 points. With the recipe held fixed, models trained on PaperGym-20k win 58.1% of three-way comparisons, against 28.2% for RubricHub Science. The trained Qwen3-8B reaches 73.48 on ResearchQA, above the far larger Kimi K2.6. We release the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.