Search papers, labs, and topics across Lattice.
This paper introduces Test-Time Policy Optimization (TTPO), a novel approach that leverages majority-vote pseudo-labels to optimize large language models during test time without relying on ground-truth labels. By recognizing the asymmetric failure modes of pseudo-labels, TTPO employs an objective that distills agreeing rollouts while penalizing those that disagree, effectively enhancing model performance even in the presence of frequent label errors. The results demonstrate that TTPO can achieve label-supervised performance on multiple benchmarks, significantly improving the Qwen3-1.7B model's test-time training outcomes and showcasing robust cross-task generalization.
TTPO achieves label-supervised performance without any ground-truth labels, outperforming traditional methods on key benchmarks.
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.