Search papers, labs, and topics across Lattice.
This paper introduces a novel framework called prediction-powered evaluation (PPE), which integrates limited human judgments with large-scale automatic scores to enhance data efficiency in system comparisons while ensuring unbiased results. By developing both parametric and non-parametric procedures, the authors analyze the efficiency trade-offs of different evaluation designs and validate their approach across six WMT datasets. The introduction of the Prediction-Powered Saving Ratio (PPSR) provides a new meta-metric that quantifies the reduction in human annotation required, demonstrating that automatic metrics can effectively complement human evaluation rather than replace it.
Automatic metrics can significantly reduce human annotation costs while maintaining unbiased evaluations, as shown by the new Prediction-Powered Saving Ratio (PPSR).
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased. We develop parametric and non-parametric procedures, analyze the efficiency trade-off between paired and unpaired designs, and validate the framework on six WMT datasets. We further introduce the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric can save when used within prediction-powered evaluation. PPSR directly targets metric utility for prediction-powered evaluation and yields more discriminative and stable metric rankings than existing system-level meta-metrics. Overall, our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment, and applies broadly to non-verifiable tasks.