Search papers, labs, and topics across Lattice.
This study evaluates the predictive validity of commonsense benchmarks by testing 23 models across various benchmark types and downstream tasks that require implicit reasoning. The findings reveal that while revised benchmarks maintain original model rankings, they fail to enhance predictive power for downstream tasks, indicating that commonsense benchmarks only demonstrate consistent validity for a limited range of applications. Ultimately, the research highlights that standardized commonsense benchmarks offer task-specific insights rather than a comprehensive measure of downstream commonsense competence.
Revised commonsense benchmarks don't boost predictive power for downstream tasks, revealing a critical limitation in their utility for assessing LLM capabilities.
Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.