Search papers, labs, and topics across Lattice.
Max Planck Institute for Intelligent Systems, Tübingen, Germany, Tübingen AI Center
2
0
3
Elo rankings can accurately reflect model performance, achieving over 90% correlation with ground-truth accuracy, even amidst stylistic biases.
Current ML benchmarks may be ungameable in theory, as they can lack a stable equilibrium where developers are incentivized to improve true model quality rather than just leaderboard scores.