Search papers, labs, and topics across Lattice.
This paper introduces a deterministic algorithm for the exact computation of local Real Log Canonical Thresholds (RLCTs) for two-dimensional singular models, addressing the limitations of classical information criteria like the Bayesian Information Criterion (BIC) in deep learning contexts. By deriving a complexity bound for the algorithm, the authors demonstrate its effectiveness across a wide range of models, including polynomial neural networks, and highlight its ability to reveal algebraic structures in learning coefficients that are obscured by sampling methods. The key result shows that this exact computation outperforms sampling-based estimators, particularly in the shallow regime, providing a more reliable foundation for model selection in singular scenarios.
Exact computation of learning coefficients reveals hidden algebraic structures that sampling methods miss, leading to more accurate model selection in deep learning.
Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning. The Widely Applicable Bayesian Information Criterion (WBIC) relies on local learning coefficients $\lambda$, which in the analytic case coincides with local Real Log Canonical Thresholds (RLCT) of the Kullback-Leibler divergence of the model, to capture correct marginal likelihood asymptotics. Exact computation of the learning coefficients has been limited to special cases, and only sampling-based estimation methods are generally applicable. We present the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynomial neural networks. Beyond providing ground truth to calibrate sampling-based estimators, exact computation reveals algebraic structure in learning coefficients that sampling cannot and out-speeds it in the shallow regime.