Search papers, labs, and topics across Lattice.
This paper introduces the Descriptive-Complexity Information Criterion (DCIC), which addresses the challenges of model selection in the presence of strong predictor dependence and uncertainty across model classes. By employing Kraft-admissible code lengths, the DCIC ensures selection consistency and provides nonasymptotic oracle risk bounds, even under model misspecification. The proposed framework not only facilitates class-model recovery but also allows for a complexity-guided search that balances computation and statistical efficiency, demonstrating stable performance in numerical experiments.
Regularizing model selection with a coding-theoretic approach reveals a new pathway to consistent performance even under complex dependencies and uncertainties.
Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty, especially when there are exponentially many models. We propose a Descriptive-Complexity Information Criterion (DCIC) that regularizes large candidate model collections through Kraft-admissible code lengths. Under sub-Weibull noise, we establish selection consistency through approximation-error separation without relying on RIP-type conditions, together with nonasymptotic oracle risk bounds that remain valid under model misspecification. The same coding principle places heterogeneous classes on a common complexity scale at a small additional class-identification cost. This extension yields class--model recovery under suitable identifiability conditions and risk adaptation across classes. We further develop a complexity-guided search path that makes the computation--statistics trade-off explicit. Large penalties yield polynomial-size retained search regions with high probability, whereas smaller penalties sharpen the oracle risk benchmark. Numerical experiments illustrate stable support recovery and favorable estimation performance under strong dependence and model-class uncertainty.