Search papers, labs, and topics across Lattice.
This study compares the predictive accuracy of full-feature versus limited-input machine learning models for residential energy estimation using two datasets: the survey-based RECS and the simulation-based ResStock. While full-feature models achieved high performance benchmarks (R2 = 0.90 for ResStock), restricting inputs to ten accessible variables resulted in a significant drop in accuracy (R2 = 0.61 for RECS and R2 = 0.62 for ResStock), highlighting the limitations of algorithmic complexity in the absence of comprehensive data. Notably, a targeted model for a specific subset of homes improved accuracy to R2 = 0.85, underscoring the importance of dataset characteristics in model performance.
Limiting input variables to just ten homeowner-accessible features can drastically reduce predictive accuracy, revealing the critical role of comprehensive data in energy estimation models.
Residential energy estimates are often needed before detailed envelope characteristics, equipment efficiencies, infiltration, sensor, or billing data are available. This study quantifies the trade-off between predictive accuracy and input accessibility using two nationally representative U.S. residential-energy datasets: the survey-based Residential Energy Consumption Survey (RECS) and the simulation-based ResStock dataset. Full-feature models were first used to establish dataset-specific performance benchmarks. For total-energy estimation, the models were subsequently restricted to ten low-burden variables obtainable from occupants, administrative records, or location-based weather data without an on-site energy audit. Among CatBoost, XGBoost, LightGBM, Random Forest, and Neural Networks, CatBoost consistently achieved the highest predictive performance for the full-feature analysis, reaching R2 = 0.90 for ResStock and R2 = 0.73 for RECS. When the feature set was restricted to ten homeowner-accessible inputs to simulate realistic deployment conditions, model performance converged to R2 = 0.61 for RECS and R2 = 0.62 for ResStock, showing that algorithmic complexity cannot fully compensate for missing physical and behavioral information. However, for a more homogeneous ResStock cohort consisting of single-family detached, natural-gas-heated homes in Climate Zone 6A constructed between 2000 and 2010, a reduced-input model improved accuracy to R2 = 0.85, demonstrating the value of targeted modeling for homogeneous populations. The results indicate that tree-based ensemble models can serve as high-fidelity emulators of national-scale residential energy datasets. However, careful consideration of feature availability, dataset origin (empirical vs. synthetic), and applicable use cases are also important.