Search papers, labs, and topics across Lattice.
This study employs functional analysis of variance (fANOVA) to assess the impact of various design choices on the performance of deep learning models for multi-label classification of remote sensing images. By analyzing 48 and 20 different models across seven datasets, the authors identify how factors like network architecture, fine-tuning strategy, and initialization interact to influence model performance. The results reveal distinct sensitivity profiles for datasets based on their intrinsic properties, highlighting that fine-tuning and architecture are crucial for large datasets, while initialization plays a key role in data-limited scenarios.
Dataset properties dictate how design choices impact model performance, revealing that fine-tuning and architecture dominate in large datasets, while initialization is critical in data-scarce environments.
Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize beyond the evaluated datasets. In this work, we move beyond rankings by employing functional analysis of variance (fANOVA) to systematically quantify the contributions of individual design choices and their interactions to performance variability. We conduct two empirical analyses covering 48 and 20 DL models, respectively, spanning design choices such as network architecture, fine-tuning strategy, learning strategy, and initialization. By applying fANOVA across seven MLC RSI datasets, we construct dataset meta-representations that capture design-choice sensitivity profiles. Hierarchical clustering of these meta-representations reveals that datasets naturally group according to how they respond to design decisions, with patterns strongly linked to intrinsic dataset properties such as scale, spatial resolution, and label space complexity. Our findings show that for large-scale datasets, fine-tuning strategy and architecture are dominant factors, while in data-limited regimes, initialization becomes decisive. For intermediate regimes, the interaction between architecture and learning strategy governs performance.