Search papers, labs, and topics across Lattice.
This paper introduces PatchGen, a novel approach that identifies and utilizes intra-image predictive subsets to enhance visual classification under various data shifts. By focusing on a sample-adaptive oracle subset, the method maintains the Bayes risk of full-patch representations while improving generalization and performance across different contexts. Experimental results demonstrate that PatchGen outperforms baseline models in diverse settings, particularly in histopathology, by prioritizing tumor-related features over irrelevant contextual information.
Identifying intra-image predictive subsets can significantly boost visual classification performance, especially in challenging data shift scenarios.
Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image predictive subset sufficient for label prediction, while the remaining patches form non-essential complementary context that may correlate with the label. The theoretical analysis shows that restricting prediction to this oracle subset preserves the Bayes risk achievable by the full-patch representation while admitting a complexity bound that tightens with the oracle-subset size. Based on this view, we propose PatchGen, a text-free module that learns a sample-dependent soft predictive-subset mask as a task-driven proxy for the unobserved oracle subset mask. Specifically, histopathology visualizations suggest that PatchGen assigns higher scores to tumor-consistent regions than to some frequently co-occurring inflammatory context. Extensive experiments on natural and histopathological image benchmarks spanning all three shift settings show that PatchGen improves average performance over matched-backbone baselines in most evaluated configurations, enhances generalization to unknown classes, and remains competitive with vision-language methods without text supervision.