Search papers, labs, and topics across Lattice.
This paper surveys the evolution of object-counting methods from class-specific density regression to open-vocabulary counters supported by foundation models, highlighting the need for a robust evaluative framework. Despite significant advancements, the authors reveal that existing benchmarks are often exploited for statistical regularities, leading to systematic failures in areas such as semantic grounding and spatial reasoning. They introduce a five-axis taxonomy to classify these methods and propose a roadmap for addressing the identified challenges, emphasizing the importance of distinguishing between genuine generalization and benchmark-specific optimization.
Systematic failures in object counting reveal that current benchmarks may mislead progress claims, necessitating a reevaluation of evaluation standards.
Object-counting methods have rapidly shifted from class-specific density regression to open-vocabulary, foundation-model-backed counters. These methods now enumerate instances from various visual and textual prompts. While this shift marks major conceptual progress, our survey argues that claims of universal generality have outpaced the evaluative infrastructure. Most progress metrics rely on a few saturated benchmarks that models exploit for statistical regularities. Newly introduced diagnostic datasets reveal systematic failures in semantic grounding, temporal identity, and spatial reasoning with occlusion. To address these failures, we introduce a five-axis taxonomy (modality, mechanism, prompting, supervision level, and generalization setting). We use this taxonomy to audit the literature across application domains, including microscopy, remote sensing, crowd counting, and agriculture. This formalizes prevailing challenges into six structural contradictions. From these, we propose a roadmap for compositional scene understanding, active counting agents, and unified multimodal evaluation protocols. The main imperative is to build a robust evaluation infrastructure to distinguish open-world generalization from benchmark-specific optimization, rather than simple incremental engineering.