Search papers, labs, and topics across Lattice.
This paper introduces WADE, a novel benchmark for evaluating compact vision-language models (VLMs) in the context of multi-instance floating waste detection in rural Bangladesh, featuring 2,167 images and 13,608 bounding boxes across ten waste categories. The benchmark includes reasoning-annotated data that provides class-level recognition rules, enhancing the models' ability to localize, classify, and explain floating waste. Results show that fine-tuning the Qwen3-VL-2B model significantly improves recall and F1 scores, though a substantial number of instances remain undetected, highlighting the benchmark's difficulty and the need for further advancements in VLM capabilities.
Fine-tuning compact vision-language models on WADE boosts detection performance but still leaves over 75% of floating waste instances undetected.
Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic coverage, sparse multi-instance annotations, and little supervision beyond boxes and labels. Compact vision-language models (VLMs) therefore remain insufficiently evaluated for jointly localizing, classifying, counting, and explaining floating waste. We introduce WADE, a reasoning-annotated benchmark containing 2,167 images from rural Bangladesh, 13,608 bounding boxes, and ten waste categories. Each annotation is associated with class-level recognition rules covering visual cues, likely confusions, and discriminative features. We evaluate six VLMs under zero-shot, two-shot, reasoning-guided, and fine-tuned settings using detection, counting, and hallucination metrics. For resource-efficient adaptation, we jointly fine-tune Qwen3-VL-2B on boxes, labels, and reasoning chains using QLoRA. Fine-tuning increases recall from 0.0248 to 0.2339 and F1 from 0.0257 to 0.2163, while reducing image-level hallucination from 0.6836 to 0.0883. However, over three-quarters of instances remain undetected, establishing WADE as a challenging benchmark for dense floating-waste grounding with compact VLMs.