Search papers, labs, and topics across Lattice.
This study audits the performance of various zero-shot vision-language models (VLMs) in recognizing Bangladeshi freshwater fish, revealing that BioCLIP2 significantly outperforms generic CLIP models, achieving accuracies of 72.36% and 68.91% with English and scientific names, respectively. The results indicate that model performance is influenced not only by visual recognition capabilities but also by factors such as multilingual alignment and prompt formulation. Additionally, the research highlights the limitations of Bengali prompts, which yield near-chance performance, underscoring the complexity of context sensitivity in biological recognition tasks.
BioCLIP2 outperforms generic CLIP models by over 45% in recognizing Bangladeshi freshwater fish, revealing the critical role of nomenclature and context sensitivity in zero-shot learning.
Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 images). BioCLIP2 reaches 72.36% on BFF-15 with English common names and 68.91% on SylFishBD with scientific names, versus 25.15% and 14.40% for generic CLIP. BioCLIP2 Bengali prompts are near chance in balanced accuracy (14.22-14.29%); Jina partially recovers Bengali discrimination to 21.89% and 16.36%, but bare Bengali names return to 14.29% on both sources. Paired SylFishBD interventions show no significant weak-blur effect, modest losses from stronger blur/gray masking, a larger white-mask artifact, and strong species dependence. Zero-shot biological VLM scores therefore jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.