Search papers, labs, and topics across Lattice.
This paper introduces ACF-Net, an innovative framework designed for asymmetric audio-visual fine-grained visual categorization (FGVC), addressing the challenges posed by weakly synchronized audio and video inputs. By incorporating Optical Flow-Guided Motion (OFGM) to enhance dynamic representation and Asymmetric CrossModal Adaptive Fusion (ACAF) to adaptively merge modalities based on reliability, ACF-Net significantly improves category-level recognition. The authors validate their approach on the newly constructed BirdPro benchmark, achieving state-of-the-art performance with notable increases over existing methods in both fused and mismatched settings.
ACF-Net outperforms existing methods in fine-grained visual categorization by effectively handling the complexities of asymmetric audio-visual inputs.
Audio-visual cross-modal Fine-Grained Visual Categorization (FGVC) aims to identify fine-grained categories by jointly leveraging visual and auditory information. However, FGVC under asymmetric cross-modal scenarios has received limited attention, where paired video and audio are not strictly synchronized and may not even correspond to the same individual or moment. Such weak and ambiguous cross-modal correspondence poses substantial challenges to effective representation learning and modality alignment. To address these issues, we propose ACF-Net, a novel optical flow-guided framework for asymmetric audio-visual fine-grained learning. ACF-Net consists of two key modules: Optical Flow-Guided Motion (OFGM) and Asymmetric CrossModal Adaptive Fusion (ACAF). OFGM captures motion-sensitive visual cues and suppresses irrelevant background interference, thereby enhancing discriminative dynamic representations in videos. ACAF estimates modality reliability under weakly matched audio-video pairs and performs uncertainty-aware adaptive fusion to improve category-level recognition robustness. To support research on asymmetric cross-modal FGVC, we further construct BirdPro, a new bird-oriented audio-visual benchmark, since existing datasets often lack large-scale category-level audio-video associations under non-strict temporal and instance correspondence. BirdPro contains 1,919 audio recordings and 11,965 videos covering 194 bird species. Extensive experiments show that ACF-Net achieves the best results compared with representative baseline methods, outperforming the strongest baselines by 2.97% and 1.92% in the fused and mismatched settings, respectively.