Search papers, labs, and topics across Lattice.
The paper introduces TF-CADE, a novel approach for Zero-Shot Temporal Action Detection (ZSTAD) that enhances the alignment between textual descriptions and action-relevant video segments. By employing Action Concentrate Aggregation (ACA), TF-CADE extracts and aggregates temporally informative video segments, resulting in a foreground-weighted embedding that improves semantic consistency and inter-class discriminability. Extensive evaluations demonstrate that TF-CADE achieves state-of-the-art performance in both in-distribution settings and cross-dataset generalization for unseen action categories.
Aligning text with action-relevant video segments, TF-CADE significantly boosts performance in zero-shot action detection, outperforming existing methods.
Zero-Shot Temporal Action Detection (ZSTAD) aims to lo- calize and recognize action instances from unseen action categories in untrimmed videos. Although existing meth- ods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing se- mantic distinctions between action classes, resulting in text- irrelevant predictions. To address this issue, we propose a Text-Foreground Concentrated Alignment for zero-shot temporal action DEtector (TF-CADE) that explicitly aligns textual information with action-relevant foreground regions. Specifically, we introduce Action Concentrate Aggregation (ACA), which extracts action concentrate scores to aggregate temporally informative video segments into a foreground- weighted video embedding. This foreground concentrated alignment enhances the semantic consistency between text and video features and improves inter-class discriminabil- ity. In addition, a Certainty-based Confidence Re-weighting (CCR) strategy refines per-snippet confidence scores by lever- aging foreground-aware similarity, effectively suppressing irrelevant action classes during inference. Extensive evalua- tions show that our TF-CADE not only achieves state-of-the- art performance under in-distribution settings but also excels in cross-dataset generalization to unseen action classes.