Search papers, labs, and topics across Lattice.
To overcome the severe boundary ambiguity in timestamp-supervised action segmentation, this work develops a boundary voting network that enhances feature discriminability in transition zones by injecting global temporal priors. The framework gathers key action representations as collaborative "votes" across the video to hierarchically sharpen representations in action-transiting regions and stabilize pseudo-label generation. Across standard benchmarks including GTEA, 50Salads, and Breakfast, this voting mechanism consistently refines boundary localization and achieves state-of-the-art segmentation performance.
Timestamp-supervised video segmentation consistently fails at transition boundaries, but aggregating video-wide keyframe "votes" resolves local ambiguity without requiring dense manual annotations.
Timestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a random frame annotated per action. Precisely localizing action boundaries from timestamp annotations is crucial for this setting, as it enables generating framewise pseudo-labels and applying the well-explored fully-supervised training. However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions. This imprecise boundary estimation significantly reduces the stability and reliability of the generated pseudo-labels in ambiguous action-transiting regions, consequently resulting in performance deterioration of the trained segmentation models. In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. By generating key action representations as votes throughout the video and targeting action-transiting regions, all votes collaboratively contribute to action-transiting feature enhancement and boundary localization refinement. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets.