Search papers, labs, and topics across Lattice.
This paper introduces ST-ColoNet, a two-stage deep learning framework for colo-segment recognition in colonoscopy videos, addressing the limitations of existing image-based methods by incorporating temporal information. The framework includes a Colorlaus module for edge-guided spatial feature extraction via metric learning and a Full-Temp module that combines self-attention patterns for improved temporal feature aggregation. Experiments on a newly curated dataset demonstrate state-of-the-art performance, achieving 81.0% accuracy and 70.7% F1-score, significantly outperforming existing methods.
Achieve state-of-the-art colo-segment recognition by combining edge-guided spatial feature extraction with a novel temporal attention mechanism, outperforming existing methods by a large margin.
Colo-segment recognition in colonoscopy videos is a key requirement for many downstream tasks, but existing automatic recognition methods only use colonoscopy images without fully exploiting the use of temporal information, leading to poor performance. Additionally, relevant public video-based datasets are in scarcity. To tackle this problem, we curate and release a labeled dataset specifically for the task of colo-segment recognition. In addition, we propose a two-stage deep learning-based framework, Colo-Segment Recognition via SpatioTemporal Network (ST-ColoNet), for the task of colo-segment recognition from colonoscopy videos which includes the Colorlaus module that uses metric learning to optimize edge-mediated spatial feature extraction, as well as the Full-Temp module which combines three self-attention patterns to better approximate full self-attention on long colonoscopy sequences and optimize temporal feature aggregation. Through extensive ablation experiments, we show that our framework is capable of achieving state-of-the-art performance on the task of colo-segment recognition, achieving an accuracy of 81.0% and F1-score of 70.7%, which is a tremendous improvement over state-of-the-art methods.