Search papers, labs, and topics across Lattice.
This paper introduces VSANet, a novel view-aware sparse attention network designed specifically for light field (LF) image denoising. By employing a view-aware sparse attention (VSA) block that utilizes locality-sensitive hashing for efficient cross-view aggregation, the model achieves linear complexity while effectively leveraging the strong correlations present in LF data. Experimental results show that VSANet significantly outperforms existing state-of-the-art LF denoising methods, highlighting its effectiveness in addressing the challenges posed by high-dimensional LF structures.
VSANet achieves linear complexity in LF image denoising while outperforming the best existing methods by effectively utilizing cross-view correlations.
Light field (LF) image denoising is challenging due to the high-dimensional structure of LF data. While noise is independent across sub-aperture images, scene content exhibits strong cross-view correlations. We introduce VSANet, a view-aware sparse attention network for LF denoising. Specifically, we propose a view-aware sparse attention (VSA) block that represents the 4D LF feature map as a unified spatial-angular token space and performs cross-view aggregation via locality-sensitive hashing-based sparse attention. This enables global feature interactions with linear complexity, effectively exploiting LF correlations across views and spatial locations. In addition, we design a feature refinement (FR) block to emphasize informative features in spatial, angular, and epipolar subspaces. The VSA and FR blocks are integrated within a sequential attention refinement module, forming the core of VSANet. Experiments demonstrate VSANet outperforms stateof-the-art LF denoising methods.