Search papers, labs, and topics across Lattice.
This paper addresses the challenge of deploying large-kernel Convolutional Neural Networks (CNNs) on resource-constrained edge devices by introducing a Channel Group-Shared (CGS) low-rank approximation method. By leveraging a novel Singular Value Decomposition (SVD)-based parameter-sharing strategy, the authors significantly reduce the parameter volume of pointwise convolutions, which account for over 87% of the parameters in models like RepLKNet-31B. Experimental results show that CGS enables competitive performance while drastically lowering storage costs, making high-performance vision models practical for deployment on devices with limited RAM.
Pointwise convolutions, which dominate parameter volume in large-kernel CNNs, can be drastically reduced through a novel group-sharing strategy, enabling efficient deployment on edge devices.
Large-kernel Convolutional Neural Networks (CNNs) deliver remarkable performance in vision tasks by significantly expanding receptive fields, yet their quadratic parameter growth critically impedes storage-efficient edge deployment. While existing efficient architectures adopt parameter-efficient depthwise separable convolution backbones that leverage techniques like low-rank approximation and weight sharing to compress depthwise convolutions, we identify a critical oversight: pointwise convolutions dominate parameter volume (>87% in models like RepLKNet-31B) and constitute the primary deployment bottleneck on resource-constrained edge devices. This results in prohibitive storage costs and severe memory-loading constraints on resource-limited devices (e.g., smartphones with 4-12 GB Random Access Memory (RAM)). To overcome this, we propose Channel Group-Shared (CGS) low-rank approximation, a novel Singular Value Decomposition (SVD)-based parameter-sharing strategy. CGS constructs a structured low-rank paradigm isomorphic to SVD decomposition, comprising shared (high-parameter-cost) down/up-projection matrices across channel groups within a layer and channel-group-specific (low-parameter-cost) scalable diagonal matrices. This group-sharing design achieves significant parameter reduction. Extensive experiments demonstrate that large-kernel CNNs (RepLKNet, ConvNeXt, SLaK) enhanced with CGS strike an empirically favorable balance between competitive performance and substantially reduced storage costs. Crucially, by alleviating storage constraints, reducing memory bandwidth pressure during loading, and minimizing model loading latency, CGS enables the feasible deployment of pre-trained large-kernel CNN models on edge devices, thereby bridging the gap between high-performance vision models and practical edge deployment.