Search papers, labs, and topics across Lattice.
This paper introduces DeGS, a novel architecture designed to enhance the scalability of 3D Gaussian Splatting (3DGS) by addressing the inefficiencies of the traditional tightly coupled dataflow. By decoupling the workload parsing and reorganization from the blending process, DeGS effectively reduces spatial and temporal redundancies, leading to substantial improvements in processing element (PE) utilization. The architecture achieves remarkable performance gains, with throughput increases of 2.36x to 7.25x and energy efficiency improvements of 1.59x to 4.42x compared to existing 3DGS accelerators, while maintaining over 80% PE utilization even at high resolutions.
DeGS redefines the performance landscape for 3D Gaussian Splatting, achieving up to 7.25x throughput improvements while maintaining high processing element utilization.
3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of PEs leads to marginal performance improvement during rendering. We identify that the root cause is the tightly coupled ``checking-while-blending''dataflow, which exacerbates PE underutilization caused by spatial redundancy from irregular Gaussian coverage and temporal redundancy from asynchronous pixel-wise termination under parallel execution. To address this issue, we propose DeGS, a scalable architecture for efficient 3DGS inference. To systematically eliminate the redundancies inherent in rendering, DeGS exploits a decoupled dataflow, restructuring the coupled $\alpha$-checking, transmittance checking, and $\alpha$-blending of the standard rendering process into consecutive workload parsing, reorganization, and blending stages. This allows the fragmented, length-variable, and temporal-dependent workloads to be reorganized into compact, conflict-free, and dense workloads prior to blending, thereby significantly improving PE utilization during parallel blending. Implemented in 28 nm technology, DeGS achieves 2.36$\times$--7.25$\times$ throughput, 1.82$\times$--6.02$\times$ end-to-end speedup, and 1.59$\times$--4.42$\times$ energy efficiency over state-of-the-art 3DGS accelerators (GSCore, GBU, GCC) across diverse scenes and resolutions (720p to 8K). Moreover, scaling from 16 to 1024 PEs, DeGS maintains over 80\% PE utilization at high resolutions, significantly outperforming existing accelerators.