Search papers, labs, and topics across Lattice.
This paper introduces ACA-GS, an adaptive-capacity framework for 4D Gaussian Splatting that optimizes the balance between motion expressiveness and storage efficiency by dynamically adjusting the number of Neural Gaussians and their feature capacities based on local spatiotemporal demands. By employing Adaptive Anchor Cardinality and Adaptive Anchor Feature Masking, the method effectively concentrates resources in areas of high complexity while minimizing redundancy in simpler regions. Experimental results show that ACA-GS achieves up to 1.5x higher compression on complex motion sequences compared to existing anchor-based methods, all while maintaining visual fidelity.
Achieving 1.5x better compression on complex motion sequences without sacrificing visual quality could redefine efficiency in real-time spatiotemporal rendering.
Recent advances in 4D Gaussian Splatting (4DGS) enable high-fidelity, real-time spatiotemporal rendering, but expose a fundamental trade-off between motion expressiveness and storage efficiency. While anchor-based designs achieve compactness through anchor-level parameter sharing, their rigid uniform parametrization enforces fixed Neural Gaussian counts and feature budgets per anchor. Consequently, insufficient fidelity is addressed by excessive anchor density, rather than lightweight, targeted increases in Neural Gaussian count or feature capacity, resulting in memory waste. To overcome this rigidity, we introduce an adaptive-capacity anchor-based framework that dynamically allocates the representational capacity based on local spatiotemporal demands. Adaptive Anchor Cardinality varies the number of Neural Gaussians per anchor, concentrating primitives in regions of high geometric or motion complexity while suppressing redundancy. In parallel, Adaptive Anchor Feature Masking modulates anchor-level feature channels, assigning rich features to complex regions and lightweight representations to simpler ones. Experiments on MPEG, Panoptic Sports, and N3DV datasets demonstrate substantial storage reduction without degrading visual quality. Notably, on challenging MPEG sequences with complex motion, our method achieves up to 1.5x higher compression than state-of-the-art anchor-based methods while preserving comparable quality.