Search papers, labs, and topics across Lattice.
This paper evaluates the environmental break-even point for machine learning-based data compression algorithms by analyzing the carbon footprint of the infrastructure required for training and inference against the carbon savings achieved through reduced disk storage. The study specifically focuses on a lossless compression algorithm to quantify these impacts, revealing critical insights into when the environmental costs of ML training are offset by storage savings. The findings indicate that there exists a specific threshold where the benefits of data compression align with sustainability goals, making it essential for future ML applications.
ML-based data compression can be environmentally sustainable, but only if it surpasses a critical break-even point in carbon savings.
We summarise the outcome of two summer internship projects based at the University of Manchester, focused on the break-even point in terms of environmental sustainability for ML-based data compression algorithms. Using the example of a ML-based lossless compression algorithm, we compare estimates for the carbon-equivalent of the infrastructure needed for ML training and inference with the carbon-equivalent savings from reduced disk storage requirements, and discuss their break-even point.