Search papers, labs, and topics across Lattice.
This paper investigates the power dynamics of AI data centers running mixed batch and inference workloads, revealing a decoupling between aggregate power variability and short-horizon ramping. Using a trace-calibrated framework, the study finds that variability exhibits a U-shaped pattern while ramping displays a hump-shaped pattern as the inference share increases, with the magnitude and turning points dependent on system loading. The underlying mechanism involves queued batch jobs filling capacity left idle by fluctuating inference demand, reducing aggregate power variability, while inference-side fluctuations propagate more directly into realized power, elevating short-horizon ramping.
The mix of batch and inference workloads in AI data centers creates surprising power dynamics, where smoothing aggregate power demand paradoxically *increases* short-horizon ramping.
Artificial intelligence (AI) is driving rapid growth in electricity demand, yet the grid-facing power dynamics of AI data centers remain poorly understood. Here we show that, in shared-GPU systems, the composition of batch and inference workloads decouples aggregate power variability from short-horizon ramping. As the inference share rises, variability becomes U-shaped, whereas ramping becomes hump-shaped, particularly under higher loading. The magnitude and turning points of these patterns also depend on system loading. Using a trace-calibrated framework linking workload arrivals, queueing, scheduling, and GPU power, we show that the underlying mechanism is asymmetric. At intermediate workload mixes, queued batch jobs fill capacity left idle by fluctuating inference demand, reducing aggregate power variability. However, short-horizon ramping remains elevated because inference-side fluctuations propagate more directly into realized power. AI data centers should therefore be understood as dynamic systems whose workload composition shapes their grid impact.