Search papers, labs, and topics across Lattice.
This paper introduces HyperFlux, a novel ultralight VM substrate that enables microsecond-scale elasticity in core allocation across colocated lightweight VMs, addressing the challenge of maintaining high deployment density while effectively managing tail latency during traffic bursts. By allowing physical cores to be dynamically shifted between VMs in just 13 microseconds, HyperFlux significantly outperforms traditional methods, which typically rely on millisecond-scale adjustments. The results demonstrate that HyperFlux can reduce tail latency for high-priority VMs by up to 10 times compared to existing static core-sharing solutions, while maintaining a minimal memory footprint and rapid boot times.
HyperFlux achieves unprecedented core allocation speed, moving cores between VMs in just 13 microseconds, revolutionizing tail latency management in serverless environments.
Serverless platforms commonly colocate many diverse workloads, each in a fast-booting, memory-lean virtual machine (VM), to improve deployment density. Overprovisioning each VM for its peak protects tail latency during traffic bursts but hurts density; maintaining high density while effectively protecting tail latency requires the infrastructure to be able to shift physical cores, at a microsecond timescale, to whichever latency-sensitive VM is bursting and reclaim them as the burst subsides. No VM substrate delivers this: conventional VMs resize a guest's cores only through a millisecond-scale vCPU hot-plug path, Firecracker fixes a VM's core count at boot, and the ultralight VMs that boot fastest drop multicore execution entirely. We present HyperFlux, a commodity-KVM ultralight VM substrate that makes a VM's parallelism width (the number of physical cores backing it) elastic at runtime. We show that HyperFlux can move a core across VMs in merely 13$\mu$s, even when forcibly reclaiming it from a busy donor, orders of magnitude faster than vCPU hot-plug. A HyperFlux VM incurs only a 3.2MB memory footprint and can cold-boot in 1.37ms, on par with the fastest-booting ultralight VMs, while uniquely supporting multicore parallelism. Under colocation, it can reduce high-priority VMs'tail latency by up to 10x under high load compared to static core-sharing with Firecracker and Cloud Hypervisor, and deliver a lower and more stable tail latency compared to using cgroup and vCPU hot-plug under changing load bursts.