Search papers, labs, and topics across Lattice.
This paper introduces a CPU+DCU heterogeneous parallel framework designed to enhance post-processing reconstruction in quantum circuit cutting, addressing the significant computational and storage challenges posed by large quantum circuits in the NISQ era. By reconstructing nonzero-probability states from subcircuit measurement results rather than relying on dense probability vectors, the framework achieves remarkable speedups, including a 259x improvement over a serial baseline. Experiments conducted on the Songshan supercomputer confirm that this approach not only maintains high fidelity but also scales effectively to handle reconstruction tasks involving hundreds of qubits.
Achieving up to 259x speedup in quantum circuit reconstruction could revolutionize how we handle large-scale quantum computations.
In the NISQ era, limited qubit resources make it difficult to execute large quantum circuits directly on real hardware. Quantum circuit cutting mitigates this limitation by decomposing a large circuit into smaller subcircuits, but it shifts substantial overhead to classical post-processing. As circuit size, complexity, and cut count increase, reconstruction becomes a major computational and storage bottleneck. This paper presents a CPU+DCU heterogeneous parallel framework for circuit-cutting post-processing reconstruction. Instead of constructing a dense $2^n$-dimensional probability vector or returning only high-probability states, the framework reconstructs the nonzero-probability states in the original output distribution from subcircuit measurement results. It combines heterogeneous CPU+DCU execution with a high/low-word integer representation for global basis-state indices beyond 64 bits and a three-level cooperative storage mechanism spanning device memory, host memory, and out-of-core storage. Experiments on the Songshan supercomputer show that the framework maintains high reconstruction fidelity while achieving up to $259\times$ speedup over an optimized serial baseline on linear-cluster states and up to $4\times$ speedup over a homogeneous CPU-parallel method on random circuits. The framework can also complete reconstruction tasks at the hundred-qubit scale. These results demonstrate that HPC-oriented heterogeneous reconstruction can effectively alleviate the classical post-processing bottleneck and improve reconstruction scalability.