Required prerequisites
Describe the bug
We ran into an OOM-kill while running a large-scale VQE job on our node (CUDA-Q 0.12.0). This was unexpected, since the node has about 125 GB of free RAM.
We traced it to two separate host-memory leaks triggered by repeated cudaq.observe() calls: one in the uccsd() kernel builtin, and one in how cudaq.observe() handles a Hamiltonian built via cudaq.chemistry.create_molecular_hamiltonian().
We built two minimal reproducers and confirmed both leaks on CUDA-Q 0.15.0 and on CUDA-Q amd64-cu12-latest (52f4442), so this isn't specific to 0.12.0.
Steps to reproduce the bug
import argparse
import resource
import cudaq
from cudaq import spin
from cudaq.kernels import uccsd, uccsd_num_parameters
def main():
parser = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
parser.add_argument('--n-electrons', type=int, default=2)
parser.add_argument('--n-qubits', type=int, default=4)
parser.add_argument('--n-calls', type=int, default=300)
parser.add_argument('--log-every', type=int, default=20)
args = parser.parse_args()
cudaq.set_target('nvidia', option='fp64')
n_e, num_qubits = args.n_electrons, args.n_qubits
num_params = uccsd_num_parameters(n_e, num_qubits)
@cudaq.kernel
def ansatz(thetas: list[float]):
q = cudaq.qvector(num_qubits)
for i in range(n_e):
x(q[i])
uccsd(q, thetas, n_e, num_qubits)
# Arbitrary, fixed, non-chemistry Hamiltonian -- any valid spin_op on
# `num_qubits` qubits reproduces this; the specific terms don't matter.
hamiltonian = spin.z(0) * spin.z(1) + spin.x(1) * spin.x(num_qubits - 1) + spin.z(num_qubits - 1)
theta = [0.0] * num_params # FIXED for the entire run -- never changes, no optimizer involved
print(f"n_electrons={n_e} n_qubits={num_qubits} num_params={num_params}")
print(f"{'call':>6} {'rss_MB':>10}")
for i in range(1, args.n_calls + 1):
cudaq.observe(ansatz, hamiltonian, theta)
if i == 1 or i % args.log_every == 0:
rss_mb = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
print(f"{i:6d} {rss_mb:10.1f}")
if __name__ == "__main__":
main()
import argparse
import resource
import cudaq
def main():
parser = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
parser.add_argument('--n-calls', type=int, default=2000)
parser.add_argument('--log-every', type=int, default=200)
args = parser.parse_args()
cudaq.set_target('nvidia', option='fp64')
# H2/STO-3G: smallest possible nontrivial molecular Hamiltonian.
geometry = [('H', (0.0, 0.0, 0.0)), ('H', (0.0, 0.0, 0.74))]
molecule, _ = cudaq.chemistry.create_molecular_hamiltonian(
geometry=geometry, basis='sto-3g', multiplicity=1, charge=0,
n_active_electrons=2, n_active_orbitals=2,
)
num_qubits = molecule.qubit_count
# Deliberately NOT uccsd() -- a trivial fixed circuit, to isolate the
# Hamiltonian object as the only chemistry-related ingredient.
@cudaq.kernel
def ansatz(theta: float):
q = cudaq.qvector(num_qubits)
for i in range(num_qubits):
ry(theta, q[i])
for i in range(num_qubits - 1):
cx(q[i], q[i + 1])
theta = 0.0 # FIXED for the entire run -- never changes, no optimizer involved
print(f"qubits={num_qubits}")
print(f"{'call':>6} {'rss_MB':>10}")
for i in range(1, args.n_calls + 1):
cudaq.observe(ansatz, molecule, theta)
if i == 1 or i % args.log_every == 0:
rss_mb = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
print(f"{i:6d} {rss_mb:10.1f}")
if __name__ == "__main__":
main()
Expected behavior
ru_maxrss should not increase without bound across repeated, identical cudaq.observe() calls.
Is this a regression? If it is, put the last known working version (or commit) here.
Not a regression
Environment
- CUDA-Q version: 0.12.0, 0.15.0, 52f4442
- Python version: 3.10.12, 3.11.15, 3.12.3
- C++ compiler:
- Operating system: Ubuntu 22.04.4 LTS, Ubuntu 26.04 LTS, Ubuntu 22.04.4 LTS
Suggestions
No response
Required prerequisites
Describe the bug
We ran into an OOM-kill while running a large-scale VQE job on our node (CUDA-Q 0.12.0). This was unexpected, since the node has about 125 GB of free RAM.
We traced it to two separate host-memory leaks triggered by repeated cudaq.observe() calls: one in the
uccsd()kernel builtin, and one in howcudaq.observe()handles a Hamiltonian built viacudaq.chemistry.create_molecular_hamiltonian().We built two minimal reproducers and confirmed both leaks on CUDA-Q 0.15.0 and on CUDA-Q amd64-cu12-latest (52f4442), so this isn't specific to 0.12.0.
Steps to reproduce the bug
Expected behavior
ru_maxrssshould not increase without bound across repeated, identicalcudaq.observe()calls.Is this a regression? If it is, put the last known working version (or commit) here.
Not a regression
Environment
Suggestions
No response