Skip to content

Memory leak during large scale VQE: uccsd() and molecular Hamiltonian in cudaq.observe() #5465

Description

@TempRep940

Required prerequisites

  • Consult the security policy. If reporting a security vulnerability, do not report the bug using this form. Use the process described in the policy to report the issue.
  • Make sure you've read the documentation. Your issue may be addressed there.
  • Search the issue tracker to verify that this hasn't already been reported. +1 or comment there if it has.
  • If possible, make a PR with a failing test to give us a starting point to work on!

Describe the bug

We ran into an OOM-kill while running a large-scale VQE job on our node (CUDA-Q 0.12.0). This was unexpected, since the node has about 125 GB of free RAM.

We traced it to two separate host-memory leaks triggered by repeated cudaq.observe() calls: one in the uccsd() kernel builtin, and one in how cudaq.observe() handles a Hamiltonian built via cudaq.chemistry.create_molecular_hamiltonian().

We built two minimal reproducers and confirmed both leaks on CUDA-Q 0.15.0 and on CUDA-Q amd64-cu12-latest (52f4442), so this isn't specific to 0.12.0.

Steps to reproduce the bug

import argparse
import resource

import cudaq
from cudaq import spin
from cudaq.kernels import uccsd, uccsd_num_parameters


def main():
    parser = argparse.ArgumentParser(description=__doc__,
                                      formatter_class=argparse.RawDescriptionHelpFormatter)
    parser.add_argument('--n-electrons', type=int, default=2)
    parser.add_argument('--n-qubits', type=int, default=4)
    parser.add_argument('--n-calls', type=int, default=300)
    parser.add_argument('--log-every', type=int, default=20)
    args = parser.parse_args()

    cudaq.set_target('nvidia', option='fp64')

    n_e, num_qubits = args.n_electrons, args.n_qubits
    num_params = uccsd_num_parameters(n_e, num_qubits)

    @cudaq.kernel
    def ansatz(thetas: list[float]):
        q = cudaq.qvector(num_qubits)
        for i in range(n_e):
            x(q[i])
        uccsd(q, thetas, n_e, num_qubits)

    # Arbitrary, fixed, non-chemistry Hamiltonian -- any valid spin_op on
    # `num_qubits` qubits reproduces this; the specific terms don't matter.
    hamiltonian = spin.z(0) * spin.z(1) + spin.x(1) * spin.x(num_qubits - 1) + spin.z(num_qubits - 1)
    theta = [0.0] * num_params  # FIXED for the entire run -- never changes, no optimizer involved

    print(f"n_electrons={n_e} n_qubits={num_qubits} num_params={num_params}")
    print(f"{'call':>6} {'rss_MB':>10}")
    for i in range(1, args.n_calls + 1):
        cudaq.observe(ansatz, hamiltonian, theta)
        if i == 1 or i % args.log_every == 0:
            rss_mb = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
            print(f"{i:6d} {rss_mb:10.1f}")


if __name__ == "__main__":
    main()
import argparse
import resource

import cudaq


def main():
    parser = argparse.ArgumentParser(description=__doc__,
                                      formatter_class=argparse.RawDescriptionHelpFormatter)
    parser.add_argument('--n-calls', type=int, default=2000)
    parser.add_argument('--log-every', type=int, default=200)
    args = parser.parse_args()

    cudaq.set_target('nvidia', option='fp64')

    # H2/STO-3G: smallest possible nontrivial molecular Hamiltonian.
    geometry = [('H', (0.0, 0.0, 0.0)), ('H', (0.0, 0.0, 0.74))]
    molecule, _ = cudaq.chemistry.create_molecular_hamiltonian(
        geometry=geometry, basis='sto-3g', multiplicity=1, charge=0,
        n_active_electrons=2, n_active_orbitals=2,
    )
    num_qubits = molecule.qubit_count

    # Deliberately NOT uccsd() -- a trivial fixed circuit, to isolate the
    # Hamiltonian object as the only chemistry-related ingredient.
    @cudaq.kernel
    def ansatz(theta: float):
        q = cudaq.qvector(num_qubits)
        for i in range(num_qubits):
            ry(theta, q[i])
        for i in range(num_qubits - 1):
            cx(q[i], q[i + 1])

    theta = 0.0  # FIXED for the entire run -- never changes, no optimizer involved

    print(f"qubits={num_qubits}")
    print(f"{'call':>6} {'rss_MB':>10}")
    for i in range(1, args.n_calls + 1):
        cudaq.observe(ansatz, molecule, theta)
        if i == 1 or i % args.log_every == 0:
            rss_mb = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
            print(f"{i:6d} {rss_mb:10.1f}")


if __name__ == "__main__":
    main()

Expected behavior

ru_maxrss should not increase without bound across repeated, identical cudaq.observe() calls.

Is this a regression? If it is, put the last known working version (or commit) here.

Not a regression

Environment

  • CUDA-Q version: 0.12.0, 0.15.0, 52f4442
  • Python version: 3.10.12, 3.11.15, 3.12.3
  • C++ compiler:
  • Operating system: Ubuntu 22.04.4 LTS, Ubuntu 26.04 LTS, Ubuntu 22.04.4 LTS

Suggestions

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions