Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
51 changes: 51 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -171,6 +171,37 @@ total 0
lrwxrwxrwx. 1 root root 0 Nov 24 13:33 aa618089-8b16-4d01-a136-25a0f3c73123 -> ../../../devices/pci0000:00/0000:00:03.0/0000:03:00.0/0000:04:09.0/0000:06:00.0/aa618089-8b16-4d01-a136-25a0f3c73123
```

--------------------------------------------------------------
### Preparing a GPU to be used in vGPU mode (Ada Lovelace / Hopper and newer)

Starting with vGPU release 17, GPUs based on Ada Lovelace and newer architectures (for example L40S, H100, H200) no longer expose vGPU instances through mdev. `/sys/bus/mdev` does not exist on these hosts. Instead, vGPU profiles are assigned directly on each SR-IOV Virtual Function through a vendor-specific VFIO sysfs interface, and this plugin discovers vGPUs on such hosts through that interface instead of mdev. The steps below replace the mdev steps above; do not use both on the same GPU.

##### 1. Enable SR-IOV Virtual Functions on the physical GPU.
```shell
$ /usr/lib/nvidia/sriov-manage -e 0000:41:00.0
```
**0000:41:00.0** -- The PCI BDF of the physical GPU (Physical Function).

##### 2. Find the vGPU types that can be created on a Virtual Function.
```shell
$ cat /sys/bus/pci/devices/0000\:41\:00.4/nvidia/creatable_vgpu_types
1428 : NVIDIA H200X-141C
1414 : NVIDIA H200X-1-18C
```
Each line lists a numeric type id and the corresponding profile name.

##### 3. Create a vGPU on the Virtual Function by writing its type id.
```shell
$ echo 1414 > /sys/bus/pci/devices/0000\:41\:00.4/nvidia/current_vgpu_type
```

##### 4. Confirm the vGPU was created.
```shell
$ cat /sys/bus/pci/devices/0000\:41\:00.4/nvidia/current_vgpu_type
1414
```
A non-zero value confirms the Virtual Function now has a vGPU profile configured; this plugin advertises it as a `nvidia.com/<profile-name>` resource, grouped separately from Virtual Functions configured with a different profile.

## Docs
### Deployment
The Daemonset creation yaml can be used to deploy the device plugin.
Expand All @@ -180,6 +211,26 @@ kubectl apply -f nvidia-kubevirt-gpu-device-plugin.yaml

Example YAML files for creating VMs with GPU/vGPU are in the `examples` folder

#### vGPU profile timing (vendor-specific VFIO)

By default the plugin scans the vendor-specific VFIO vGPU Virtual Functions once at startup and does not re-scan while running. Every profile a node should advertise must therefore be configured on its Virtual Functions (step 3 of "Preparing a GPU to be used in vGPU mode (Ada Lovelace / Hopper and newer)") **before** the device plugin pod starts. A Virtual Function whose profile is created or changed after startup is not picked up until the pod restarts. After a node reboot, order the plugin after whatever recreates the Virtual Functions (for example an init container that waits until the count of configured Virtual Functions is non-zero and stable) so it discovers the full set.

Optionally, set the `VFIO_VGPU_RESCAN_INTERVAL` env var to a Go duration (for example `30s`) to enable periodic rediscovery. The plugin then re-runs the same scan on that interval and, while running, starts advertising a newly configured profile, updates the device count of a profile whose set of Virtual Functions changed, and stops advertising a profile once none of its Virtual Functions remain configured. Rediscovery is disabled when the variable is unset, empty or non-positive, so the default behavior above is unchanged; a positive value below five seconds is clamped up to five seconds. A periodic rescan is used rather than an inotify watch because inotify does not fire reliably for the sysfs attribute writes that (re)configure a vGPU, and SR-IOV Virtual Function creation/removal moves whole device directories that are awkward to watch correctly. On a fully consumed card each rescan may consult NVML (see the NVML fallback section below), so choose an interval no shorter than you need.

#### NVML fallback (fully consumed nodes)

The profile name of a configured Virtual Function is normally read from its card's `creatable_vgpu_types` catalog. On a fully consumed card that catalog is reduced to its header on every function, so the plugin resolves the configured type ids through NVML instead (`GetSupportedVgpus`, whose list does not shrink as capacity is allocated). This path is reached only on fully consumed nodes; the default manifest does not enable it.

NVML needs two things inside the container:

- **`libnvidia-ml.so.1`** from the host driver, loadable by the dynamic linker. Bind the **single file**, not the host library directory: putting the host lib directory on `LD_LIBRARY_PATH` drags the host glibc into the container and breaks the dynamic linker. This single-file requirement was confirmed on hardware.
- **The NVIDIA device nodes** the queries touch: `/dev/nvidiactl` (the NVML control device) and the per-GPU `/dev/nvidiaN` for each physical card. The management-only queries this plugin makes (`DeviceGetHandleByPciBusId`, `GetSupportedVgpus`, `GetName`) do not open a CUDA context or MIG capabilities, so `/dev/nvidia-uvm`, `/dev/nvidia-uvm-tools` and `/dev/nvidia-caps` are not needed. `/dev/nvidiactl` plus the per-GPU nodes were confirmed on hardware; the narrower "uvm/caps not required" scope is inferred from the set of NVML calls above, not separately tested.

There are two supported ways to provide these:

1. **NVIDIA Container Toolkit (`runtimeClassName: nvidia`) — minimally privileged, recommended.** When the toolkit is installed on the node, set `runtimeClassName: nvidia` and the container env `NVIDIA_VISIBLE_DEVICES=all` and `NVIDIA_DRIVER_CAPABILITIES=utility` (the `utility` capability is the one that provides NVML). The toolkit's runtime hook then injects `libnvidia-ml.so.1` and the device nodes and adds the matching device-cgroup rules, so the container stays non-privileged. This is the standard toolkit mechanism; the exact `runtimeClassName` form was not part of this feature's hardware validation.
2. **hostPath — privileged.** `manifests/nvidia-kubevirt-gpu-device-plugin-nvml.yaml` single-file-binds `libnvidia-ml.so.1`, sets `LD_LIBRARY_PATH`, mounts only `/dev/nvidiactl` and the per-GPU `/dev/nvidiaN` nodes (not the whole host `/dev`) and runs `privileged: true`. A hostPath device node is not added to the container's device cgroup, so a non-privileged container is denied when it opens the node regardless of the mount, and a plain pod spec has no field to grant a single host device node; `privileged: true` is the only pod-spec-native way to make host device nodes usable without the toolkit. The hardware validation of the fallback used this hostPath + privileged mechanism with the whole host `/dev` mounted; this manifest narrows that to `/dev/nvidiactl` plus the per-GPU nodes, which are the only device nodes the NVML calls above use. Point the `libnvidia-ml.so.1` hostPath at wherever the host driver installed it (find it with `ldconfig -p | grep libnvidia-ml`).

### Build

Build executable binary using make:
Expand Down
Loading