diff --git a/docs/installation/configure-cdi.md b/docs/installation/configure-cdi.md new file mode 100644 index 00000000..65d79ba3 --- /dev/null +++ b/docs/installation/configure-cdi.md @@ -0,0 +1,146 @@ +--- +title: Enable NVIDIA CDI support for HAMi +sidebar_label: NVIDIA CDI support +--- + +This page explains what CDI does and which Helm chart values enable HAMi's NVIDIA CDI integration. + +## What is CDI? + +[CDI (Container Device Interface)](https://github.com/cncf-tags/container-device-interface) is a specification for describing the OCI configuration that a container needs to use a device. A CDI specification can define device nodes, mounts, environment variables, and OCI hooks in JSON or YAML. + +CDI is vendor-neutral. HAMi currently supports CDI device injection only for NVIDIA GPUs. The HAMi configuration and troubleshooting guidance on this page do not apply to devices from other vendors. + +CDI identifies each device by a fully qualified name: + +```text +vendor.com/class=device-name +``` + +When a container runtime receives this name, it finds the matching CDI specification in a directory such as `/etc/cdi` or `/var/run/cdi` and applies the specified changes to the container's OCI runtime specification. + +CDI does not schedule devices, allocate resources, or enforce quotas. In HAMi's NVIDIA CDI integration, the scheduler and NVIDIA Device Plugin continue to select and allocate GPUs; CDI injects the allocated GPUs and their runtime requirements into containers. + +## What problems does CDI solve? + +Injecting a complex device often requires more than mounting a single `/dev` node. A GPU container may also need driver libraries, several device nodes, environment variables, and lifecycle hooks. Without a common interface, vendors must adapt this logic to individual container runtimes, and runtimes may need vendor-specific implementations. + +CDI moves these OCI changes into a common specification. It provides: + +- A standard format for device nodes, mounts, environment variables, and hooks. +- Stable device names for passing allocation results from Device Plugins to container runtimes. +- Less dependency on runtime-specific device injection implementations. +- An inspectable record of the device configuration applied to a container. + +## How HAMi uses CDI for NVIDIA GPUs + +With `cdi-annotations` enabled for the NVIDIA Device Plugin, HAMi uses the following injection flow: + +1. The HAMi scheduler and NVIDIA Device Plugin select and allocate a GPU. +2. The Device Plugin generates a CDI specification and writes it to `/var/run/cdi` on the node. +3. The Device Plugin returns the allocated device name through a CDI annotation. +4. A CDI-aware container runtime resolves the annotation and applies the matching CDI specification to the container's OCI runtime specification. + +HAMi uses the following CDI kind for NVIDIA GPUs: + +```text +k8s.device-plugin.nvidia.com/gpu +``` + +## Prerequisites + +Before enabling NVIDIA CDI support in HAMi, confirm that: + +- The NVIDIA driver and NVIDIA Container Toolkit are installed on each GPU node. +- The container runtime supports CDI and has CDI enabled. +- The container runtime can read CDI specifications from `/var/run/cdi`. +- The actual NVIDIA driver root and `nvidia-ctk` path on each node are known. + +For runtime-specific instructions, see the [upstream CDI configuration guide](https://github.com/cncf-tags/container-device-interface#how-to-configure-cdi). + +## Configure the Helm chart + +Save the following values as `values-cdi.yaml`: + +```yaml +devicePlugin: + deviceListStrategy: "cdi-annotations" + nvidiaDriverRoot: "" + nvidiaHookPath: "" +``` + +- `devicePlugin.deviceListStrategy`: Set this to `cdi-annotations` so that the HAMi Device Plugin passes device names through CDI annotations. +- `devicePlugin.nvidiaDriverRoot`: The NVIDIA driver root on the node. HAMi uses this directory to discover driver files and device nodes. +- `devicePlugin.nvidiaHookPath`: The `nvidia-ctk` executable used on the node. HAMi writes this path into the generated CDI specification. + +Replace the two path placeholders according to the deployment: + +- Host-installed NVIDIA driver and Container Toolkit: use `/` for `` and `/usr/bin/nvidia-ctk` for ``. +- GPU Operator-managed driver and Container Toolkit: use `/run/nvidia/driver` for `` and `/usr/local/nvidia/toolkit/nvidia-ctk` for ``. + +:::warning + +The two path values must match the actual node filesystem. Do not copy values from a different deployment model. `nvidiaHookPath` must point to an executable that is available to the container runtime on the node. + +::: + +Install or upgrade HAMi: + +```bash +helm upgrade --install hami hami-charts/hami \ + --namespace kube-system \ + --create-namespace \ + --values values-cdi.yaml +``` + +## Verify the configuration + +First, confirm that Helm applied the expected values: + +```bash +helm get values hami -n kube-system +``` + +On a GPU node, check the selected `nvidia-ctk` executable and the CDI specification generated by HAMi. The following example uses the host-installed path; replace `NVIDIA_CTK_PATH` in GPU Operator environments: + +```bash +NVIDIA_CTK_PATH=/usr/bin/nvidia-ctk +sudo test -x "$NVIDIA_CTK_PATH" +"$NVIDIA_CTK_PATH" --version +sudo ls -l /var/run/cdi +sudo jq . /var/run/cdi/k8s.device-plugin.nvidia.com-gpu.json +``` + +Confirm that: + +- The CDI kind is `k8s.device-plugin.nvidia.com/gpu`. +- The allocated GPU UUID is present in the device list. +- Every device node and hook path in the specification exists on the node. + +Finally, check the Device Plugin logs: + +```bash +kubectl logs -n kube-system \ + -l app.kubernetes.io/component=hami-device-plugin \ + --tail=200 | grep -i cdi +``` + +The logs should not contain errors about CDI specification generation, unresolved devices, or missing hook paths. + +## Troubleshooting + +### `failed to stat CDI host device "/dev/nvidia-modeset": no such file or directory` + +`devicePlugin.nvidiaDriverRoot` does not match the driver layout on the node. Correct the value using the path rules above, upgrade HAMi, and inspect the regenerated CDI specification. + +### `stderr: No help topic for 'disable-device-node-modification'` + +The NVIDIA Container Toolkit executable referenced by the CDI specification does not support the `disable-device-node-modification` command. This command requires NVIDIA Container Toolkit 1.18.0 or later. + +Upgrade NVIDIA Container Toolkit and verify both the hook path recorded in the CDI specification and the version of the executable at that path. Host-installed and GPU Operator-managed Toolkit installations are independent; upgrading one does not update the other. + +[GPU Operator v25.10.1](https://github.com/NVIDIA/gpu-operator/blob/v25.10.1/deployments/gpu-operator/values.yaml#L220-L224) uses NVIDIA Container Toolkit v1.18.1. For related details, see [Toolkit version parser broken for casual kind demo user](https://github.com/kubernetes-sigs/dra-driver-nvidia-gpu/issues/568). + +### `error running createContainer hook #0: fork/exec /usr/bin/nvidia-ctk: no such file or directory` + +The hook path recorded in the CDI specification does not exist on the node. Locate the actual `nvidia-ctk` executable, correct `devicePlugin.nvidiaHookPath`, and upgrade HAMi to regenerate the CDI specification. diff --git a/docs/installation/how-to-use-hami-dra.md b/docs/installation/how-to-use-hami-dra.md index 2695115e..6db129c5 100644 --- a/docs/installation/how-to-use-hami-dra.md +++ b/docs/installation/how-to-use-hami-dra.md @@ -18,7 +18,7 @@ HAMi has provided support for K8s [DRA](https://kubernetes.io/docs/concepts/sche ## Prerequisites - Kubernetes version >= 1.34 with DRA Consumable Capacity [featuregate](https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates/) enabled -- [CDI](https://github.com/cncf-tags/container-device-interface?tab=readme-ov-file#how-to-configure-cdi) must be enabled in the underlying container runtime (such as containerd or CRI-O) +- [CDI](./configure-cdi.md) must be enabled in the underlying container runtime (such as containerd or CRI-O) - NVIDIA GPU Driver 440 or later ## Installation diff --git a/docs/installation/prerequisites.md b/docs/installation/prerequisites.md index 22d26669..59a4e880 100644 --- a/docs/installation/prerequisites.md +++ b/docs/installation/prerequisites.md @@ -17,6 +17,8 @@ This guide assumes pre-installation of NVIDIA drivers and the `nvidia-container- For details see [Installing the NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html). +To inject NVIDIA GPUs through CDI, follow [Enable NVIDIA CDI support for HAMi](./configure-cdi.md) before installing HAMi. Verify the container runtime, driver root, and NVIDIA Container Toolkit path first. + ### Example for debian-based systems with Docker and containerd #### Install the `nvidia-container-toolkit` diff --git a/i18n/zh/docusaurus-plugin-content-docs/current/installation/configure-cdi.md b/i18n/zh/docusaurus-plugin-content-docs/current/installation/configure-cdi.md new file mode 100644 index 00000000..6ccdae28 --- /dev/null +++ b/i18n/zh/docusaurus-plugin-content-docs/current/installation/configure-cdi.md @@ -0,0 +1,147 @@ +--- +title: 为 HAMi 启用 NVIDIA CDI 支持 +sidebar_label: NVIDIA CDI 支持 +translated: true +--- + +本文介绍 CDI 的作用,以及为 HAMi 配置 NVIDIA CDI 支持所需的 Helm Chart 参数。 + +## CDI 是什么 + +[CDI(Container Device Interface)](https://github.com/cncf-tags/container-device-interface)是一项容器设备接口规范。它使用 JSON 或 YAML 文件描述容器使用某个设备时需要应用的 OCI 配置,包括设备节点、文件挂载、环境变量和 OCI hook。 + +CDI 是与设备厂商无关的通用规范。HAMi 当前仅支持通过 CDI 注入 NVIDIA GPU;本文中的 HAMi 配置和排错方法不适用于其他厂商设备。 + +CDI 使用完全限定的设备名称标识设备: + +```text +vendor.com/class=device-name +``` + +容器运行时收到设备名称后,会从 `/etc/cdi`、`/var/run/cdi` 等目录查找对应的 CDI spec,并将其中的配置写入容器的 OCI runtime spec。 + +CDI 本身不负责设备调度、资源分配或配额管理。在 HAMi 的 NVIDIA CDI 集成中,调度器和 NVIDIA Device Plugin 仍然负责选择、分配 GPU;CDI 负责把已经分配的 GPU 及其运行环境注入容器。 + +## CDI 解决什么问题 + +复杂设备通常不能只靠挂载一个 `/dev` 设备节点完成注入。GPU 容器还可能需要驱动库、多个设备节点、环境变量和容器生命周期 hook。没有统一接口时,设备厂商需要分别适配不同的容器运行时,运行时中也容易出现厂商专用逻辑。 + +CDI 将这些设备相关的 OCI 配置集中写入 spec,主要解决以下问题: + +- 使用统一格式描述设备节点、挂载、环境变量和 hook。 +- 使用稳定的设备名称在 Device Plugin 与容器运行时之间传递分配结果。 +- 减少设备注入逻辑对特定容器运行时实现的依赖。 +- 便于检查容器最终使用了哪些设备配置。 + +## HAMi 如何为 NVIDIA GPU 使用 CDI + +为 NVIDIA Device Plugin 启用 `cdi-annotations` 后,HAMi 的设备注入流程如下: + +1. HAMi 调度器和 NVIDIA Device Plugin 完成 GPU 选择与分配。 +2. NVIDIA Device Plugin 生成 HAMi 使用的 CDI spec,并写入节点的 `/var/run/cdi` 目录。 +3. Device Plugin 通过 CDI annotation 返回分配到的设备名称。 +4. 支持 CDI 的容器运行时读取 annotation,查找对应的 CDI spec,并将设备配置加入容器的 OCI runtime spec。 + +HAMi 生成的 NVIDIA GPU CDI kind 为: + +```text +k8s.device-plugin.nvidia.com/gpu +``` + +## 前置条件 + +启用 HAMi 的 NVIDIA CDI 支持前,需要确认以下条件: + +- 节点已经安装 NVIDIA 驱动和 NVIDIA Container Toolkit。 +- 容器运行时支持并已启用 CDI。 +- 容器运行时能够读取 `/var/run/cdi` 中的 CDI spec。 +- 已确认 NVIDIA 驱动和 `nvidia-ctk` 在节点上的实际路径。 + +不同容器运行时的 CDI 配置方式,参阅 [CDI 上游配置说明](https://github.com/cncf-tags/container-device-interface#how-to-configure-cdi)。 + +## 配置 Helm Chart + +将以下内容保存为 `values-cdi.yaml`: + +```yaml +devicePlugin: + deviceListStrategy: "cdi-annotations" + nvidiaDriverRoot: "" + nvidiaHookPath: "" +``` + +- `devicePlugin.deviceListStrategy`:固定设置为 `cdi-annotations`,让 HAMi Device Plugin 通过 CDI annotation 传递设备名称。 +- `devicePlugin.nvidiaDriverRoot`:指定节点上的 NVIDIA 驱动根目录,HAMi 使用该目录发现驱动文件和设备节点。 +- `devicePlugin.nvidiaHookPath`:指定节点上实际执行的 `nvidia-ctk`,HAMi 将该路径写入生成的 CDI spec。 + +根据部署方式替换两个路径占位符: + +- NVIDIA 驱动和 NVIDIA Container Toolkit 直接安装在宿主机:`` 使用 `/`,`` 使用 `/usr/bin/nvidia-ctk`。 +- GPU Operator 管理驱动和 NVIDIA Container Toolkit:`` 使用 `/run/nvidia/driver`,`` 使用 `/usr/local/nvidia/toolkit/nvidia-ctk`。 + +:::warning + +两个路径参数必须按节点上的实际文件布局设置,不能直接套用另一种部署方式的值。`nvidiaHookPath` 必须是容器运行时所在节点能够执行的路径。 + +::: + +安装或升级 HAMi: + +```bash +helm upgrade --install hami hami-charts/hami \ + --namespace kube-system \ + --create-namespace \ + --values values-cdi.yaml +``` + +## 验证配置 + +先确认 Helm 使用了预期配置: + +```bash +helm get values hami -n kube-system +``` + +然后在 GPU 节点上检查所选 `nvidia-ctk` 路径和 HAMi 生成的 CDI spec。下面以宿主机安装路径为例;GPU Operator 环境需要替换 `NVIDIA_CTK_PATH`: + +```bash +NVIDIA_CTK_PATH=/usr/bin/nvidia-ctk +sudo test -x "$NVIDIA_CTK_PATH" +"$NVIDIA_CTK_PATH" --version +sudo ls -l /var/run/cdi +sudo jq . /var/run/cdi/k8s.device-plugin.nvidia.com-gpu.json +``` + +重点确认以下内容: + +- CDI kind 为 `k8s.device-plugin.nvidia.com/gpu`。 +- 分配使用的 GPU UUID 出现在设备列表中。 +- spec 中记录的设备节点和 hook 路径在节点上真实存在。 + +最后检查 Device Plugin 日志: + +```bash +kubectl logs -n kube-system \ + -l app.kubernetes.io/component=hami-device-plugin \ + --tail=200 | grep -i cdi +``` + +日志中不应出现 CDI spec 生成失败、设备无法解析或 hook 路径不存在等错误。 + +## 常见问题 + +### `failed to stat CDI host device "/dev/nvidia-modeset": no such file or directory` + +`devicePlugin.nvidiaDriverRoot` 与节点的驱动布局不一致。按前文的路径规则修正该参数,升级 HAMi,然后重新检查生成的 CDI spec。 + +### `stderr: No help topic for 'disable-device-node-modification'` + +该错误表示实际执行的 NVIDIA Container Toolkit 不支持 CDI spec 中使用的 `disable-device-node-modification` 命令。该命令需要 NVIDIA Container Toolkit 1.18.0 或更高版本。 + +将 NVIDIA Container Toolkit 升级到 1.18.0 或更高版本,并检查 CDI spec 中记录的 hook 路径及其实际版本。如果同时存在宿主机安装和 GPU Operator 容器化安装,需要分别检查;升级其中一套不会更新另一套。 + +[GPU Operator v25.10.1](https://github.com/NVIDIA/gpu-operator/blob/v25.10.1/deployments/gpu-operator/values.yaml#L220-L224) 使用 NVIDIA Container Toolkit v1.18.1。相关问题参阅 [Toolkit version parser broken for casual kind demo user](https://github.com/kubernetes-sigs/dra-driver-nvidia-gpu/issues/568)。 + +### `error running createContainer hook #0: fork/exec /usr/bin/nvidia-ctk: no such file or directory` + +CDI spec 中记录的 hook 路径在节点上不存在。确认 `nvidia-ctk` 的实际位置,修正 `devicePlugin.nvidiaHookPath`,升级 HAMi 以重新生成 CDI spec。 diff --git a/i18n/zh/docusaurus-plugin-content-docs/current/installation/how-to-use-hami-dra.md b/i18n/zh/docusaurus-plugin-content-docs/current/installation/how-to-use-hami-dra.md index 7293b584..ba8593af 100644 --- a/i18n/zh/docusaurus-plugin-content-docs/current/installation/how-to-use-hami-dra.md +++ b/i18n/zh/docusaurus-plugin-content-docs/current/installation/how-to-use-hami-dra.md @@ -18,7 +18,7 @@ HAMi 已经提供了对 K8s [DRA](https://kubernetes.io/docs/concepts/scheduling ## 前提条件 - Kubernetes 版本 >= 1.34 并且 DRA Consumable Capacity [featuregate](https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates/) 已启用 -- 底层容器运行时(例如 containerd 或 CRI-O)必须启用 [CDI](https://github.com/cncf-tags/container-device-interface?tab=readme-ov-file#how-to-configure-cdi) +- 底层容器运行时(例如 containerd 或 CRI-O)必须启用 [CDI](./configure-cdi.md) - NVIDIA GPU 驱动版本 440 及以上 ## 安装 diff --git a/i18n/zh/docusaurus-plugin-content-docs/current/installation/prerequisites.md b/i18n/zh/docusaurus-plugin-content-docs/current/installation/prerequisites.md index 5bd5d3dd..4cf563d5 100644 --- a/i18n/zh/docusaurus-plugin-content-docs/current/installation/prerequisites.md +++ b/i18n/zh/docusaurus-plugin-content-docs/current/installation/prerequisites.md @@ -19,6 +19,8 @@ title: 前置条件 参阅[安装 NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html)。 +如需使用 CDI 注入 NVIDIA GPU,请在安装 HAMi 前参阅[为 HAMi 启用 NVIDIA CDI 支持](./configure-cdi.md),确认容器运行时、驱动根目录和 NVIDIA Container Toolkit 路径。 + ### Debian 系统示例(Docker + containerd) #### 安装 `nvidia-container-toolkit` diff --git a/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/configure-cdi.md b/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/configure-cdi.md new file mode 100644 index 00000000..c1476b1a --- /dev/null +++ b/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/configure-cdi.md @@ -0,0 +1,148 @@ +--- +title: 为 HAMi 启用 NVIDIA CDI 支持 +sidebar_label: NVIDIA CDI 支持 +translated: true +--- + +本文介绍 CDI 的作用,以及为 HAMi 配置 NVIDIA CDI 支持所需的 Helm Chart 参数。 + +## CDI 是什么 + +[CDI(Container Device Interface)](https://github.com/cncf-tags/container-device-interface)是一项容器设备接口规范。它使用 JSON 或 YAML 文件描述容器使用某个设备时需要应用的 OCI 配置,包括设备节点、文件挂载、环境变量和 OCI hook。 + +CDI 是与设备厂商无关的通用规范。HAMi 当前仅支持通过 CDI 注入 NVIDIA GPU;本文中的 HAMi 配置和排错方法不适用于其他厂商设备。 + +CDI 使用完全限定的设备名称标识设备: + +```text +vendor.com/class=device-name +``` + +容器运行时收到设备名称后,会从 `/etc/cdi`、`/var/run/cdi` 等目录查找对应的 CDI spec,并将其中的配置写入容器的 OCI runtime spec。 + +CDI 本身不负责设备调度、资源分配或配额管理。在 HAMi 的 NVIDIA CDI 集成中,调度器和 NVIDIA Device Plugin 仍然负责选择、分配 GPU;CDI 负责把已经分配的 GPU 及其运行环境注入容器。 + +## CDI 解决什么问题 + +复杂设备通常不能只靠挂载一个 `/dev` 设备节点完成注入。GPU 容器还可能需要驱动库、多个设备节点、环境变量和容器生命周期 hook。没有统一接口时,设备厂商需要分别适配不同的容器运行时,运行时中也容易出现厂商专用逻辑。 + +CDI 将这些设备相关的 OCI 配置集中写入 spec,主要解决以下问题: + +- 使用统一格式描述设备节点、挂载、环境变量和 hook。 +- 使用稳定的设备名称在 Device Plugin 与容器运行时之间传递分配结果。 +- 减少设备注入逻辑对特定容器运行时实现的依赖。 +- 便于检查容器最终使用了哪些设备配置。 + +## HAMi 如何为 NVIDIA GPU 使用 CDI + +为 NVIDIA Device Plugin 启用 `cdi-annotations` 后,HAMi 的设备注入流程如下: + +1. HAMi 调度器和 NVIDIA Device Plugin 完成 GPU 选择与分配。 +2. NVIDIA Device Plugin 生成 HAMi 使用的 CDI spec,并写入节点的 `/var/run/cdi` 目录。 +3. Device Plugin 通过 CDI annotation 返回分配到的设备名称。 +4. 支持 CDI 的容器运行时读取 annotation,查找对应的 CDI spec,并将设备配置加入容器的 OCI runtime spec。 + +HAMi 生成的 NVIDIA GPU CDI kind 为: + +```text +k8s.device-plugin.nvidia.com/gpu +``` + +## 前置条件 + +启用 HAMi 的 NVIDIA CDI 支持前,需要确认以下条件: + +- 节点已经安装 NVIDIA 驱动和 NVIDIA Container Toolkit。 +- 容器运行时支持并已启用 CDI。 +- 容器运行时能够读取 `/var/run/cdi` 中的 CDI spec。 +- 已确认 NVIDIA 驱动和 `nvidia-ctk` 在节点上的实际路径。 + +不同容器运行时的 CDI 配置方式,参阅 [CDI 上游配置说明](https://github.com/cncf-tags/container-device-interface#how-to-configure-cdi)。 + +## 配置 Helm Chart + +将以下内容保存为 `values-cdi.yaml`: + +```yaml +devicePlugin: + deviceListStrategy: "cdi-annotations" + nvidiaDriverRoot: "" + nvidiaHookPath: "" +``` + +- `devicePlugin.deviceListStrategy`:固定设置为 `cdi-annotations`,让 HAMi Device Plugin 通过 CDI annotation 传递设备名称。 +- `devicePlugin.nvidiaDriverRoot`:指定节点上的 NVIDIA 驱动根目录,HAMi 使用该目录发现驱动文件和设备节点。 +- `devicePlugin.nvidiaHookPath`:指定节点上实际执行的 `nvidia-ctk`,HAMi 将该路径写入生成的 CDI spec。 + +根据部署方式替换两个路径占位符: + +- NVIDIA 驱动和 NVIDIA Container Toolkit 直接安装在宿主机:`` 使用 `/`,`` 使用 `/usr/bin/nvidia-ctk`。 +- GPU Operator 管理驱动和 NVIDIA Container Toolkit:`` 使用 `/run/nvidia/driver`,`` 使用 `/usr/local/nvidia/toolkit/nvidia-ctk`。 + +:::warning + +两个路径参数必须按节点上的实际文件布局设置,不能直接套用另一种部署方式的值。`nvidiaHookPath` 必须是容器运行时所在节点能够执行的路径。 + +::: + +安装或升级 HAMi: + +```bash +helm upgrade --install hami hami-charts/hami \ + --version 2.8.0 \ + --namespace kube-system \ + --create-namespace \ + --values values-cdi.yaml +``` + +## 验证配置 + +先确认 Helm 使用了预期配置: + +```bash +helm get values hami -n kube-system +``` + +然后在 GPU 节点上检查所选 `nvidia-ctk` 路径和 HAMi 生成的 CDI spec。下面以宿主机安装路径为例;GPU Operator 环境需要替换 `NVIDIA_CTK_PATH`: + +```bash +NVIDIA_CTK_PATH=/usr/bin/nvidia-ctk +sudo test -x "$NVIDIA_CTK_PATH" +"$NVIDIA_CTK_PATH" --version +sudo ls -l /var/run/cdi +sudo jq . /var/run/cdi/k8s.device-plugin.nvidia.com-gpu.json +``` + +重点确认以下内容: + +- CDI kind 为 `k8s.device-plugin.nvidia.com/gpu`。 +- 分配使用的 GPU UUID 出现在设备列表中。 +- spec 中记录的设备节点和 hook 路径在节点上真实存在。 + +最后检查 Device Plugin 日志: + +```bash +kubectl logs -n kube-system \ + -l app.kubernetes.io/component=hami-device-plugin \ + --tail=200 | grep -i cdi +``` + +日志中不应出现 CDI spec 生成失败、设备无法解析或 hook 路径不存在等错误。 + +## 常见问题 + +### `failed to stat CDI host device "/dev/nvidia-modeset": no such file or directory` + +`devicePlugin.nvidiaDriverRoot` 与节点的驱动布局不一致。按前文的路径规则修正该参数,升级 HAMi,然后重新检查生成的 CDI spec。 + +### `stderr: No help topic for 'disable-device-node-modification'` + +该错误表示实际执行的 NVIDIA Container Toolkit 不支持 CDI spec 中使用的 `disable-device-node-modification` 命令。该命令需要 NVIDIA Container Toolkit 1.18.0 或更高版本。 + +将 NVIDIA Container Toolkit 升级到 1.18.0 或更高版本,并检查 CDI spec 中记录的 hook 路径及其实际版本。如果同时存在宿主机安装和 GPU Operator 容器化安装,需要分别检查;升级其中一套不会更新另一套。 + +[GPU Operator v25.10.1](https://github.com/NVIDIA/gpu-operator/blob/v25.10.1/deployments/gpu-operator/values.yaml#L220-L224) 使用 NVIDIA Container Toolkit v1.18.1。相关问题参阅 [Toolkit version parser broken for casual kind demo user](https://github.com/kubernetes-sigs/dra-driver-nvidia-gpu/issues/568)。 + +### `error running createContainer hook #0: fork/exec /usr/bin/nvidia-ctk: no such file or directory` + +CDI spec 中记录的 hook 路径在节点上不存在。确认 `nvidia-ctk` 的实际位置,修正 `devicePlugin.nvidiaHookPath`,升级 HAMi 以重新生成 CDI spec。 diff --git a/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/how-to-use-hami-dra.md b/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/how-to-use-hami-dra.md index a5fdcdfa..9bac135d 100644 --- a/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/how-to-use-hami-dra.md +++ b/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/how-to-use-hami-dra.md @@ -12,6 +12,7 @@ HAMi 已经提供了对 K8s [DRA](https://kubernetes.io/docs/concepts/scheduling ## 前提条件 - Kubernetes 版本 >= 1.34 并且 DRA Consumable Capacity [featuregate](https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates/) 启用 +- 底层容器运行时(例如 containerd 或 CRI-O)必须启用 [CDI](./configure-cdi.md) ## 安装 diff --git a/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/prerequisites.md b/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/prerequisites.md index 70966793..95b90de5 100644 --- a/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/prerequisites.md +++ b/i18n/zh/docusaurus-plugin-content-docs/version-v2.8.0/installation/prerequisites.md @@ -20,6 +20,8 @@ translated: true 参阅:[https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html) +如需使用 CDI 注入 NVIDIA GPU,请在安装 HAMi 前参阅[为 HAMi 启用 NVIDIA CDI 支持](./configure-cdi.md),确认容器运行时、驱动根目录和 NVIDIA Container Toolkit 路径。 + ### 适用于基于 Debian 系统的 `Docker` 和 `containerd` 示例 #### 安装 `nvidia-container-toolkit` diff --git a/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/configure-cdi.md b/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/configure-cdi.md new file mode 100644 index 00000000..717eb911 --- /dev/null +++ b/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/configure-cdi.md @@ -0,0 +1,148 @@ +--- +title: 为 HAMi 启用 NVIDIA CDI 支持 +sidebar_label: NVIDIA CDI 支持 +translated: true +--- + +本文介绍 CDI 的作用,以及为 HAMi 配置 NVIDIA CDI 支持所需的 Helm Chart 参数。 + +## CDI 是什么 + +[CDI(Container Device Interface)](https://github.com/cncf-tags/container-device-interface)是一项容器设备接口规范。它使用 JSON 或 YAML 文件描述容器使用某个设备时需要应用的 OCI 配置,包括设备节点、文件挂载、环境变量和 OCI hook。 + +CDI 是与设备厂商无关的通用规范。HAMi 当前仅支持通过 CDI 注入 NVIDIA GPU;本文中的 HAMi 配置和排错方法不适用于其他厂商设备。 + +CDI 使用完全限定的设备名称标识设备: + +```text +vendor.com/class=device-name +``` + +容器运行时收到设备名称后,会从 `/etc/cdi`、`/var/run/cdi` 等目录查找对应的 CDI spec,并将其中的配置写入容器的 OCI runtime spec。 + +CDI 本身不负责设备调度、资源分配或配额管理。在 HAMi 的 NVIDIA CDI 集成中,调度器和 NVIDIA Device Plugin 仍然负责选择、分配 GPU;CDI 负责把已经分配的 GPU 及其运行环境注入容器。 + +## CDI 解决什么问题 + +复杂设备通常不能只靠挂载一个 `/dev` 设备节点完成注入。GPU 容器还可能需要驱动库、多个设备节点、环境变量和容器生命周期 hook。没有统一接口时,设备厂商需要分别适配不同的容器运行时,运行时中也容易出现厂商专用逻辑。 + +CDI 将这些设备相关的 OCI 配置集中写入 spec,主要解决以下问题: + +- 使用统一格式描述设备节点、挂载、环境变量和 hook。 +- 使用稳定的设备名称在 Device Plugin 与容器运行时之间传递分配结果。 +- 减少设备注入逻辑对特定容器运行时实现的依赖。 +- 便于检查容器最终使用了哪些设备配置。 + +## HAMi 如何为 NVIDIA GPU 使用 CDI + +为 NVIDIA Device Plugin 启用 `cdi-annotations` 后,HAMi 的设备注入流程如下: + +1. HAMi 调度器和 NVIDIA Device Plugin 完成 GPU 选择与分配。 +2. NVIDIA Device Plugin 生成 HAMi 使用的 CDI spec,并写入节点的 `/var/run/cdi` 目录。 +3. Device Plugin 通过 CDI annotation 返回分配到的设备名称。 +4. 支持 CDI 的容器运行时读取 annotation,查找对应的 CDI spec,并将设备配置加入容器的 OCI runtime spec。 + +HAMi 生成的 NVIDIA GPU CDI kind 为: + +```text +k8s.device-plugin.nvidia.com/gpu +``` + +## 前置条件 + +启用 HAMi 的 NVIDIA CDI 支持前,需要确认以下条件: + +- 节点已经安装 NVIDIA 驱动和 NVIDIA Container Toolkit。 +- 容器运行时支持并已启用 CDI。 +- 容器运行时能够读取 `/var/run/cdi` 中的 CDI spec。 +- 已确认 NVIDIA 驱动和 `nvidia-ctk` 在节点上的实际路径。 + +不同容器运行时的 CDI 配置方式,参阅 [CDI 上游配置说明](https://github.com/cncf-tags/container-device-interface#how-to-configure-cdi)。 + +## 配置 Helm Chart + +将以下内容保存为 `values-cdi.yaml`: + +```yaml +devicePlugin: + deviceListStrategy: "cdi-annotations" + nvidiaDriverRoot: "" + nvidiaHookPath: "" +``` + +- `devicePlugin.deviceListStrategy`:固定设置为 `cdi-annotations`,让 HAMi Device Plugin 通过 CDI annotation 传递设备名称。 +- `devicePlugin.nvidiaDriverRoot`:指定节点上的 NVIDIA 驱动根目录,HAMi 使用该目录发现驱动文件和设备节点。 +- `devicePlugin.nvidiaHookPath`:指定节点上实际执行的 `nvidia-ctk`,HAMi 将该路径写入生成的 CDI spec。 + +根据部署方式替换两个路径占位符: + +- NVIDIA 驱动和 NVIDIA Container Toolkit 直接安装在宿主机:`` 使用 `/`,`` 使用 `/usr/bin/nvidia-ctk`。 +- GPU Operator 管理驱动和 NVIDIA Container Toolkit:`` 使用 `/run/nvidia/driver`,`` 使用 `/usr/local/nvidia/toolkit/nvidia-ctk`。 + +:::warning + +两个路径参数必须按节点上的实际文件布局设置,不能直接套用另一种部署方式的值。`nvidiaHookPath` 必须是容器运行时所在节点能够执行的路径。 + +::: + +安装或升级 HAMi: + +```bash +helm upgrade --install hami hami-charts/hami \ + --version 2.9.0 \ + --namespace kube-system \ + --create-namespace \ + --values values-cdi.yaml +``` + +## 验证配置 + +先确认 Helm 使用了预期配置: + +```bash +helm get values hami -n kube-system +``` + +然后在 GPU 节点上检查所选 `nvidia-ctk` 路径和 HAMi 生成的 CDI spec。下面以宿主机安装路径为例;GPU Operator 环境需要替换 `NVIDIA_CTK_PATH`: + +```bash +NVIDIA_CTK_PATH=/usr/bin/nvidia-ctk +sudo test -x "$NVIDIA_CTK_PATH" +"$NVIDIA_CTK_PATH" --version +sudo ls -l /var/run/cdi +sudo jq . /var/run/cdi/k8s.device-plugin.nvidia.com-gpu.json +``` + +重点确认以下内容: + +- CDI kind 为 `k8s.device-plugin.nvidia.com/gpu`。 +- 分配使用的 GPU UUID 出现在设备列表中。 +- spec 中记录的设备节点和 hook 路径在节点上真实存在。 + +最后检查 Device Plugin 日志: + +```bash +kubectl logs -n kube-system \ + -l app.kubernetes.io/component=hami-device-plugin \ + --tail=200 | grep -i cdi +``` + +日志中不应出现 CDI spec 生成失败、设备无法解析或 hook 路径不存在等错误。 + +## 常见问题 + +### `failed to stat CDI host device "/dev/nvidia-modeset": no such file or directory` + +`devicePlugin.nvidiaDriverRoot` 与节点的驱动布局不一致。按前文的路径规则修正该参数,升级 HAMi,然后重新检查生成的 CDI spec。 + +### `stderr: No help topic for 'disable-device-node-modification'` + +该错误表示实际执行的 NVIDIA Container Toolkit 不支持 CDI spec 中使用的 `disable-device-node-modification` 命令。该命令需要 NVIDIA Container Toolkit 1.18.0 或更高版本。 + +将 NVIDIA Container Toolkit 升级到 1.18.0 或更高版本,并检查 CDI spec 中记录的 hook 路径及其实际版本。如果同时存在宿主机安装和 GPU Operator 容器化安装,需要分别检查;升级其中一套不会更新另一套。 + +[GPU Operator v25.10.1](https://github.com/NVIDIA/gpu-operator/blob/v25.10.1/deployments/gpu-operator/values.yaml#L220-L224) 使用 NVIDIA Container Toolkit v1.18.1。相关问题参阅 [Toolkit version parser broken for casual kind demo user](https://github.com/kubernetes-sigs/dra-driver-nvidia-gpu/issues/568)。 + +### `error running createContainer hook #0: fork/exec /usr/bin/nvidia-ctk: no such file or directory` + +CDI spec 中记录的 hook 路径在节点上不存在。确认 `nvidia-ctk` 的实际位置,修正 `devicePlugin.nvidiaHookPath`,升级 HAMi 以重新生成 CDI spec。 diff --git a/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/how-to-use-hami-dra.md b/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/how-to-use-hami-dra.md index 7293b584..ba8593af 100644 --- a/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/how-to-use-hami-dra.md +++ b/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/how-to-use-hami-dra.md @@ -18,7 +18,7 @@ HAMi 已经提供了对 K8s [DRA](https://kubernetes.io/docs/concepts/scheduling ## 前提条件 - Kubernetes 版本 >= 1.34 并且 DRA Consumable Capacity [featuregate](https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates/) 已启用 -- 底层容器运行时(例如 containerd 或 CRI-O)必须启用 [CDI](https://github.com/cncf-tags/container-device-interface?tab=readme-ov-file#how-to-configure-cdi) +- 底层容器运行时(例如 containerd 或 CRI-O)必须启用 [CDI](./configure-cdi.md) - NVIDIA GPU 驱动版本 440 及以上 ## 安装 diff --git a/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/prerequisites.md b/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/prerequisites.md index a865a4a4..b90c5dd5 100644 --- a/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/prerequisites.md +++ b/i18n/zh/docusaurus-plugin-content-docs/version-v2.9.0/installation/prerequisites.md @@ -19,6 +19,8 @@ title: 前置条件 参阅[安装 NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html)。 +如需使用 CDI 注入 NVIDIA GPU,请在安装 HAMi 前参阅[为 HAMi 启用 NVIDIA CDI 支持](./configure-cdi.md),确认容器运行时、驱动根目录和 NVIDIA Container Toolkit 路径。 + ### Debian 系统示例(Docker + containerd) #### 安装 `nvidia-container-toolkit` diff --git a/sidebars.js b/sidebars.js index db9a89b4..828ff4e7 100644 --- a/sidebars.js +++ b/sidebars.js @@ -49,6 +49,7 @@ module.exports = { }, items: [ "installation/prerequisites", + "installation/configure-cdi", "installation/online-installation", "installation/offline-installation", "installation/upgrade", diff --git a/versioned_docs/version-v2.8.0/installation/configure-cdi.md b/versioned_docs/version-v2.8.0/installation/configure-cdi.md new file mode 100644 index 00000000..d323f4d3 --- /dev/null +++ b/versioned_docs/version-v2.8.0/installation/configure-cdi.md @@ -0,0 +1,147 @@ +--- +title: Enable NVIDIA CDI support for HAMi +sidebar_label: NVIDIA CDI support +--- + +This page explains what CDI does and which Helm chart values enable HAMi's NVIDIA CDI integration. + +## What is CDI? + +[CDI (Container Device Interface)](https://github.com/cncf-tags/container-device-interface) is a specification for describing the OCI configuration that a container needs to use a device. A CDI specification can define device nodes, mounts, environment variables, and OCI hooks in JSON or YAML. + +CDI is vendor-neutral. HAMi currently supports CDI device injection only for NVIDIA GPUs. The HAMi configuration and troubleshooting guidance on this page do not apply to devices from other vendors. + +CDI identifies each device by a fully qualified name: + +```text +vendor.com/class=device-name +``` + +When a container runtime receives this name, it finds the matching CDI specification in a directory such as `/etc/cdi` or `/var/run/cdi` and applies the specified changes to the container's OCI runtime specification. + +CDI does not schedule devices, allocate resources, or enforce quotas. In HAMi's NVIDIA CDI integration, the scheduler and NVIDIA Device Plugin continue to select and allocate GPUs; CDI injects the allocated GPUs and their runtime requirements into containers. + +## What problems does CDI solve? + +Injecting a complex device often requires more than mounting a single `/dev` node. A GPU container may also need driver libraries, several device nodes, environment variables, and lifecycle hooks. Without a common interface, vendors must adapt this logic to individual container runtimes, and runtimes may need vendor-specific implementations. + +CDI moves these OCI changes into a common specification. It provides: + +- A standard format for device nodes, mounts, environment variables, and hooks. +- Stable device names for passing allocation results from Device Plugins to container runtimes. +- Less dependency on runtime-specific device injection implementations. +- An inspectable record of the device configuration applied to a container. + +## How HAMi uses CDI for NVIDIA GPUs + +With `cdi-annotations` enabled for the NVIDIA Device Plugin, HAMi uses the following injection flow: + +1. The HAMi scheduler and NVIDIA Device Plugin select and allocate a GPU. +2. The Device Plugin generates a CDI specification and writes it to `/var/run/cdi` on the node. +3. The Device Plugin returns the allocated device name through a CDI annotation. +4. A CDI-aware container runtime resolves the annotation and applies the matching CDI specification to the container's OCI runtime specification. + +HAMi uses the following CDI kind for NVIDIA GPUs: + +```text +k8s.device-plugin.nvidia.com/gpu +``` + +## Prerequisites + +Before enabling NVIDIA CDI support in HAMi, confirm that: + +- The NVIDIA driver and NVIDIA Container Toolkit are installed on each GPU node. +- The container runtime supports CDI and has CDI enabled. +- The container runtime can read CDI specifications from `/var/run/cdi`. +- The actual NVIDIA driver root and `nvidia-ctk` path on each node are known. + +For runtime-specific instructions, see the [upstream CDI configuration guide](https://github.com/cncf-tags/container-device-interface#how-to-configure-cdi). + +## Configure the Helm chart + +Save the following values as `values-cdi.yaml`: + +```yaml +devicePlugin: + deviceListStrategy: "cdi-annotations" + nvidiaDriverRoot: "" + nvidiaHookPath: "" +``` + +- `devicePlugin.deviceListStrategy`: Set this to `cdi-annotations` so that the HAMi Device Plugin passes device names through CDI annotations. +- `devicePlugin.nvidiaDriverRoot`: The NVIDIA driver root on the node. HAMi uses this directory to discover driver files and device nodes. +- `devicePlugin.nvidiaHookPath`: The `nvidia-ctk` executable used on the node. HAMi writes this path into the generated CDI specification. + +Replace the two path placeholders according to the deployment: + +- Host-installed NVIDIA driver and Container Toolkit: use `/` for `` and `/usr/bin/nvidia-ctk` for ``. +- GPU Operator-managed driver and Container Toolkit: use `/run/nvidia/driver` for `` and `/usr/local/nvidia/toolkit/nvidia-ctk` for ``. + +:::warning + +The two path values must match the actual node filesystem. Do not copy values from a different deployment model. `nvidiaHookPath` must point to an executable that is available to the container runtime on the node. + +::: + +Install or upgrade HAMi: + +```bash +helm upgrade --install hami hami-charts/hami \ + --version 2.8.0 \ + --namespace kube-system \ + --create-namespace \ + --values values-cdi.yaml +``` + +## Verify the configuration + +First, confirm that Helm applied the expected values: + +```bash +helm get values hami -n kube-system +``` + +On a GPU node, check the selected `nvidia-ctk` executable and the CDI specification generated by HAMi. The following example uses the host-installed path; replace `NVIDIA_CTK_PATH` in GPU Operator environments: + +```bash +NVIDIA_CTK_PATH=/usr/bin/nvidia-ctk +sudo test -x "$NVIDIA_CTK_PATH" +"$NVIDIA_CTK_PATH" --version +sudo ls -l /var/run/cdi +sudo jq . /var/run/cdi/k8s.device-plugin.nvidia.com-gpu.json +``` + +Confirm that: + +- The CDI kind is `k8s.device-plugin.nvidia.com/gpu`. +- The allocated GPU UUID is present in the device list. +- Every device node and hook path in the specification exists on the node. + +Finally, check the Device Plugin logs: + +```bash +kubectl logs -n kube-system \ + -l app.kubernetes.io/component=hami-device-plugin \ + --tail=200 | grep -i cdi +``` + +The logs should not contain errors about CDI specification generation, unresolved devices, or missing hook paths. + +## Troubleshooting + +### `failed to stat CDI host device "/dev/nvidia-modeset": no such file or directory` + +`devicePlugin.nvidiaDriverRoot` does not match the driver layout on the node. Correct the value using the path rules above, upgrade HAMi, and inspect the regenerated CDI specification. + +### `stderr: No help topic for 'disable-device-node-modification'` + +The NVIDIA Container Toolkit executable referenced by the CDI specification does not support the `disable-device-node-modification` command. This command requires NVIDIA Container Toolkit 1.18.0 or later. + +Upgrade NVIDIA Container Toolkit and verify both the hook path recorded in the CDI specification and the version of the executable at that path. Host-installed and GPU Operator-managed Toolkit installations are independent; upgrading one does not update the other. + +[GPU Operator v25.10.1](https://github.com/NVIDIA/gpu-operator/blob/v25.10.1/deployments/gpu-operator/values.yaml#L220-L224) uses NVIDIA Container Toolkit v1.18.1. For related details, see [Toolkit version parser broken for casual kind demo user](https://github.com/kubernetes-sigs/dra-driver-nvidia-gpu/issues/568). + +### `error running createContainer hook #0: fork/exec /usr/bin/nvidia-ctk: no such file or directory` + +The hook path recorded in the CDI specification does not exist on the node. Locate the actual `nvidia-ctk` executable, correct `devicePlugin.nvidiaHookPath`, and upgrade HAMi to regenerate the CDI specification. diff --git a/versioned_docs/version-v2.8.0/installation/how-to-use-hami-dra.md b/versioned_docs/version-v2.8.0/installation/how-to-use-hami-dra.md index 8eb39838..70661d72 100644 --- a/versioned_docs/version-v2.8.0/installation/how-to-use-hami-dra.md +++ b/versioned_docs/version-v2.8.0/installation/how-to-use-hami-dra.md @@ -12,6 +12,7 @@ By installing the [HAMi DRA webhook](https://github.com/Project-HAMi/HAMi-DRA) i ## Prerequisites * Kubernetes version >= 1.34 with DRA Consumable Capacity [featuregate](https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates/) enabled +* [CDI](./configure-cdi.md) must be enabled in the underlying container runtime (such as containerd or CRI-O) ## Installation diff --git a/versioned_docs/version-v2.8.0/installation/prerequisites.md b/versioned_docs/version-v2.8.0/installation/prerequisites.md index 01519388..596c9491 100644 --- a/versioned_docs/version-v2.8.0/installation/prerequisites.md +++ b/versioned_docs/version-v2.8.0/installation/prerequisites.md @@ -19,6 +19,8 @@ This README assumes pre-installation of NVIDIA drivers and the `nvidia-container For details see [Installing the NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html). +To inject NVIDIA GPUs through CDI, follow [Enable NVIDIA CDI support for HAMi](./configure-cdi.md) before installing HAMi. Verify the container runtime, driver root, and NVIDIA Container Toolkit path first. + ### Example for debian-based systems with Docker and containerd #### Install the `nvidia-container-toolkit` diff --git a/versioned_docs/version-v2.9.0/installation/configure-cdi.md b/versioned_docs/version-v2.9.0/installation/configure-cdi.md new file mode 100644 index 00000000..772b7667 --- /dev/null +++ b/versioned_docs/version-v2.9.0/installation/configure-cdi.md @@ -0,0 +1,147 @@ +--- +title: Enable NVIDIA CDI support for HAMi +sidebar_label: NVIDIA CDI support +--- + +This page explains what CDI does and which Helm chart values enable HAMi's NVIDIA CDI integration. + +## What is CDI? + +[CDI (Container Device Interface)](https://github.com/cncf-tags/container-device-interface) is a specification for describing the OCI configuration that a container needs to use a device. A CDI specification can define device nodes, mounts, environment variables, and OCI hooks in JSON or YAML. + +CDI is vendor-neutral. HAMi currently supports CDI device injection only for NVIDIA GPUs. The HAMi configuration and troubleshooting guidance on this page do not apply to devices from other vendors. + +CDI identifies each device by a fully qualified name: + +```text +vendor.com/class=device-name +``` + +When a container runtime receives this name, it finds the matching CDI specification in a directory such as `/etc/cdi` or `/var/run/cdi` and applies the specified changes to the container's OCI runtime specification. + +CDI does not schedule devices, allocate resources, or enforce quotas. In HAMi's NVIDIA CDI integration, the scheduler and NVIDIA Device Plugin continue to select and allocate GPUs; CDI injects the allocated GPUs and their runtime requirements into containers. + +## What problems does CDI solve? + +Injecting a complex device often requires more than mounting a single `/dev` node. A GPU container may also need driver libraries, several device nodes, environment variables, and lifecycle hooks. Without a common interface, vendors must adapt this logic to individual container runtimes, and runtimes may need vendor-specific implementations. + +CDI moves these OCI changes into a common specification. It provides: + +- A standard format for device nodes, mounts, environment variables, and hooks. +- Stable device names for passing allocation results from Device Plugins to container runtimes. +- Less dependency on runtime-specific device injection implementations. +- An inspectable record of the device configuration applied to a container. + +## How HAMi uses CDI for NVIDIA GPUs + +With `cdi-annotations` enabled for the NVIDIA Device Plugin, HAMi uses the following injection flow: + +1. The HAMi scheduler and NVIDIA Device Plugin select and allocate a GPU. +2. The Device Plugin generates a CDI specification and writes it to `/var/run/cdi` on the node. +3. The Device Plugin returns the allocated device name through a CDI annotation. +4. A CDI-aware container runtime resolves the annotation and applies the matching CDI specification to the container's OCI runtime specification. + +HAMi uses the following CDI kind for NVIDIA GPUs: + +```text +k8s.device-plugin.nvidia.com/gpu +``` + +## Prerequisites + +Before enabling NVIDIA CDI support in HAMi, confirm that: + +- The NVIDIA driver and NVIDIA Container Toolkit are installed on each GPU node. +- The container runtime supports CDI and has CDI enabled. +- The container runtime can read CDI specifications from `/var/run/cdi`. +- The actual NVIDIA driver root and `nvidia-ctk` path on each node are known. + +For runtime-specific instructions, see the [upstream CDI configuration guide](https://github.com/cncf-tags/container-device-interface#how-to-configure-cdi). + +## Configure the Helm chart + +Save the following values as `values-cdi.yaml`: + +```yaml +devicePlugin: + deviceListStrategy: "cdi-annotations" + nvidiaDriverRoot: "" + nvidiaHookPath: "" +``` + +- `devicePlugin.deviceListStrategy`: Set this to `cdi-annotations` so that the HAMi Device Plugin passes device names through CDI annotations. +- `devicePlugin.nvidiaDriverRoot`: The NVIDIA driver root on the node. HAMi uses this directory to discover driver files and device nodes. +- `devicePlugin.nvidiaHookPath`: The `nvidia-ctk` executable used on the node. HAMi writes this path into the generated CDI specification. + +Replace the two path placeholders according to the deployment: + +- Host-installed NVIDIA driver and Container Toolkit: use `/` for `` and `/usr/bin/nvidia-ctk` for ``. +- GPU Operator-managed driver and Container Toolkit: use `/run/nvidia/driver` for `` and `/usr/local/nvidia/toolkit/nvidia-ctk` for ``. + +:::warning + +The two path values must match the actual node filesystem. Do not copy values from a different deployment model. `nvidiaHookPath` must point to an executable that is available to the container runtime on the node. + +::: + +Install or upgrade HAMi: + +```bash +helm upgrade --install hami hami-charts/hami \ + --version 2.9.0 \ + --namespace kube-system \ + --create-namespace \ + --values values-cdi.yaml +``` + +## Verify the configuration + +First, confirm that Helm applied the expected values: + +```bash +helm get values hami -n kube-system +``` + +On a GPU node, check the selected `nvidia-ctk` executable and the CDI specification generated by HAMi. The following example uses the host-installed path; replace `NVIDIA_CTK_PATH` in GPU Operator environments: + +```bash +NVIDIA_CTK_PATH=/usr/bin/nvidia-ctk +sudo test -x "$NVIDIA_CTK_PATH" +"$NVIDIA_CTK_PATH" --version +sudo ls -l /var/run/cdi +sudo jq . /var/run/cdi/k8s.device-plugin.nvidia.com-gpu.json +``` + +Confirm that: + +- The CDI kind is `k8s.device-plugin.nvidia.com/gpu`. +- The allocated GPU UUID is present in the device list. +- Every device node and hook path in the specification exists on the node. + +Finally, check the Device Plugin logs: + +```bash +kubectl logs -n kube-system \ + -l app.kubernetes.io/component=hami-device-plugin \ + --tail=200 | grep -i cdi +``` + +The logs should not contain errors about CDI specification generation, unresolved devices, or missing hook paths. + +## Troubleshooting + +### `failed to stat CDI host device "/dev/nvidia-modeset": no such file or directory` + +`devicePlugin.nvidiaDriverRoot` does not match the driver layout on the node. Correct the value using the path rules above, upgrade HAMi, and inspect the regenerated CDI specification. + +### `stderr: No help topic for 'disable-device-node-modification'` + +The NVIDIA Container Toolkit executable referenced by the CDI specification does not support the `disable-device-node-modification` command. This command requires NVIDIA Container Toolkit 1.18.0 or later. + +Upgrade NVIDIA Container Toolkit and verify both the hook path recorded in the CDI specification and the version of the executable at that path. Host-installed and GPU Operator-managed Toolkit installations are independent; upgrading one does not update the other. + +[GPU Operator v25.10.1](https://github.com/NVIDIA/gpu-operator/blob/v25.10.1/deployments/gpu-operator/values.yaml#L220-L224) uses NVIDIA Container Toolkit v1.18.1. For related details, see [Toolkit version parser broken for casual kind demo user](https://github.com/kubernetes-sigs/dra-driver-nvidia-gpu/issues/568). + +### `error running createContainer hook #0: fork/exec /usr/bin/nvidia-ctk: no such file or directory` + +The hook path recorded in the CDI specification does not exist on the node. Locate the actual `nvidia-ctk` executable, correct `devicePlugin.nvidiaHookPath`, and upgrade HAMi to regenerate the CDI specification. diff --git a/versioned_docs/version-v2.9.0/installation/how-to-use-hami-dra.md b/versioned_docs/version-v2.9.0/installation/how-to-use-hami-dra.md index 2695115e..6db129c5 100644 --- a/versioned_docs/version-v2.9.0/installation/how-to-use-hami-dra.md +++ b/versioned_docs/version-v2.9.0/installation/how-to-use-hami-dra.md @@ -18,7 +18,7 @@ HAMi has provided support for K8s [DRA](https://kubernetes.io/docs/concepts/sche ## Prerequisites - Kubernetes version >= 1.34 with DRA Consumable Capacity [featuregate](https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates/) enabled -- [CDI](https://github.com/cncf-tags/container-device-interface?tab=readme-ov-file#how-to-configure-cdi) must be enabled in the underlying container runtime (such as containerd or CRI-O) +- [CDI](./configure-cdi.md) must be enabled in the underlying container runtime (such as containerd or CRI-O) - NVIDIA GPU Driver 440 or later ## Installation diff --git a/versioned_docs/version-v2.9.0/installation/prerequisites.md b/versioned_docs/version-v2.9.0/installation/prerequisites.md index 24966331..cc43b3e6 100644 --- a/versioned_docs/version-v2.9.0/installation/prerequisites.md +++ b/versioned_docs/version-v2.9.0/installation/prerequisites.md @@ -19,6 +19,8 @@ This guide assumes pre-installation of NVIDIA drivers and the `nvidia-container- For details see [Installing the NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html). +To inject NVIDIA GPUs through CDI, follow [Enable NVIDIA CDI support for HAMi](./configure-cdi.md) before installing HAMi. Verify the container runtime, driver root, and NVIDIA Container Toolkit path first. + ### Example for debian-based systems with Docker and containerd #### Install the `nvidia-container-toolkit` diff --git a/versioned_sidebars/version-v2.8.0-sidebars.json b/versioned_sidebars/version-v2.8.0-sidebars.json index f6860d80..fdae2454 100644 --- a/versioned_sidebars/version-v2.8.0-sidebars.json +++ b/versioned_sidebars/version-v2.8.0-sidebars.json @@ -31,6 +31,7 @@ "label": "Installation", "items": [ "installation/prerequisites", + "installation/configure-cdi", "installation/online-installation", "installation/offline-installation", "installation/upgrade", diff --git a/versioned_sidebars/version-v2.9.0-sidebars.json b/versioned_sidebars/version-v2.9.0-sidebars.json index 43434f82..6d0cd867 100644 --- a/versioned_sidebars/version-v2.9.0-sidebars.json +++ b/versioned_sidebars/version-v2.9.0-sidebars.json @@ -54,6 +54,7 @@ }, "items": [ "installation/prerequisites", + "installation/configure-cdi", "installation/online-installation", "installation/offline-installation", "installation/upgrade",