fix(macpro31): load the NVIDIA modules and guard the CDI generator
CI / flake (push) Skipped
CI / flake (pull_request) Successful in 4m12s

The CDI generator aborted with "failed to initialize NVML: Driver Not
Loaded", taking docker.service with it (requiredBy) and failing the
switch.

Two causes. The nixpkgs NVIDIA module only adds nvidia/nvidia_modeset/
nvidia_drm to boot.kernelModules when services.xserver.enable is set,
which is false on this Wayland-only host, so load them explicitly.
nvidia_uvm stays out: the module's modprobe softdep loads it once the GPU
device exists.

The generator also runs during activation, when a module rebuilt against a
new kernel cannot be loaded until reboot -- a guaranteed failure after
every kernel bump. Guard it with ConditionPathExists on
/proc/driver/nvidia/version so it skips rather than fails; the toolkit's
udev rule restarts it when the device appears, so the specs are generated
on the next boot.
This commit is contained in:
Emma Thorpe
2026-08-17 20:47:33 +01:00
parent 0f7fb7f78a
commit d4e7475db9
2 changed files with 49 additions and 0 deletions
+26
View File
@@ -82,6 +82,32 @@ docker run --rm --device=nvidia.com/gpu=all nvidia/cuda:12.9.1-base-ubuntu24.04
- Docker socket is local-only (no TCP listener, unlike the Pi). Users need the
`docker` group; the registry already grants it.
### "Driver Not Loaded" from the CDI generator
`nvidia-container-toolkit-cdi-generator.service` fails with
`failed to initialize NVML: Driver Not Loaded` whenever the `nvidia` kernel
module is not loaded in the **running** kernel. After a kernel bump that is
unavoidable — the rebuilt module cannot load until reboot — so the unit is
guarded with `ConditionPathExists=/proc/driver/nvidia/version` and skips
instead of failing. Without that guard it also takes `docker.service`
(`requiredBy`) with it and makes `nixos-rebuild switch` exit non-zero.
**Reboot after a rebuild that touches the driver or the kernel.** The toolkit's
udev rule restarts the generator when the GPU device appears, so the CDI specs
are written on the next boot. To check the state:
```sh
lsmod | grep nvidia # nvidia, nvidia_modeset, nvidia_drm, nvidia_uvm
cat /proc/driver/nvidia/version
nvidia-smi
systemctl status nvidia-container-toolkit-cdi-generator.service
ls /var/run/cdi # the generated spec
```
If the module is genuinely absent after a reboot, check `dmesg | grep -i
nvidia` (build/version mismatch, or nouveau still bound — the module blacklists
it, so that should not happen).
## Claude Code — not installed here
The dual Harpertown Xeons are **x86-64-v1** (SSE4.1, but no SSE4.2/POPCNT) and