CI / flake (push) Skipped
CI / flake (pull_request) Successful in 4m12s
The CDI generator aborted with "failed to initialize NVML: Driver Not Loaded", taking docker.service with it (requiredBy) and failing the switch. Two causes. The nixpkgs NVIDIA module only adds nvidia/nvidia_modeset/ nvidia_drm to boot.kernelModules when services.xserver.enable is set, which is false on this Wayland-only host, so load them explicitly. nvidia_uvm stays out: the module's modprobe softdep loads it once the GPU device exists. The generator also runs during activation, when a module rebuilt against a new kernel cannot be loaded until reboot -- a guaranteed failure after every kernel bump. Guard it with ConditionPathExists on /proc/driver/nvidia/version so it skips rather than fails; the toolkit's udev rule restarts it when the device appears, so the specs are generated on the next boot.
67 lines
3.2 KiB
Nix
67 lines
3.2 KiB
Nix
# NVIDIA Quadro P400 (Pascal, GP108) on the Mac Pro 3,1: proprietary driver for
|
|
# the Sway desktop, plus Docker with GPU/CUDA access for containers.
|
|
#
|
|
# Driver branch: 580 (nvidiaPackages.legacy_580), NOT the nixpkgs default
|
|
# (`production`, currently 595.x). 580 is the last branch that supports
|
|
# Maxwell/Pascal/Volta -- NVIDIA keeps it as an LTS branch to Aug 2028 -- and a
|
|
# newer branch simply will not drive this card.
|
|
#
|
|
# The driver is unfree, so it is not in the binary cache: the kernel module is
|
|
# compiled locally. On this machine's 2008 Xeons expect the first rebuild after
|
|
# a kernel bump to take a long while.
|
|
{ config, ... }:
|
|
|
|
{
|
|
# Selects the proprietary driver; the module blacklists nouveau/nvidiafb and
|
|
# loads nvidia-uvm (needed by CUDA) via modprobe softdep. Naming is historical
|
|
# -- this option drives the kernel/driver choice on Wayland hosts too, which
|
|
# is why it is set on a machine that runs no X server.
|
|
services.xserver.videoDrivers = [ "nvidia" ];
|
|
|
|
hardware.nvidia = {
|
|
package = config.boot.kernelPackages.nvidiaPackages.legacy_580;
|
|
# Required for Wayland: sets nvidia-drm.modeset=1 (and fbdev=1), without
|
|
# which wlroots gets no GBM device and Sway/cage fail to start.
|
|
modesetting.enable = true;
|
|
# The open kernel modules need Turing or later; Pascal must use the closed
|
|
# ones. Explicit because the option has no default on driver >= 560.
|
|
open = false;
|
|
};
|
|
|
|
# The NVIDIA module only puts these in boot.kernelModules when
|
|
# services.xserver.enable is true, which is false on this Wayland-only host --
|
|
# so load them explicitly rather than relying on udev modalias autoloading.
|
|
# nvidia_uvm (needed by CUDA) is deliberately absent: the module's modprobe
|
|
# softdep pulls it in after the GPU device exists, which is the supported
|
|
# ordering.
|
|
boot.kernelModules = [
|
|
"nvidia"
|
|
"nvidia_modeset"
|
|
"nvidia_drm"
|
|
];
|
|
|
|
# wlroots refuses the proprietary NVIDIA driver unless told to proceed. The
|
|
# greeter's compositor (cage) has no such check; only Sway needs the flag,
|
|
# which the module bakes into the wrapper the session's .desktop file runs.
|
|
programs.sway.extraOptions = [ "--unsupported-gpu" ];
|
|
|
|
virtualisation.docker.enable = true;
|
|
|
|
# CDI-based GPU access for containers: generates /var/run/cdi specs from the
|
|
# host driver at boot and turns on Docker's CDI feature. Run GPU workloads
|
|
# with `docker run --device=nvidia.com/gpu=all ...`. The deprecated
|
|
# virtualisation.docker.enableNvidia runtime wrapper is deliberately not used.
|
|
hardware.nvidia-container-toolkit.enable = true;
|
|
|
|
# The generator needs a loaded kernel module: without one it aborts with
|
|
# "failed to initialize NVML: Driver Not Loaded". That is guaranteed after a
|
|
# kernel bump, where the rebuilt module cannot load until reboot -- and since
|
|
# the unit is requiredBy docker.service and wantedBy multi-user.target, the
|
|
# failure takes Docker down and makes `nixos-rebuild switch` exit non-zero.
|
|
# Skip the run instead when no driver is loaded; the toolkit's udev rule
|
|
# restarts the unit as soon as the nvidia device appears, so the CDI specs are
|
|
# still generated on the next boot.
|
|
systemd.services.nvidia-container-toolkit-cdi-generator.unitConfig.ConditionPathExists =
|
|
"/proc/driver/nvidia/version";
|
|
}
|