feat(macpro31): NVIDIA P400 with CUDA Docker, and a fleet-wide CPU capability gate #94

Merged
lyrathorpe merged 3 commits from feat/macpro31-nvidia-cuda into main 2026-08-17 20:59:21 +01:00
3 Commits
Author SHA1 Message Date
Emma Thorpe d4e7475db9 fix(macpro31): load the NVIDIA modules and guard the CDI generator
CI / flake (push) Skipped
CI / flake (pull_request) Successful in 4m12s
The CDI generator aborted with "failed to initialize NVML: Driver Not
Loaded", taking docker.service with it (requiredBy) and failing the
switch.

Two causes. The nixpkgs NVIDIA module only adds nvidia/nvidia_modeset/
nvidia_drm to boot.kernelModules when services.xserver.enable is set,
which is false on this Wayland-only host, so load them explicitly.
nvidia_uvm stays out: the module's modprobe softdep loads it once the GPU
device exists.

The generator also runs during activation, when a module rebuilt against a
new kernel cannot be loaded until reboot -- a guaranteed failure after
every kernel bump. Guard it with ConditionPathExists on
/proc/driver/nvidia/version so it skips rather than fails; the toolkit's
udev rule restarts it when the device appears, so the specs are generated
on the next boot.
2026-08-17 20:47:33 +01:00
Emma Thorpe 0f7fb7f78a feat(macpro31): NVIDIA Quadro P400 driver and CUDA-enabled Docker
CI / flake (push) Skipped
CI / flake (pull_request) Successful in 4m11s
The stock GPU has been replaced with a Quadro P400 (Pascal, GP108). Add
hosts/MacPro31/nvidia.nix:

- Driver branch 580 (nvidiaPackages.legacy_580), not the nixpkgs default
  production branch (595.x). 580 is the last branch supporting
  Maxwell/Pascal/Volta and is an LTS branch until Aug 2028; a newer one
  does not drive this card.
- modesetting.enable for Wayland (nvidia-drm.modeset=1), open = false
  (the open kernel modules need Turing or later), and sway
  --unsupported-gpu, which wlroots requires with the proprietary driver.
- Docker with GPU access via CDI (hardware.nvidia-container-toolkit),
  rather than the deprecated virtualisation.docker.enableNvidia runtime
  wrapper. Containers run with --device=nvidia.com/gpu=all and must ship
  a CUDA 12.x or older runtime: CUDA 13 dropped sm_61.

The driver packages are unfree, so allowlist them in unfreePackages; they
are not cached and the kernel module builds on the host.

Also declare features.cpu.microarchLevel = 1 for this machine: the
Harpertown Xeons have SSE4.1 but no SSE4.2/POPCNT, which switches off
Claude Code through the fleet-wide gate.
2026-08-17 20:35:39 +01:00
Emma Thorpe 0d13581896 feat(features): gate Claude Code on the host CPU microarchitecture level
Claude Code runs on Node, whose V8 build requires SSE4.2 and POPCNT
(x86-64-v2). On an older x86_64 CPU it does not run, so it must not be
installed there in the first place.

Nix cannot detect the CPU (pure evaluation, hosts often built elsewhere),
so add features.cpu.microarchLevel: the psABI level a host declares about
itself, defaulting to 2. features.claudeCode.enable derives from it, and
home/claude.nix reads that through home-manager's osConfig and installs
nothing -- CLI, CLAUDE.md, output style or memory symlink -- when it is
off. Hosts without the option (Darwin, the standalone homeConfigurations)
keep the tool enabled.

An assertion fails evaluation if a host force-enables the flag below the
required level, so the mistake surfaces in nix flake check rather than as
an illegal-instruction crash on the machine.
2026-08-17 20:35:29 +01:00