← Lab notes

Lab notes #4: Four things that broke a two-node GPU Kubernetes cluster before vLLM would serve

Validating a private-inference stack — k3s, NVIDIA GPU Operator, vLLM, DCGM — on two cloud A10s took five attempts. A CUDA 13 image on a 570 driver, broken -cu129 image variants, a CDI hook missing from the host toolkit, and a Service named vllm. Every failure was a version or naming mismatch between open-source parts, and most left no logs.

Many of the companies we talk to want to run inference on their own GPUs, inside their own network. The usual plan is to take well-known open-source parts and put them together: Kubernetes, NVIDIA's GPU Operator, vLLM, DCGM metrics.

We wanted to know how much can go wrong with that plan, so we built the smallest realistic version on two rented GPUs.

The good news: the final stack passed every check in ten minutes and cost $0.45.

The real lesson: it took five attempts to get there. None of the four problems was a bug in our code or a hardware fault. Each was a version or naming mismatch between correct, popular components, and two of them failed without a single log line.

The setup

Machines 2 × Lambda Cloud gpu_1x_a10 (24 GB), us-east-1, private network between them
Host Lambda's Ubuntu image: NVIDIA driver 570.148.08, NVIDIA Container Toolkit 1.17.8 preinstalled
Cluster k3s v1.36 (1 server + 1 agent)
GPU stack NVIDIA GPU Operator via Helm, using the preinstalled driver and toolkit
Inference vLLM, 2 replicas with 1 GPU each, behind a Service; model Qwen2.5-3B-Instruct
Checks request spread across replicas, DCGM metrics from every node, TokenWatt probe on every node, drain/uncordon drill

This validates management, not performance. vLLM ran with --enforce-eager, for reasons explained below.

The run that passed

Step Time Result
GPUs visible, private network 3 s both nodes
k3s server + agent Ready 20 s ✔
GPU Operator: 1 allocatable GPU per node 31 s ✔
vLLM Deployment, 2/2 ready on 2 nodes 162 s ✔ (including a 9 GB image pull)
Service spreads requests 50 s 40/40 answered, split 16 / 24
DCGM metrics (utilization, power, memory) 3 s every node
TokenWatt probe 14 s every node
Drain node 2 → service stays up → uncordon 97 s 10/10 answered during drain; replica back after

Five attempts plus one single-machine debugging session cost $2.73 in total. Every instance was terminated automatically.

What broke, in order

1. A CUDA 13 image on a CUDA 12.8 driver

vLLM's default image is now built for CUDA 13, and the hosts we got had driver 570, which supports CUDA 12.8. The container still starts, because CUDA's forward-compatibility libraries take over. But CUDA-graph capture fails in that mode with operation not permitted. We'd hit the same thing on an H100 in lab notes #2.

Eager mode works. For benchmarks, where CUDA graphs matter, we skip hosts with drivers older than 580 instead.

2. The CUDA 12.9 images were broken

The obvious fix is vLLM's CUDA 12.9 build. We tried two:

  • v0.31.0-cu129 ships a torchcodec built against CUDA 13: libtorchcodec_image.so can't find libcudart.so.13. vLLM imports torchcodec at startup and exits.
  • v0.30.0-x86_64-cu129 fails with operator torchvision::nms does not exist, a torchvision built for a different torch.

We didn't find a working -cu129 image. We reported the torchcodec mismatch as vllm#60406; the torchvision one was already reported in vllm#59057.

3. The GPU Operator wrote a hook the host couldn't run — and nothing logged it

With the default image and eager mode, the pods still never started. They showed CrashLoopBackOff and RunContainerError, and kubectl logs --previous was empty. Running the same image with plain Docker on the same machine worked.

The difference only showed up in the pod events:

error running createRuntime hook #0: exit status 3,
stderr: No help topic for 'apply-cuda-memory-limits'

The GPU Operator's device plugin writes CDI specs that call nvidia-ctk hook apply-cuda-memory-limits. That subcommand exists in newer NVIDIA Container Toolkit releases, but the host had 1.17.8, so every GPU container died before its process started. Docker uses a different injection path and never reads those specs.

The fix is to install the operator with cdi.enabled=false, or upgrade the host toolkit to match it.

4. A Service named vllm crashed vLLM

Now the containers started, loaded the model, and crashed during engine startup. The cause was the Service name. Kubernetes injects environment variables for every Service in the namespace. A Service called vllm produces VLLM_PORT=tcp://10.43.x.x:8000, and VLLM_PORT is also one of vLLM's own settings, which it expects to be an integer.

The fix is enableServiceLinks: false on the pod, plus a different Service name. After that, everything passed.

What this means for private inference

  • Each component was fine on its own. The driver, toolkit, GPU Operator, vLLM image and Kubernetes all work as documented. The failures came from how they were combined: versions that drifted apart and names that collided.
  • The worst failures were silent. Two of the four left nothing in the container logs. You find them by reading pod events, or by running the same image outside Kubernetes and comparing.
  • The fixes don't stay fixed. The next GPU Operator release, the next vLLM image or the next host image can break this again. A stack like this needs automated acceptance checks after every upgrade — the same checks we ran here.

That's the case for treating acceptance and operations as part of the platform, not an afterthought. It's also why we keep these checks as code.

Reproduce

cd bench
scripts/lambda/k8s_validate.py --type gpu_1x_a10 --max-minutes 75 --yes

The script launches both instances, runs every step, writes a report with timings, and terminates the instances whatever happens.

Next: the same validation with Slurm, then on AMD Instinct with the AMD GPU Operator.