Topics

GPUs on the line

Accelerated inference in production is mostly a driver problem. CUDA versions, TensorRT engines, Triton serving and GPU sharing on edge nodes that nobody is allowed to reboot.
Draft
This page is seeded structure, not finished writing. The author has not rewritten it yet.

Getting a model to run on a GPU takes an afternoon. Keeping it running on dozens of edge nodes, through driver updates and hardware refreshes, is the actual job.

The driver is part of the application

Driver, CUDA, cuDNN and TensorRT versions form one compatibility matrix, and an engine built against one combination does not load on another. Treating the host driver as someone else's problem is how an OS patch takes a line down.

Knowing when to reach for the NVIDIA stack

TensorRT, DeepStream and Triton are the right answer for some workloads and a heavy dependency for others. ONNX Runtime or a plain CPU path is sometimes the better call — and the decision should be made on measured latency, not on which SDK is fashionable.

Serving on Kubernetes at the edge

The GPU operator, device plugins, time-slicing and MIG: how several models share one card on a k3s node without starving each other, and how that gets rolled out through GitOps like everything else.

Engines are build artefacts

A TensorRT engine is tied to the GPU it was built on. Building, caching and shipping engines per hardware target belongs in the pipeline, not in a startup script that silently takes ten minutes on first boot.

Related work

k3s
A certified Kubernetes distribution small enough to run in a plant cabinet.
© 2026 Mathieu Sabatier