GPUs on the line
Getting a model to run on a GPU takes an afternoon. Keeping it running on dozens of edge nodes, through driver updates and hardware refreshes, is the actual job.
The driver is part of the application
Driver, CUDA, cuDNN and TensorRT versions form one compatibility matrix, and an engine built against one combination does not load on another. Treating the host driver as someone else's problem is how an OS patch takes a line down.
Knowing when to reach for the NVIDIA stack
TensorRT, DeepStream and Triton are the right answer for some workloads and a heavy dependency for others. ONNX Runtime or a plain CPU path is sometimes the better call — and the decision should be made on measured latency, not on which SDK is fashionable.
Serving on Kubernetes at the edge
The GPU operator, device plugins, time-slicing and MIG: how several models share one card on a k3s node without starving each other, and how that gets rolled out through GitOps like everything else.
Engines are build artefacts
A TensorRT engine is tied to the GPU it was built on. Building, caching and shipping engines per hardware target belongs in the pipeline, not in a startup script that silently takes ten minutes on first boot.