[{"data":1,"prerenderedAt":89},["ShallowReactive",2],{"navigation":3,"topic-inference":14,"topic-inference-ventures":74,"topic-inference-oss":75,"topic-inference-posts":88},[4],{"title":5,"path":6,"stem":7,"children":8,"page":13},"Blog","\u002Fblog","blog",[9],{"title":10,"path":11,"stem":12},"Kubernetes 101","\u002Fblog\u002Fkubernetes-introduction","blog\u002Fkubernetes-introduction",false,{"id":15,"title":16,"body":17,"description":61,"draftNotice":62,"extension":63,"meta":64,"navigation":62,"order":65,"path":66,"pillar":67,"related":68,"seo":71,"stem":72,"__hash__":73},"topics\u002Ftopics\u002Finference.md","GPUs on the line",{"type":18,"value":19,"toc":53},"minimark",[20,24,29,32,36,39,43,46,50],[21,22,23],"p",{},"Getting a model to run on a GPU takes an afternoon. Keeping it running on dozens of edge nodes, through driver updates and hardware refreshes, is the actual job.",[25,26,28],"h2",{"id":27},"the-driver-is-part-of-the-application","The driver is part of the application",[21,30,31],{},"Driver, CUDA, cuDNN and TensorRT versions form one compatibility matrix, and an engine built against one combination does not load on another. Treating the host driver as someone else's problem is how an OS patch takes a line down.",[25,33,35],{"id":34},"knowing-when-to-reach-for-the-nvidia-stack","Knowing when to reach for the NVIDIA stack",[21,37,38],{},"TensorRT, DeepStream and Triton are the right answer for some workloads and a heavy dependency for others. ONNX Runtime or a plain CPU path is sometimes the better call — and the decision should be made on measured latency, not on which SDK is fashionable.",[25,40,42],{"id":41},"serving-on-kubernetes-at-the-edge","Serving on Kubernetes at the edge",[21,44,45],{},"The GPU operator, device plugins, time-slicing and MIG: how several models share one card on a k3s node without starving each other, and how that gets rolled out through GitOps like everything else.",[25,47,49],{"id":48},"engines-are-build-artefacts","Engines are build artefacts",[21,51,52],{},"A TensorRT engine is tied to the GPU it was built on. Building, caching and shipping engines per hardware target belongs in the pipeline, not in a startup script that silently takes ten minutes on first boot.",{"title":54,"searchDepth":55,"depth":55,"links":56},"",2,[57,58,59,60],{"id":27,"depth":55,"text":28},{"id":34,"depth":55,"text":35},{"id":41,"depth":55,"text":42},{"id":48,"depth":55,"text":49},"Accelerated inference in production is mostly a driver problem. CUDA versions, TensorRT engines, Triton serving and GPU sharing on edge nodes that nobody is allowed to reboot.",true,"md",{},6,"\u002Ftopics\u002Finference","inference",{"oss":69},[70],"k3s",{"title":16,"description":61},"topics\u002Finference","qXpEvPvjEaeEUYaL3gBH9Th3s3ZM9QDCVJ0QVNxT0vs",[],[76],{"id":77,"description":78,"extension":79,"featured":13,"language":80,"meta":81,"name":70,"order":82,"repo":83,"role":84,"slug":70,"stem":85,"why":86,"__hash__":87},"oss\u002Fwork\u002Foss\u002Fk3s.yml","A certified Kubernetes distribution small enough to run in a plant cabinet.","yml","Go",{},5,"https:\u002F\u002Fgithub.com\u002Fk3s-io\u002Fk3s","watching","work\u002Foss\u002Fk3s","It is the reason \"if it needs a datacenter it does not belong at the edge\" is a design constraint rather than a complaint.","VFzFEQeRQHsxeOcQknHeoEoI6ks_dqacKmCHVZfZj7c",[],1791414106762]