[{"data":1,"prerenderedAt":77},["ShallowReactive",2],{"navigation":3,"topic-ml":14,"topic-ml-ventures":74,"topic-ml-oss":75,"topic-ml-posts":76},[4],{"title":5,"path":6,"stem":7,"children":8,"page":13},"Blog","\u002Fblog","blog",[9],{"title":10,"path":11,"stem":12},"Kubernetes 101","\u002Fblog\u002Fkubernetes-introduction","blog\u002Fkubernetes-introduction",false,{"id":15,"title":16,"body":17,"description":61,"draftNotice":62,"extension":63,"meta":64,"navigation":62,"order":65,"path":66,"pillar":67,"related":68,"seo":71,"stem":72,"__hash__":73},"topics\u002Ftopics\u002Fml.md","Models under a latency budget",{"type":18,"value":19,"toc":53},"minimark",[20,24,29,32,36,39,43,46,50],[21,22,23],"p",{},"On a production line the cycle time is fixed. The model either answers inside it or it does not ship, and no amount of accuracy on a validation set changes that.",[25,26,28],"h2",{"id":27},"measure-before-you-optimise","Measure before you optimise",[21,30,31],{},"Most slow pipelines are slow in pre-processing, copies between host and device, or a Python loop around the model — not in the model itself. Profiling the whole path first is how you avoid spending a week quantising something that was never the bottleneck.",[25,33,35],{"id":34},"every-optimisation-has-a-price","Every optimisation has a price",[21,37,38],{},"FP16, INT8, pruning, distillation, smaller input resolutions. Each one trades accuracy, engineering time or maintainability for latency. The skill is knowing which trade the use case can afford, and proving it on real plant data rather than a benchmark.",[25,40,42],{"id":41},"batching-versus-latency","Batching versus latency",[21,44,45],{},"Throughput and latency pull in opposite directions. A camera triggering per part and a batch inspection of a full tray are different serving problems, and the right batch size is a property of the process, not of the GPU.",[25,47,49],{"id":48},"mlops-that-survives-the-plant","MLOps that survives the plant",[21,51,52],{},"Shipping a model where R&D can iterate daily, on infrastructure that is not allowed to be down. Versioned models, reproducible exports, and a rollback that takes seconds.",{"title":54,"searchDepth":55,"depth":55,"links":56},"",2,[57,58,59,60],{"id":27,"depth":55,"text":28},{"id":34,"depth":55,"text":35},{"id":41,"depth":55,"text":42},{"id":48,"depth":55,"text":49},"A model that is accurate in a notebook and too slow on the line is not a model, it is a demo. ML engineering for performance — and knowing which optimisation is worth what it costs.",true,"md",{},5,"\u002Ftopics\u002Fml","ml",{"ventures":69,"oss":70},[],[],{"title":16,"description":61},"topics\u002Fml","ukInvv4VI67lz4vDb1AahLVUzt6Qs82UZjbJvqUR8eg",[],[],[],1791414106637]