Models under a latency budget
On a production line the cycle time is fixed. The model either answers inside it or it does not ship, and no amount of accuracy on a validation set changes that.
Measure before you optimise
Most slow pipelines are slow in pre-processing, copies between host and device, or a Python loop around the model — not in the model itself. Profiling the whole path first is how you avoid spending a week quantising something that was never the bottleneck.
Every optimisation has a price
FP16, INT8, pruning, distillation, smaller input resolutions. Each one trades accuracy, engineering time or maintainability for latency. The skill is knowing which trade the use case can afford, and proving it on real plant data rather than a benchmark.
Batching versus latency
Throughput and latency pull in opposite directions. A camera triggering per part and a batch inspection of a full tray are different serving problems, and the right batch size is a property of the process, not of the GPU.
MLOps that survives the plant
Shipping a model where R&D can iterate daily, on infrastructure that is not allowed to be down. Versioned models, reproducible exports, and a rollback that takes seconds.