Topics

Models under a latency budget

A model that is accurate in a notebook and too slow on the line is not a model, it is a demo. ML engineering for performance — and knowing which optimisation is worth what it costs.
Draft
This page is seeded structure, not finished writing. The author has not rewritten it yet.

On a production line the cycle time is fixed. The model either answers inside it or it does not ship, and no amount of accuracy on a validation set changes that.

Measure before you optimise

Most slow pipelines are slow in pre-processing, copies between host and device, or a Python loop around the model — not in the model itself. Profiling the whole path first is how you avoid spending a week quantising something that was never the bottleneck.

Every optimisation has a price

FP16, INT8, pruning, distillation, smaller input resolutions. Each one trades accuracy, engineering time or maintainability for latency. The skill is knowing which trade the use case can afford, and proving it on real plant data rather than a benchmark.

Batching versus latency

Throughput and latency pull in opposite directions. A camera triggering per part and a batch inspection of a full tray are different serving problems, and the right batch size is a property of the process, not of the GPU.

MLOps that survives the plant

Shipping a model where R&D can iterate daily, on infrastructure that is not allowed to be down. Versioned models, reproducible exports, and a rollback that takes seconds.

© 2026 Mathieu Sabatier