Topics

Observability

Industrial systems are all metrics and no traces. Cost-controlled OpenTelemetry, tagset governance across independent modules, and knowing which signals earn their storage.
Draft
This page is seeded structure, not finished writing. The author has not rewritten it yet.

A distributed application watching a critical manufacturing process across a hundred hosts generates enormous amounts of telemetry and, usually, very little understanding.

Which metrics earn their storage

Most metrics are collected because collecting them was easy. Cardinality is a budget, and spending it without a decision in mind is how observability bills get out of control.

Tagsets across independent modules

The moment more than one team emits telemetry, labels drift. Enforcement has to happen at ingestion or at CI — asking politely does not work at scale.

The missing spans

Industrial systems trace almost nothing. A camera triggers, inference runs, a result reaches a display, and the only evidence is three unrelated metrics. End-to-end latency is a span problem, and treating it as a metrics problem is why bottlenecks stay invisible.

Instrumenting the whole chain

What it took to instrument camera to inference to display, and what surfaced immediately once the chain was visible.

Related work

OpAMP
Open Agent Management Protocol — remote configuration, health reporting and upgrade of telemetry agents at fleet scale.
OpenTelemetry
The vendor-neutral standard for traces, metrics and logs.
© 2026 Mathieu Sabatier