Observability
A distributed application watching a critical manufacturing process across a hundred hosts generates enormous amounts of telemetry and, usually, very little understanding.
Which metrics earn their storage
Most metrics are collected because collecting them was easy. Cardinality is a budget, and spending it without a decision in mind is how observability bills get out of control.
Tagsets across independent modules
The moment more than one team emits telemetry, labels drift. Enforcement has to happen at ingestion or at CI — asking politely does not work at scale.
The missing spans
Industrial systems trace almost nothing. A camera triggers, inference runs, a result reaches a display, and the only evidence is three unrelated metrics. End-to-end latency is a span problem, and treating it as a metrics problem is why bottlenecks stay invisible.
Instrumenting the whole chain
What it took to instrument camera to inference to display, and what surfaced immediately once the chain was visible.