vLLM metrics into VCF Operations, part 1: don't ship the histogram
vLLM exposes ~200 Prometheus series per model. Pushing them raw into VCF Operations is a cardinality bomb. The design of a pipeline that interpolates percentiles client-side, computes live rates statefully, derives a saturation score and drops what can’t be graphed.