<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Prometheus on The Nested Lab</title>
    <link>https://thenestedlab.com/tags/prometheus/</link>
    <description>Recent content in Prometheus on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/tags/prometheus/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>vLLM metrics into VCF Operations, part 1: don&#39;t ship the histogram</title>
      <link>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-1/</link>
      <pubDate>Wed, 30 Sep 2026 00:00:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-1/</guid>
      <description>vLLM exposes ~200 Prometheus series per model. Pushing them raw into VCF Operations is a cardinality bomb. The design of a pipeline that interpolates percentiles client-side, computes live rates statefully, derives a saturation score and drops what can&amp;rsquo;t be graphed.</description>
      <content:encoded><![CDATA[<p>vLLM&rsquo;s <code>/metrics</code> endpoint is generous. Every model instance exposes a few
hundred Prometheus series: counters, gauges, and — the problem — histograms
with dozens of buckets each. Point a naïve scraper at it and push
everything into VCF Operations and you get two things: a TSDB full of
<code>_bucket{le=&quot;0.05&quot;}</code> series nobody will ever graph, and a dashboard that
tells an operator nothing.</p>
<p>This two-parter is the pipeline that fixed that. Part 1 is the design:
what to compute, what to drop, and why. <a href="/series/llm-ops-on-vcf/">Part 2</a>
is the operations story — scheduling, self-monitoring, alerts.</p>
<h2 id="the-core-decision-percentiles-are-computed-here-not-there">The core decision: percentiles are computed here, not there</h2>
<p>A Prometheus histogram is a set of cumulative bucket counters:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.1&#34;}   1203
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.25&#34;}  4871
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.5&#34;}   7120
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;+Inf&#34;}  7412
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_sum                1812.4
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_count              7412
</span></span></code></pre></div><p>Prometheus turns that into a P95 at query time with <code>histogram_quantile</code>.
VCF Operations has no such function — it stores gauges. So the pipeline
does the interpolation <em>before</em> pushing:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="kd">function</span><span class="w"> </span><span class="nb">Get-Percentile</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="k">param</span><span class="p">([</span><span class="no">array</span><span class="p">]</span><span class="nv">$Buckets</span><span class="p">,</span> <span class="p">[</span><span class="no">double</span><span class="p">]</span><span class="nv">$Total</span><span class="p">,</span> <span class="p">[</span><span class="no">double</span><span class="p">]</span><span class="nv">$Percentile</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="nv">$target</span> <span class="p">=</span> <span class="nv">$Percentile</span> <span class="p">*</span> <span class="nv">$Total</span>
</span></span><span class="line"><span class="cl">    <span class="nv">$prevLe</span> <span class="p">=</span> <span class="mf">0.0</span><span class="p">;</span> <span class="nv">$prevCount</span> <span class="p">=</span> <span class="mf">0.0</span>
</span></span><span class="line"><span class="cl">    <span class="k">foreach</span> <span class="p">(</span><span class="nv">$b</span> <span class="k">in</span> <span class="nv">$Buckets</span><span class="p">)</span> <span class="p">{</span>                    <span class="c"># sorted by le, +Inf last</span>
</span></span><span class="line"><span class="cl">        <span class="k">if</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Count</span> <span class="o">-ge</span> <span class="nv">$target</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">            <span class="k">if</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Le</span> <span class="o">-eq</span> <span class="p">[</span><span class="no">double</span><span class="p">]::</span><span class="n">PositiveInfinity</span><span class="p">)</span> <span class="p">{</span> <span class="k">return</span> <span class="nv">$prevLe</span> <span class="p">}</span>
</span></span><span class="line"><span class="cl">            <span class="nv">$frac</span> <span class="p">=</span> <span class="p">(</span><span class="nv">$target</span> <span class="p">-</span> <span class="nv">$prevCount</span><span class="p">)</span> <span class="p">/</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Count</span> <span class="p">-</span> <span class="nv">$prevCount</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">            <span class="k">return</span> <span class="nv">$prevLe</span> <span class="p">+</span> <span class="nv">$frac</span> <span class="p">*</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Le</span> <span class="p">-</span> <span class="nv">$prevLe</span><span class="p">)</span>   <span class="c"># linear within the bucket</span>
</span></span><span class="line"><span class="cl">        <span class="p">}</span>
</span></span><span class="line"><span class="cl">        <span class="nv">$prevLe</span> <span class="p">=</span> <span class="nv">$b</span><span class="p">.</span><span class="n">Le</span><span class="p">;</span> <span class="nv">$prevCount</span> <span class="p">=</span> <span class="nv">$b</span><span class="p">.</span><span class="py">Count</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>Linear interpolation within the bucket that crosses the target rank —
the same approximation Prometheus makes. Out come <strong>four flat gauges per
histogram</strong>: <code>p50_ms</code>, <code>p95_ms</code>, <code>p99_ms</code>, <code>avg_ms</code> (from <code>_sum/_count</code>).
Twenty-odd series become four, and they&rsquo;re the four an operator reads.</p>
<p>Applied to: TTFT, end-to-end latency, inter-token latency, time per output
token, prefill time, decode time, inference time, queue time.</p>
<h2 id="rates-need-memory">Rates need memory</h2>
<p>Counters (<code>generation_tokens_total</code>, <code>request_success_total</code>) are useless
as absolute values. What you want is tokens <em>per second</em> — which needs the
previous sample. So the script is stateful: <code>llm-metrics-state.json</code> holds
the last counters and timestamp per target, and each run computes deltas:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">tokens_per_sec = (tokens_now - tokens_prev) / (t_now - t_prev)
</span></span></code></pre></div><p>Same trick, one step further, for <strong>live latency</strong>: <code>_sum</code> and <code>_count</code>
are both counters, so <code>Δsum / Δcount</code> is the <em>mean over the last interval</em>
— not the all-time mean the histogram gives you. That&rsquo;s how you get
<code>live_avg_ttft_ms</code> that reflects the last 60 seconds instead of the last
fortnight.</p>
<p>Two guards make this safe:</p>
<ul>
<li><strong>Stale-state protection.</strong> If the gap since the last run exceeds 300 s
(scheduler stopped, server rebooted), the baseline is dropped rather
than producing a diluted &ldquo;per-second&rdquo; rate averaged over an hour.</li>
<li><strong>Restart detection.</strong> A counter that went <em>down</em> means vLLM restarted;
the delta is discarded for that cycle.</li>
</ul>
<h2 id="derived-metrics-what-the-raw-numbers-wont-tell-you">Derived metrics: what the raw numbers won&rsquo;t tell you</h2>
<p>Two computed values earn their place on the top of the dashboard.</p>
<p><strong>Queue pressure ratio</strong> — <code>(waiting + swapped) / (running + 1)</code>. Above
1.0 means more requests are waiting than being served: scale out.</p>
<p><strong>System saturation score</strong> (0–100):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="nv">$saturation</span> <span class="p">=</span> <span class="p">(</span><span class="nv">$kvCachePct</span> <span class="p">*</span> <span class="mf">0.5</span><span class="p">)</span> <span class="p">+</span> <span class="p">(</span><span class="nv">$queuePressure</span> <span class="p">*</span> <span class="mf">25.0</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="k">if</span> <span class="p">(</span><span class="nv">$saturation</span> <span class="o">-gt</span> <span class="mf">100</span><span class="p">)</span> <span class="p">{</span> <span class="nv">$saturation</span> <span class="p">=</span> <span class="mf">100</span> <span class="p">}</span>
</span></span></code></pre></div><p>Full KV cache alone scores 50; queue pressure of 2.0 scores the other 50.
It&rsquo;s a heuristic, and it&rsquo;s deliberately one number: a dial that goes red
when the engine is about to start swapping requests to CPU memory, which
is the moment latency falls off a cliff. Warning at 75, critical at 90.</p>
<p><img alt="vllm|perf|system_saturation_score over the last hour in VCF Operations" loading="lazy" src="/images/ui/o6-ops-vllm-saturation-score.jpg">
<em>The dial, as Ops draws it: one gauge climbing towards the warning line as KV-cache use and queue pressure rise together. (Test-mode data — see part 2.)</em></p>
<p>Also derived: prefix-cache hit rate (live and all-time), average request
size from the dropped <code>http_request_size_bytes</code> <code>_sum</code>, uptime in days,
and <code>is_up = 1</code> on every successful scrape — so <em>absence</em> of the metric is
the alert.</p>
<h2 id="what-gets-dropped-and-why">What gets dropped, and why</h2>
<p>The <code>$DropPrefixes</code> list is as important as anything computed:</p>
<table>
	<thead>
			<tr>
					<th>Dropped</th>
					<th>Why</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td><code>python_gc_*</code>, <code>python_info</code></td>
					<td>runtime noise; static text can&rsquo;t be graphed</td>
			</tr>
			<tr>
					<td><code>process_max_fds</code>, <code>vllm:cache_config_info</code>, <code>vllm:engine_sleep_state</code></td>
					<td>static configuration, not performance</td>
			</tr>
			<tr>
					<td><code>vllm:request_params_*</code></td>
					<td><strong>cardinality explosion</strong> — a series per distinct <code>max_tokens</code>/<code>n</code></td>
			</tr>
			<tr>
					<td><code>vllm:iteration_tokens_total</code></td>
					<td>redundant with tokens/s</td>
			</tr>
			<tr>
					<td><code>http_request_duration_highr_seconds</code></td>
					<td>&ldquo;high-resolution&rdquo; = hundreds of buckets; the standard one suffices</td>
			</tr>
			<tr>
					<td><code>http_request/response_size_bytes</code> buckets</td>
					<td>dropped, but <code>_sum</code> harvested for an average</td>
			</tr>
			<tr>
					<td><code>*_created</code></td>
					<td>bucket-initialisation timestamps; pure noise</td>
			</tr>
	</tbody>
</table>
<p>Everything that survives is truncated to two decimals before push — a
small mercy for the TSDB.</p>
<h2 id="the-key-hierarchy">The key hierarchy</h2>
<p>Ops shows metrics as a tree, so the names are designed to browse:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vllm|system|is_up
</span></span><span class="line"><span class="cl">vllm|throughput|total_tokens_per_sec
</span></span><span class="line"><span class="cl">vllm|perf|live_avg_ttft_ms
</span></span><span class="line"><span class="cl">vllm|perf|ttft|p95_ms
</span></span><span class="line"><span class="cl">vllm|queue|pressure_ratio
</span></span><span class="line"><span class="cl">vllm|memory|kv_cache_pct
</span></span><span class="line"><span class="cl">vllm|cache|live_prefix_hit_rate_pct
</span></span><span class="line"><span class="cl">vllm|process|rss_memory_gb
</span></span></code></pre></div><p><code>vllm | category | metric [| percentile]</code>. An operator who has never seen
vLLM can find &ldquo;time to first token, 95th percentile&rdquo; without a manual.</p>
<p><img alt="The key hierarchy as it lands in VCF Operations" loading="lazy" src="/images/ui/o5-ops-vllm-metric-tree.jpg">
<em><code>vllm | category | metric</code> in the Ops metric picker — see <a href="/series/llm-ops-on-vcf/">part 2</a> for how this was captured.</em></p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Organisations putting language models into service quickly discover that
&ldquo;is it up?&rdquo; isn&rsquo;t the question. The questions are: how long are users
waiting for the first word, is the service about to run out of memory, and
do we need another GPU before Friday? This pipeline answers them inside the
same VCF Operations console the infrastructure team already lives in, so
AI services get the same capacity planning, alerting and dashboards as
everything else — no second monitoring stack, no new team to staff it.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li><strong>Never push raw histogram buckets</strong> into a gauge-oriented TSDB.
Interpolate P50/P95/P99 client-side and push four gauges.</li>
<li>Counters need <strong>state</strong>: keep the last sample, compute deltas, and drop
the baseline after a long gap (300 s) rather than dilute the rate.</li>
<li><code>Δsum/Δcount</code> gives you <em>live</em> mean latency — far more useful than the
all-time mean.</li>
<li>One derived <strong>saturation score</strong> beats six raw gauges on the top of a
dashboard. Make it explainable (KV% × 0.5 + pressure × 25).</li>
<li>The drop-list is a design artefact, not housekeeping. <code>request_params_*</code>
alone can double your series count.</li>
<li>Name for browsing: <code>product | category | metric</code>.</li>
</ul>
<p><em>Part 2: <a href="/series/llm-ops-on-vcf/">running it as a service, monitoring the monitor, and the four
alerts</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Config shown is sanitised; the script&rsquo;s
test mode reads a local metrics file instead of a live endpoint.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
