<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Private AI on The Nested Lab</title>
    <link>https://thenestedlab.com/products/private-ai/</link>
    <description>Recent content in Private AI on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/products/private-ai/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>vLLM metrics into VCF Operations, part 1: don&#39;t ship the histogram</title>
      <link>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-1/</link>
      <pubDate>Wed, 30 Sep 2026 00:00:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-1/</guid>
      <description>vLLM exposes about 200 Prometheus series per model, and pushing them raw into VCF Operations is a cardinality bomb. A pipeline that computes the useful numbers client-side and drops what can&amp;rsquo;t be graphed.</description>
      <content:encoded><![CDATA[<p>vLLM&rsquo;s <code>/metrics</code> endpoint is generous. Every model instance exposes a few
hundred Prometheus series: counters, gauges and histograms. The histograms
are the problem, with dozens of buckets each.</p>
<p>Point a naïve scraper at it, push everything into VCF Operations, and you
get two things. One is a time-series database (TSDB) full of
<code>_bucket{le=&quot;0.05&quot;}</code> series nobody will ever graph. The other is a
dashboard that tells an operator nothing, in tremendous detail.</p>
<p>This two-parter is the pipeline that fixed that. Part 1 is the design:
what to compute, what to drop, and why. <a href="/series/llm-ops-on-vcf/">Part 2</a>
is the operations story: scheduling, self-monitoring and alerts.</p>
<h2 id="the-core-decision-percentiles-are-computed-here-not-there">The core decision: percentiles are computed here, not there</h2>
<p>A Prometheus histogram is a set of cumulative bucket counters:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.1&#34;}   1203
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.25&#34;}  4871
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.5&#34;}   7120
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;+Inf&#34;}  7412
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_sum                1812.4
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_count              7412
</span></span></code></pre></div><p>Prometheus turns that into a P95 at query time with <code>histogram_quantile</code>.
VCF Operations has no such function: it stores gauges. So the pipeline
does the interpolation <em>before</em> pushing:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="kd">function</span><span class="w"> </span><span class="nb">Get-Percentile</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="k">param</span><span class="p">([</span><span class="no">array</span><span class="p">]</span><span class="nv">$Buckets</span><span class="p">,</span> <span class="p">[</span><span class="no">double</span><span class="p">]</span><span class="nv">$Total</span><span class="p">,</span> <span class="p">[</span><span class="no">double</span><span class="p">]</span><span class="nv">$Percentile</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="nv">$target</span> <span class="p">=</span> <span class="nv">$Percentile</span> <span class="p">*</span> <span class="nv">$Total</span>
</span></span><span class="line"><span class="cl">    <span class="nv">$prevLe</span> <span class="p">=</span> <span class="mf">0.0</span><span class="p">;</span> <span class="nv">$prevCount</span> <span class="p">=</span> <span class="mf">0.0</span>
</span></span><span class="line"><span class="cl">    <span class="k">foreach</span> <span class="p">(</span><span class="nv">$b</span> <span class="k">in</span> <span class="nv">$Buckets</span><span class="p">)</span> <span class="p">{</span>                    <span class="c"># sorted by le, +Inf last</span>
</span></span><span class="line"><span class="cl">        <span class="k">if</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Count</span> <span class="o">-ge</span> <span class="nv">$target</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">            <span class="k">if</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Le</span> <span class="o">-eq</span> <span class="p">[</span><span class="no">double</span><span class="p">]::</span><span class="n">PositiveInfinity</span><span class="p">)</span> <span class="p">{</span> <span class="k">return</span> <span class="nv">$prevLe</span> <span class="p">}</span>
</span></span><span class="line"><span class="cl">            <span class="nv">$frac</span> <span class="p">=</span> <span class="p">(</span><span class="nv">$target</span> <span class="p">-</span> <span class="nv">$prevCount</span><span class="p">)</span> <span class="p">/</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Count</span> <span class="p">-</span> <span class="nv">$prevCount</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">            <span class="k">return</span> <span class="nv">$prevLe</span> <span class="p">+</span> <span class="nv">$frac</span> <span class="p">*</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Le</span> <span class="p">-</span> <span class="nv">$prevLe</span><span class="p">)</span>   <span class="c"># linear within the bucket</span>
</span></span><span class="line"><span class="cl">        <span class="p">}</span>
</span></span><span class="line"><span class="cl">        <span class="nv">$prevLe</span> <span class="p">=</span> <span class="nv">$b</span><span class="p">.</span><span class="n">Le</span><span class="p">;</span> <span class="nv">$prevCount</span> <span class="p">=</span> <span class="nv">$b</span><span class="p">.</span><span class="py">Count</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>That&rsquo;s linear interpolation within the bucket that crosses the target
rank, the same approximation Prometheus makes. Out come <strong>four flat gauges
per histogram</strong>: <code>p50_ms</code>, <code>p95_ms</code>, <code>p99_ms</code> and <code>avg_ms</code> (from
<code>_sum/_count</code>). Twenty-odd series become four, and they&rsquo;re the four an
operator reads.</p>
<p>It&rsquo;s applied to time to first token (TTFT), end-to-end latency,
inter-token latency, time per output token, prefill time, decode time,
inference time and queue time.</p>
<h2 id="rates-need-memory">Rates need memory</h2>
<p>Counters (<code>generation_tokens_total</code>, <code>request_success_total</code>) are useless
as absolute values. What you want is tokens <em>per second</em>, and that needs
the previous sample. So the script keeps state. <code>llm-metrics-state.json</code>
holds the last counters and timestamp for each target, and each run works
out the deltas:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">tokens_per_sec = (tokens_now - tokens_prev) / (t_now - t_prev)
</span></span></code></pre></div><p>The same trick goes one step further for <strong>live latency</strong>. <code>_sum</code> and
<code>_count</code> are both counters, so <code>Δsum / Δcount</code> is the <em>mean over the last
interval</em>, not the all-time mean the histogram gives you. That&rsquo;s how you
get a <code>live_avg_ttft_ms</code> that reflects the last 60 seconds instead of the
last fortnight.</p>
<p>Two guards make this safe:</p>
<ul>
<li><strong>Stale-state protection.</strong> If the gap since the last run is over 300 s
(scheduler stopped, server rebooted), the baseline is dropped. Otherwise
you&rsquo;d get a diluted &ldquo;per-second&rdquo; rate, averaged over an hour.</li>
<li><strong>Restart detection.</strong> A counter that went <em>down</em> means vLLM restarted,
because counters don&rsquo;t go backwards for fun. The delta for that cycle is
discarded.</li>
</ul>
<h2 id="derived-metrics-what-the-raw-numbers-wont-tell-you">Derived metrics: what the raw numbers won&rsquo;t tell you</h2>
<p>Two computed values earn their place at the top of the dashboard.</p>
<p><strong>Queue pressure ratio:</strong> <code>(waiting + swapped) / (running + 1)</code>. Above
1.0, more requests are waiting than being served, so scale out.</p>
<p><strong>System saturation score</strong> (0–100):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="nv">$saturation</span> <span class="p">=</span> <span class="p">(</span><span class="nv">$kvCachePct</span> <span class="p">*</span> <span class="mf">0.5</span><span class="p">)</span> <span class="p">+</span> <span class="p">(</span><span class="nv">$queuePressure</span> <span class="p">*</span> <span class="mf">25.0</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="k">if</span> <span class="p">(</span><span class="nv">$saturation</span> <span class="o">-gt</span> <span class="mf">100</span><span class="p">)</span> <span class="p">{</span> <span class="nv">$saturation</span> <span class="p">=</span> <span class="mf">100</span> <span class="p">}</span>
</span></span></code></pre></div><p>A full KV (key-value) cache alone scores 50, and queue pressure of 2.0
scores the other 50. It&rsquo;s a heuristic, and it&rsquo;s deliberately one number.</p>
<p>Think of it as a dial that goes red when the engine is about to start
swapping requests to CPU memory. That&rsquo;s the moment latency falls off a
cliff. Warning is at 75, critical at 90.</p>
<p><img alt="vllm|perf|system_saturation_score over the last hour in VCF Operations" loading="lazy" src="/images/ui/o6-ops-vllm-saturation-score.jpg">
<em>The dial, as Ops draws it: one gauge climbing towards the warning line as KV-cache use and queue pressure rise together. (Test-mode data: see part 2.)</em></p>
<p>Also derived:</p>
<ul>
<li>prefix-cache hit rate, live and all-time;</li>
<li>average request size, from the dropped <code>http_request_size_bytes</code> <code>_sum</code>;</li>
<li>uptime in days;</li>
<li><code>is_up = 1</code> on every successful scrape, so the <em>absence</em> of the metric
is the alert.</li>
</ul>
<p>That last one does its most useful work by not turning up.</p>
<h2 id="what-gets-dropped-and-why">What gets dropped, and why</h2>
<p>The <code>$DropPrefixes</code> list is as important as anything computed:</p>
<table>
	<thead>
			<tr>
					<th>Dropped</th>
					<th>Why</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td><code>python_gc_*</code>, <code>python_info</code></td>
					<td>runtime noise; static text can&rsquo;t be graphed</td>
			</tr>
			<tr>
					<td><code>process_max_fds</code>, <code>vllm:cache_config_info</code>, <code>vllm:engine_sleep_state</code></td>
					<td>static configuration, not performance</td>
			</tr>
			<tr>
					<td><code>vllm:request_params_*</code></td>
					<td><strong>cardinality explosion</strong> — a series per distinct <code>max_tokens</code>/<code>n</code></td>
			</tr>
			<tr>
					<td><code>vllm:iteration_tokens_total</code></td>
					<td>redundant with tokens/s</td>
			</tr>
			<tr>
					<td><code>http_request_duration_highr_seconds</code></td>
					<td>&ldquo;high-resolution&rdquo; = hundreds of buckets; the standard one suffices</td>
			</tr>
			<tr>
					<td><code>http_request/response_size_bytes</code> buckets</td>
					<td>dropped, but <code>_sum</code> harvested for an average</td>
			</tr>
			<tr>
					<td><code>*_created</code></td>
					<td>bucket-initialisation timestamps; pure noise</td>
			</tr>
	</tbody>
</table>
<p>Everything that survives is truncated to two decimals before the push.
It&rsquo;s a small mercy for the TSDB.</p>
<h2 id="the-key-hierarchy">The key hierarchy</h2>
<p>Ops shows metrics as a tree, so the names are designed for browsing:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vllm|system|is_up
</span></span><span class="line"><span class="cl">vllm|throughput|total_tokens_per_sec
</span></span><span class="line"><span class="cl">vllm|perf|live_avg_ttft_ms
</span></span><span class="line"><span class="cl">vllm|perf|ttft|p95_ms
</span></span><span class="line"><span class="cl">vllm|queue|pressure_ratio
</span></span><span class="line"><span class="cl">vllm|memory|kv_cache_pct
</span></span><span class="line"><span class="cl">vllm|cache|live_prefix_hit_rate_pct
</span></span><span class="line"><span class="cl">vllm|process|rss_memory_gb
</span></span></code></pre></div><p><code>vllm | category | metric [| percentile]</code>. An operator who has never seen
vLLM can find &ldquo;time to first token, 95th percentile&rdquo; without a manual.</p>
<p><img alt="The key hierarchy as it lands in VCF Operations" loading="lazy" src="/images/ui/o5-ops-vllm-metric-tree.jpg">
<em><code>vllm | category | metric</code> in the Ops metric picker — see <a href="/series/llm-ops-on-vcf/">part 2</a> for how this was captured.</em></p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Organisations putting language models into service soon discover that &ldquo;is
it up?&rdquo; isn&rsquo;t the question. The real questions are these. How long are
users waiting for the first word? Is the service about to run out of
memory? Do we need another GPU before Friday?</p>
<p>This pipeline answers them inside the same VCF Operations console the
infrastructure team already lives in. So AI services get the same capacity
planning, alerting and dashboards as everything else. There&rsquo;s no second
monitoring stack, and no new team to staff it.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li><strong>Never push raw histogram buckets</strong> into a gauge-oriented TSDB.
Interpolate P50/P95/P99 client-side and push four gauges.</li>
<li>Counters need <strong>state</strong>: keep the last sample, compute deltas, and drop
the baseline after a long gap (300 s) rather than dilute the rate.</li>
<li><code>Δsum/Δcount</code> gives you <em>live</em> mean latency, which is far more useful
than the all-time mean.</li>
<li>One derived <strong>saturation score</strong> beats six raw gauges at the top of a
dashboard. Make it explainable (KV% × 0.5 + pressure × 25).</li>
<li>The drop-list is a design artefact, not housekeeping. <code>request_params_*</code>
alone can double your series count.</li>
<li>Name for browsing: <code>product | category | metric</code>.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/configuring-super-metrics/super-metrics-tab/super-metric-functions-and-operators.html">Super Metric Functions and Operators</a>: the functions an Ops formula can use; none of them is a percentile</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/administration-sdks-cli-and-tools/understanding-the-vr-ops-api/using-the-api-with-vrealize-operations-manager.html">Using the API with VCF Operations</a>: the REST API the pipeline pushes through, and the Swagger reference on the appliance</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/administration-sdks-cli-and-tools/understanding-the-vr-ops-api/getting-started-with-the-api/generate-a-list-of-all-metrics-for-the-object.html">Generate a List of All Metrics for the Object</a>: reading an object&rsquo;s stat keys back, grouped with <code>|</code> as in <code>mem|host_workload</code></li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/configuring-alerts-and-actions/symptom-definitions.html">Symptom Definitions in VCF Operations</a>: metric symptoms at Warning and Critical levels, as for the saturation score</li>
</ul>
<p><em>Part 2: <a href="/series/llm-ops-on-vcf/">running it as a service, monitoring the monitor, and the four
alerts</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Config shown is sanitised; the script&rsquo;s
test mode reads a local metrics file instead of a live endpoint.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
