<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Alerting on The Nested Lab</title>
    <link>https://thenestedlab.com/tags/alerting/</link>
    <description>Recent content in Alerting on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Wed, 07 Oct 2026 00:00:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/tags/alerting/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>vLLM metrics into VCF Operations, part 2: monitor the monitor</title>
      <link>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-2/</link>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-2/</guid>
      <description>Part 1&amp;rsquo;s design only helps if it runs every minute, for ever, and says when it stops. Self-monitoring through Telegraf, batched pushes, the four alerts for day one, and an offline test mode.</description>
      <content:encoded><![CDATA[<p><a href="/posts/vllm-metrics-vcf-ops-part-1/">Part 1</a> was about <em>what</em> to push.
This one is about making it boring: scheduled, self-monitoring, alertable,
and testable without a live model. In monitoring, boring is the highest
compliment going.</p>
<h2 id="execution-pipeline">Execution pipeline</h2>
<p>Each run does seven things:</p>
<ol>
<li><strong>Authenticate once</strong> to VCF Operations. PowerCLI&rsquo;s Ops module gets the
token; native <code>Invoke-RestMethod</code> does the rest, far faster.</li>
<li><strong>Iterate targets</strong> from the config file.</li>
<li><strong>Resolve the resource:</strong> look up the Ops object by name. The target
must already exist in Ops (a VM, a container object, a custom
application). Exact match first, then a partial match with a warning.</li>
<li><strong>Scrape</strong> the Prometheus text from <code>:8000/metrics</code>.</li>
<li><strong>Parse and transform:</strong> gauges, counters, histograms; percentiles,
deltas, derived scores; drop-list applied.</li>
<li><strong>Push in batches of 1000</strong> stats per request. Ops rejects oversized
payloads. A busy model produces ~150 stats per run, which fits easily;
a config with ten targets doesn&rsquo;t.</li>
<li><strong>Persist state:</strong> counters and timestamps for each target. Targets
removed from the config are pruned from the state file automatically.</li>
</ol>
<h2 id="the-stdout-contract">The stdout contract</h2>
<p>The single most useful design choice is this. When run non-interactively,
<strong>the script prints exactly one integer</strong>: the total number of stats
pushed. It&rsquo;s not a chatty script, but it is honest. Mostly; more on that
below.</p>
<p>That makes it a Telegraf <code>inputs.exec</code> plugin with zero glue:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-toml" data-lang="toml"><span class="line"><span class="cl"><span class="p">[[</span><span class="nx">inputs</span><span class="p">.</span><span class="nx">exec</span><span class="p">]]</span>
</span></span><span class="line"><span class="cl">  <span class="nx">commands</span> <span class="p">=</span> <span class="p">[</span><span class="s1">&#39;powershell.exe -NoProfile -File C:\monitoring\Push-LLMMetrics.ps1&#39;</span><span class="p">]</span>
</span></span><span class="line"><span class="cl">  <span class="nx">timeout</span> <span class="p">=</span> <span class="s2">&#34;50s&#34;</span>
</span></span><span class="line"><span class="cl">  <span class="nx">interval</span> <span class="p">=</span> <span class="s2">&#34;60s&#34;</span>
</span></span><span class="line"><span class="cl">  <span class="nx">data_format</span> <span class="p">=</span> <span class="s2">&#34;value&#34;</span>
</span></span><span class="line"><span class="cl">  <span class="nx">data_type</span> <span class="p">=</span> <span class="s2">&#34;integer&#34;</span>
</span></span><span class="line"><span class="cl">  <span class="nx">name_override</span> <span class="p">=</span> <span class="s2">&#34;llm_metrics_pushed&#34;</span>
</span></span></code></pre></div><p>Telegraf now has a series called <code>llm_metrics_pushed</code> that should read
~150 every minute. If it reads 0, the scrape failed. If it&rsquo;s <em>missing</em>, the
script didn&rsquo;t run. The integration monitors itself by construction, and
the thing watching it is the same Telegraf you already run.</p>
<p>(Task Scheduler or cron work too, running
<code>powershell.exe -File Push-LLMMetrics.ps1</code> every 60 s. You just lose the
free self-monitoring.)</p>
<h2 id="configuration">Configuration</h2>
<p>Credentials and targets live in a JSON file beside the script, never in it:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="p">{</span>
</span></span><span class="line"><span class="cl">  <span class="nt">&#34;vrops&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;server&#34;</span><span class="p">:</span> <span class="s2">&#34;vcfops.example.lab&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;user&#34;</span><span class="p">:</span> <span class="s2">&#34;svc_vllm_metrics&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;password&#34;</span><span class="p">:</span> <span class="s2">&#34;&lt;from-credential-store&gt;&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;auth_source&#34;</span><span class="p">:</span> <span class="s2">&#34;local&#34;</span>
</span></span><span class="line"><span class="cl">  <span class="p">},</span>
</span></span><span class="line"><span class="cl">  <span class="nt">&#34;targets&#34;</span><span class="p">:</span> <span class="p">[</span>
</span></span><span class="line"><span class="cl">    <span class="p">{</span> <span class="nt">&#34;subject&#34;</span><span class="p">:</span> <span class="s2">&#34;LLM01&#34;</span><span class="p">,</span> <span class="nt">&#34;metrics_url&#34;</span><span class="p">:</span> <span class="s2">&#34;http://10.0.0.5:8000/metrics&#34;</span> <span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="p">{</span> <span class="nt">&#34;subject&#34;</span><span class="p">:</span> <span class="s2">&#34;LLM02&#34;</span><span class="p">,</span> <span class="nt">&#34;metrics_url&#34;</span><span class="p">:</span> <span class="s2">&#34;http://10.0.0.6:8000/metrics&#34;</span> <span class="p">}</span>
</span></span><span class="line"><span class="cl">  <span class="p">]</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p><code>subject</code> must match the Ops resource name exactly. <code>-ConfigPath</code> overrides
the location for scheduled tasks. <code>auth_source</code> is <code>local</code> or your Active
Directory domain.</p>
<h2 id="logging-that-doesnt-fill-the-disk">Logging that doesn&rsquo;t fill the disk</h2>
<p>Interactive runs get colour-coded console output. Silent runs copy
everything to <code>llm-metrics.log</code>, which rotates itself to <code>.old</code> at 5 MB. A
log that fills the disk would turn the monitoring into the incident. The
three warnings you&rsquo;ll actually see:</p>
<table>
	<thead>
			<tr>
					<th>Log line</th>
					<th>Meaning</th>
					<th>Action</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td><code>Resource NOT found with exact name 'LLM01' - looking for partial match</code></td>
					<td>Ops object name doesn&rsquo;t match <code>subject</code></td>
					<td>create the object or fix the spelling</td>
			</tr>
			<tr>
					<td><code>Rates are exactly 0</code></td>
					<td>no traffic since last run</td>
					<td>normal when idle</td>
			</tr>
			<tr>
					<td><code>Time gap too large (&gt;300s). Resetting baseline</code></td>
					<td>scheduler paused / host rebooted</td>
					<td>normal; rates resume next run</td>
			</tr>
	</tbody>
</table>
<p>And one error: <code>API Error on batch: 401 Unauthorized</code>. The service account
has expired or been locked. It&rsquo;s the only failure that needs a human.</p>
<h2 id="the-offline-test-mode">The offline test mode</h2>
<p>You do not need a GPU to build this. <code>metrics_url</code> accepts <code>file://</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="p">{</span> <span class="nt">&#34;subject&#34;</span><span class="p">:</span> <span class="s2">&#34;LLM01&#34;</span><span class="p">,</span> <span class="nt">&#34;metrics_url&#34;</span><span class="p">:</span> <span class="s2">&#34;file://C:/temp/llm_metrics.txt&#34;</span> <span class="p">}</span>
</span></span></code></pre></div><p>Save one real <code>/metrics</code> scrape to a text file and point the config at it.
Then iterate on parsing and Ops resource mapping at your desk.
To test the rate maths, bump the counter values in the file between runs.
The script can&rsquo;t tell the difference, and neither can Ops.</p>
<h2 id="the-four-alerts-to-configure-on-day-one">The four alerts to configure on day one</h2>
<p>Everything in part 1 was designed so these four symptoms are one metric
each:</p>
<table>
	<thead>
			<tr>
					<th>Alert</th>
					<th>Symptom</th>
					<th>What it means</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td><strong>vLLM Engine Down</strong></td>
					<td><code>vllm|system|is_up</code> missing or &lt; 1</td>
					<td>container crashed or <code>/metrics</code> unreachable</td>
			</tr>
			<tr>
					<td><strong>OOM Imminent</strong></td>
					<td><code>vllm|perf|system_saturation_score</code> &gt; 90</td>
					<td>KV cache full, queue stacking — new prompts will be rejected</td>
			</tr>
			<tr>
					<td><strong>Severe End-User Lag</strong></td>
					<td><code>vllm|perf|live_avg_tpot_ms</code> &gt; 100</td>
					<td>generating slower than people read (~50 ms/token); users see stutter</td>
			</tr>
			<tr>
					<td><strong>Queue Spillage</strong></td>
					<td><code>vllm|queue|requests_swapped</code> &gt; 0</td>
					<td>engine evicting live requests to CPU RAM — should <em>always</em> be 0</td>
			</tr>
	</tbody>
</table>
<p>Suggested thresholds for the rest, from the field:</p>
<ul>
<li><code>live_avg_ttft_ms</code>: warn 1000, critical 3000 (a slow first token drives users to give up)</li>
<li><code>queue|pressure_ratio</code>: warn 0.5, critical 1.0 (more waiting than running means scale out)</li>
<li><code>memory|kv_cache_pct</code>: warn 85, critical 95</li>
<li><code>queue|preemptions_per_sec</code>: warn above 0</li>
<li><code>http|requests_per_sec</code> well above <code>throughput|requests_per_sec</code>: users are getting 429/503 at the front door</li>
</ul>
<h2 id="what-running-it-against-a-fresh-ops-taught-me">What running it against a fresh Ops taught me</h2>
<p>I re-ran the collector in <code>file://</code> mode against a lab VCF Operations 9.1
that had never seen it, pointing at an existing VM object. Three findings
in the first ten minutes, which is either very efficient or slightly
worrying. All three are worth knowing before you deploy:</p>
<ol>
<li><strong>The stdout integer lied once.</strong> The first two runs logged
<code>[ERR] The SSL connection could not be established</code> on the push, and
still printed <code>Total stats pushed: 76</code>. The counter tallies stats
<em>prepared</em>, not batches <em>accepted</em>. Telegraf would have seen a healthy
76 while nothing landed. The fix for v6.4: increment only on a 2xx from
the batch push, and print 0 on any batch error. Until then, alert on
the log&rsquo;s <code>[ERR]</code> lines too.</li>
<li><strong>The push path needs its own certificate handling.</strong> The PowerCLI
connect honours <code>Set-PowerCLIConfiguration -InvalidCertificateAction Ignore</code>; the native <code>Invoke-RestMethod</code> pushes don&rsquo;t. On PowerShell 7, set
<code>$PSDefaultParameterValues['Invoke-RestMethod:SkipCertificateCheck']=$true</code>
before the run, or trust the Ops CA properly. Otherwise auth succeeds
and every push fails.</li>
<li><strong>Metric names drift between vLLM versions.</strong> Newer vLLM exposes
<code>vllm:time_per_output_token_seconds</code>, but the key builder was written
for <code>vllm:request_time_per_output_token_seconds</code>. So the newer name fell
through the hierarchy and landed as a raw top-level metric in Ops
(<code>vllm_time_per_output_token_seconds</code>) instead of <code>vllm|perf|tpot|*</code>.
Everything else arrived exactly where the design says:
<code>vllm|perf|ttft|p95_ms</code>, <code>vllm|throughput|total_tokens_per_sec</code>,
<code>vllm|perf|system_saturation_score</code> and the <code>http|request_duration</code>
percentiles. Add the new name to the mapping, and treat &ldquo;unexpected
top-level metric&rdquo; as a check in the smoke test.</li>
</ol>
<p>Everything else worked first time: the resource resolved by name, 89 stats
per run, batches accepted, and rates flowing from the second run onwards.</p>
<p><img alt="VCF Operations: the vllm metric tree on the target object — requests, system|is_up, throughput, tokens — with prompt_total and live_avg_generation_tokens_per_req charted over the last hour" loading="lazy" src="/images/ui/o5-ops-vllm-metric-tree.jpg">
<em>The tree as an operator sees it, browsable without a manual. Captured from the <code>file://</code> test run above. The pipeline and the Ops side are real; the model behind the numbers was a saved scrape with advancing counters.</em></p>
<h2 id="what-id-change">What I&rsquo;d change</h2>
<p>Two things, honestly:</p>
<ul>
<li>It&rsquo;s PowerShell because the monitoring host was Windows and PowerCLI was
already there. The same design ports to Python in an afternoon. The
value is in the <em>transformation</em>, not the language.</li>
<li>Resolving the Ops resource by <em>name</em> is fragile. Resolving it by an
identifier, stored in the config after the first lookup, would survive a
rename.</li>
</ul>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>The operational payoff is an AI service that can be <em>run</em>, not just
hosted. You get four alerts that fire before users complain, and a
saturation dial that tells capacity planners when to add a GPU. You also
get an integration that reports its own health, so silence is never
mistaken for calm.</p>
<p>For a customer, that&rsquo;s the difference between an LLM pilot and an LLM in
production. And the design ports to any workload that speaks Prometheus
but has no VCF Operations adapter of its own.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Make the collector&rsquo;s <strong>stdout a single integer</strong> and let the scheduler
(Telegraf) monitor the integration for free.</li>
<li>Batch pushes (1000 stats). Authenticate once per run, not per push.</li>
<li><code>file://</code> targets let you develop the whole pipeline with zero GPUs.</li>
<li>Four alerts cover 90 % of incidents: down, saturated, laggy, swapping.
Everything else is a threshold on an existing gauge.</li>
<li>Log rotation is part of the script, not an ops afterthought.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/workload-monitoring-and-observability/os-and-application-monitoring/steps-to-monitor-your-applications/additional-operations-from-the-manage-agents-tab/custom-script.html">Custom Script</a>: the Telegraf agent&rsquo;s own exec-based script check, which also expects a single integer</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/administration-sdks-cli-and-tools/understanding-the-vr-ops-api/getting-started-with-the-api/acquire-an-authentication-token.html">Acquire an Authentication Token</a>: <code>POST /suite-api/api/auth/token/acquire</code>, the <code>authSource</code> field, and a token reusable for six hours</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/configuring-alerts-and-actions/alert-definitions.html">Alert Definitions in VCF Operations</a>: alerts built from symptoms and recommendations</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/configuring-alerts-and-actions/symptom-definitions.html">Symptom Definitions in VCF Operations</a>: metric symptoms at Warning and Critical levels, for the thresholds</li>
</ul>
<p><em>Previously: <a href="/posts/vllm-metrics-vcf-ops-part-1/">don&rsquo;t ship the histogram</a>.
Related: <a href="/series/observability-on-vcf/">Telegraf on Windows Server 2025</a> —
where this script actually runs.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Config sanitised.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
