<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>VCF Operations on The Nested Lab</title>
    <link>https://thenestedlab.com/products/vcf-operations/</link>
    <description>Recent content in VCF Operations on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/products/vcf-operations/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>vLLM metrics into VCF Operations, part 1: don&#39;t ship the histogram</title>
      <link>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-1/</link>
      <pubDate>Wed, 30 Sep 2026 00:00:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/vllm-metrics-vcf-ops-part-1/</guid>
      <description>vLLM exposes about 200 Prometheus series per model, and pushing them raw into VCF Operations is a cardinality bomb. A pipeline that computes the useful numbers client-side and drops what can&amp;rsquo;t be graphed.</description>
      <content:encoded><![CDATA[<p>vLLM&rsquo;s <code>/metrics</code> endpoint is generous. Every model instance exposes a few
hundred Prometheus series: counters, gauges and histograms. The histograms
are the problem, with dozens of buckets each.</p>
<p>Point a naïve scraper at it, push everything into VCF Operations, and you
get two things. One is a time-series database (TSDB) full of
<code>_bucket{le=&quot;0.05&quot;}</code> series nobody will ever graph. The other is a
dashboard that tells an operator nothing, in tremendous detail.</p>
<p>This two-parter is the pipeline that fixed that. Part 1 is the design:
what to compute, what to drop, and why. <a href="/series/llm-ops-on-vcf/">Part 2</a>
is the operations story: scheduling, self-monitoring and alerts.</p>
<h2 id="the-core-decision-percentiles-are-computed-here-not-there">The core decision: percentiles are computed here, not there</h2>
<p>A Prometheus histogram is a set of cumulative bucket counters:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.1&#34;}   1203
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.25&#34;}  4871
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;0.5&#34;}   7120
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_bucket{le=&#34;+Inf&#34;}  7412
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_sum                1812.4
</span></span><span class="line"><span class="cl">vllm:time_to_first_token_seconds_count              7412
</span></span></code></pre></div><p>Prometheus turns that into a P95 at query time with <code>histogram_quantile</code>.
VCF Operations has no such function: it stores gauges. So the pipeline
does the interpolation <em>before</em> pushing:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="kd">function</span><span class="w"> </span><span class="nb">Get-Percentile</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="k">param</span><span class="p">([</span><span class="no">array</span><span class="p">]</span><span class="nv">$Buckets</span><span class="p">,</span> <span class="p">[</span><span class="no">double</span><span class="p">]</span><span class="nv">$Total</span><span class="p">,</span> <span class="p">[</span><span class="no">double</span><span class="p">]</span><span class="nv">$Percentile</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="nv">$target</span> <span class="p">=</span> <span class="nv">$Percentile</span> <span class="p">*</span> <span class="nv">$Total</span>
</span></span><span class="line"><span class="cl">    <span class="nv">$prevLe</span> <span class="p">=</span> <span class="mf">0.0</span><span class="p">;</span> <span class="nv">$prevCount</span> <span class="p">=</span> <span class="mf">0.0</span>
</span></span><span class="line"><span class="cl">    <span class="k">foreach</span> <span class="p">(</span><span class="nv">$b</span> <span class="k">in</span> <span class="nv">$Buckets</span><span class="p">)</span> <span class="p">{</span>                    <span class="c"># sorted by le, +Inf last</span>
</span></span><span class="line"><span class="cl">        <span class="k">if</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Count</span> <span class="o">-ge</span> <span class="nv">$target</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">            <span class="k">if</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Le</span> <span class="o">-eq</span> <span class="p">[</span><span class="no">double</span><span class="p">]::</span><span class="n">PositiveInfinity</span><span class="p">)</span> <span class="p">{</span> <span class="k">return</span> <span class="nv">$prevLe</span> <span class="p">}</span>
</span></span><span class="line"><span class="cl">            <span class="nv">$frac</span> <span class="p">=</span> <span class="p">(</span><span class="nv">$target</span> <span class="p">-</span> <span class="nv">$prevCount</span><span class="p">)</span> <span class="p">/</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Count</span> <span class="p">-</span> <span class="nv">$prevCount</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">            <span class="k">return</span> <span class="nv">$prevLe</span> <span class="p">+</span> <span class="nv">$frac</span> <span class="p">*</span> <span class="p">(</span><span class="nv">$b</span><span class="p">.</span><span class="py">Le</span> <span class="p">-</span> <span class="nv">$prevLe</span><span class="p">)</span>   <span class="c"># linear within the bucket</span>
</span></span><span class="line"><span class="cl">        <span class="p">}</span>
</span></span><span class="line"><span class="cl">        <span class="nv">$prevLe</span> <span class="p">=</span> <span class="nv">$b</span><span class="p">.</span><span class="n">Le</span><span class="p">;</span> <span class="nv">$prevCount</span> <span class="p">=</span> <span class="nv">$b</span><span class="p">.</span><span class="py">Count</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>That&rsquo;s linear interpolation within the bucket that crosses the target
rank, the same approximation Prometheus makes. Out come <strong>four flat gauges
per histogram</strong>: <code>p50_ms</code>, <code>p95_ms</code>, <code>p99_ms</code> and <code>avg_ms</code> (from
<code>_sum/_count</code>). Twenty-odd series become four, and they&rsquo;re the four an
operator reads.</p>
<p>It&rsquo;s applied to time to first token (TTFT), end-to-end latency,
inter-token latency, time per output token, prefill time, decode time,
inference time and queue time.</p>
<h2 id="rates-need-memory">Rates need memory</h2>
<p>Counters (<code>generation_tokens_total</code>, <code>request_success_total</code>) are useless
as absolute values. What you want is tokens <em>per second</em>, and that needs
the previous sample. So the script keeps state. <code>llm-metrics-state.json</code>
holds the last counters and timestamp for each target, and each run works
out the deltas:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">tokens_per_sec = (tokens_now - tokens_prev) / (t_now - t_prev)
</span></span></code></pre></div><p>The same trick goes one step further for <strong>live latency</strong>. <code>_sum</code> and
<code>_count</code> are both counters, so <code>Δsum / Δcount</code> is the <em>mean over the last
interval</em>, not the all-time mean the histogram gives you. That&rsquo;s how you
get a <code>live_avg_ttft_ms</code> that reflects the last 60 seconds instead of the
last fortnight.</p>
<p>Two guards make this safe:</p>
<ul>
<li><strong>Stale-state protection.</strong> If the gap since the last run is over 300 s
(scheduler stopped, server rebooted), the baseline is dropped. Otherwise
you&rsquo;d get a diluted &ldquo;per-second&rdquo; rate, averaged over an hour.</li>
<li><strong>Restart detection.</strong> A counter that went <em>down</em> means vLLM restarted,
because counters don&rsquo;t go backwards for fun. The delta for that cycle is
discarded.</li>
</ul>
<h2 id="derived-metrics-what-the-raw-numbers-wont-tell-you">Derived metrics: what the raw numbers won&rsquo;t tell you</h2>
<p>Two computed values earn their place at the top of the dashboard.</p>
<p><strong>Queue pressure ratio:</strong> <code>(waiting + swapped) / (running + 1)</code>. Above
1.0, more requests are waiting than being served, so scale out.</p>
<p><strong>System saturation score</strong> (0–100):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="nv">$saturation</span> <span class="p">=</span> <span class="p">(</span><span class="nv">$kvCachePct</span> <span class="p">*</span> <span class="mf">0.5</span><span class="p">)</span> <span class="p">+</span> <span class="p">(</span><span class="nv">$queuePressure</span> <span class="p">*</span> <span class="mf">25.0</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="k">if</span> <span class="p">(</span><span class="nv">$saturation</span> <span class="o">-gt</span> <span class="mf">100</span><span class="p">)</span> <span class="p">{</span> <span class="nv">$saturation</span> <span class="p">=</span> <span class="mf">100</span> <span class="p">}</span>
</span></span></code></pre></div><p>A full KV (key-value) cache alone scores 50, and queue pressure of 2.0
scores the other 50. It&rsquo;s a heuristic, and it&rsquo;s deliberately one number.</p>
<p>Think of it as a dial that goes red when the engine is about to start
swapping requests to CPU memory. That&rsquo;s the moment latency falls off a
cliff. Warning is at 75, critical at 90.</p>
<p><img alt="vllm|perf|system_saturation_score over the last hour in VCF Operations" loading="lazy" src="/images/ui/o6-ops-vllm-saturation-score.jpg">
<em>The dial, as Ops draws it: one gauge climbing towards the warning line as KV-cache use and queue pressure rise together. (Test-mode data: see part 2.)</em></p>
<p>Also derived:</p>
<ul>
<li>prefix-cache hit rate, live and all-time;</li>
<li>average request size, from the dropped <code>http_request_size_bytes</code> <code>_sum</code>;</li>
<li>uptime in days;</li>
<li><code>is_up = 1</code> on every successful scrape, so the <em>absence</em> of the metric
is the alert.</li>
</ul>
<p>That last one does its most useful work by not turning up.</p>
<h2 id="what-gets-dropped-and-why">What gets dropped, and why</h2>
<p>The <code>$DropPrefixes</code> list is as important as anything computed:</p>
<table>
	<thead>
			<tr>
					<th>Dropped</th>
					<th>Why</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td><code>python_gc_*</code>, <code>python_info</code></td>
					<td>runtime noise; static text can&rsquo;t be graphed</td>
			</tr>
			<tr>
					<td><code>process_max_fds</code>, <code>vllm:cache_config_info</code>, <code>vllm:engine_sleep_state</code></td>
					<td>static configuration, not performance</td>
			</tr>
			<tr>
					<td><code>vllm:request_params_*</code></td>
					<td><strong>cardinality explosion</strong> — a series per distinct <code>max_tokens</code>/<code>n</code></td>
			</tr>
			<tr>
					<td><code>vllm:iteration_tokens_total</code></td>
					<td>redundant with tokens/s</td>
			</tr>
			<tr>
					<td><code>http_request_duration_highr_seconds</code></td>
					<td>&ldquo;high-resolution&rdquo; = hundreds of buckets; the standard one suffices</td>
			</tr>
			<tr>
					<td><code>http_request/response_size_bytes</code> buckets</td>
					<td>dropped, but <code>_sum</code> harvested for an average</td>
			</tr>
			<tr>
					<td><code>*_created</code></td>
					<td>bucket-initialisation timestamps; pure noise</td>
			</tr>
	</tbody>
</table>
<p>Everything that survives is truncated to two decimals before the push.
It&rsquo;s a small mercy for the TSDB.</p>
<h2 id="the-key-hierarchy">The key hierarchy</h2>
<p>Ops shows metrics as a tree, so the names are designed for browsing:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vllm|system|is_up
</span></span><span class="line"><span class="cl">vllm|throughput|total_tokens_per_sec
</span></span><span class="line"><span class="cl">vllm|perf|live_avg_ttft_ms
</span></span><span class="line"><span class="cl">vllm|perf|ttft|p95_ms
</span></span><span class="line"><span class="cl">vllm|queue|pressure_ratio
</span></span><span class="line"><span class="cl">vllm|memory|kv_cache_pct
</span></span><span class="line"><span class="cl">vllm|cache|live_prefix_hit_rate_pct
</span></span><span class="line"><span class="cl">vllm|process|rss_memory_gb
</span></span></code></pre></div><p><code>vllm | category | metric [| percentile]</code>. An operator who has never seen
vLLM can find &ldquo;time to first token, 95th percentile&rdquo; without a manual.</p>
<p><img alt="The key hierarchy as it lands in VCF Operations" loading="lazy" src="/images/ui/o5-ops-vllm-metric-tree.jpg">
<em><code>vllm | category | metric</code> in the Ops metric picker — see <a href="/series/llm-ops-on-vcf/">part 2</a> for how this was captured.</em></p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Organisations putting language models into service soon discover that &ldquo;is
it up?&rdquo; isn&rsquo;t the question. The real questions are these. How long are
users waiting for the first word? Is the service about to run out of
memory? Do we need another GPU before Friday?</p>
<p>This pipeline answers them inside the same VCF Operations console the
infrastructure team already lives in. So AI services get the same capacity
planning, alerting and dashboards as everything else. There&rsquo;s no second
monitoring stack, and no new team to staff it.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li><strong>Never push raw histogram buckets</strong> into a gauge-oriented TSDB.
Interpolate P50/P95/P99 client-side and push four gauges.</li>
<li>Counters need <strong>state</strong>: keep the last sample, compute deltas, and drop
the baseline after a long gap (300 s) rather than dilute the rate.</li>
<li><code>Δsum/Δcount</code> gives you <em>live</em> mean latency, which is far more useful
than the all-time mean.</li>
<li>One derived <strong>saturation score</strong> beats six raw gauges at the top of a
dashboard. Make it explainable (KV% × 0.5 + pressure × 25).</li>
<li>The drop-list is a design artefact, not housekeeping. <code>request_params_*</code>
alone can double your series count.</li>
<li>Name for browsing: <code>product | category | metric</code>.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/configuring-super-metrics/super-metrics-tab/super-metric-functions-and-operators.html">Super Metric Functions and Operators</a>: the functions an Ops formula can use; none of them is a percentile</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/administration-sdks-cli-and-tools/understanding-the-vr-ops-api/using-the-api-with-vrealize-operations-manager.html">Using the API with VCF Operations</a>: the REST API the pipeline pushes through, and the Swagger reference on the appliance</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/administration-sdks-cli-and-tools/understanding-the-vr-ops-api/getting-started-with-the-api/generate-a-list-of-all-metrics-for-the-object.html">Generate a List of All Metrics for the Object</a>: reading an object&rsquo;s stat keys back, grouped with <code>|</code> as in <code>mem|host_workload</code></li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/configuring-alerts-and-actions/symptom-definitions.html">Symptom Definitions in VCF Operations</a>: metric symptoms at Warning and Critical levels, as for the saturation score</li>
</ul>
<p><em>Part 2: <a href="/series/llm-ops-on-vcf/">running it as a service, monitoring the monitor, and the four
alerts</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Config shown is sanitised; the script&rsquo;s
test mode reads a local metrics file instead of a live endpoint.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Telegraf on VKS: the dependency that isn&#39;t in the README</title>
      <link>https://thenestedlab.com/posts/telegraf-vks-management-proxy/</link>
      <pubDate>Wed, 23 Sep 2026 06:10:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/telegraf-vks-management-proxy/</guid>
      <description>On VCF 9.1, the supported route installs Telegraf in new VKS clusters for you. Underneath sits a dependency the README never mentions, two secrets from the Supervisor Management Proxy, and without them every Telegraf pod is stuck.</description>
      <content:encoded><![CDATA[<p>The <a href="/posts/telegraf-windows-2025/">Windows half of this series</a> was about
an agent that works on an OS the vendor hasn&rsquo;t listed. This one is the
opposite: a package on a fully supported platform that does nothing,
silently, because of a dependency its own documentation doesn&rsquo;t mention.</p>
<p>The supported route comes first. The rest is the path underneath, as it
ran on f06 (VCF 9.1).</p>
<h2 id="the-symptom">The symptom</h2>
<p>Install the Telegraf package on a VKS cluster with the metric proxy flag
set (<code>isMetricProxyConfigured: true</code>, which you want when metrics go to
VCF Operations). Every Telegraf pod then stays in <code>ContainerCreating</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">NAME           PACKAGE NAME                         PACKAGE VERSION         DESCRIPTION                                                            AGE   PAUSED
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">telegraf       telegraf.kubernetes.vmware.com       1.34.4+vmware.2-vks.1   Reconcile failed: Error (see .status.usefulErrorMessage for details)   15m
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">kapp: Error: waiting on reconcile deployment/telegraf (apps/v1) namespace: tanzu-system-telegraf:
</span></span><span class="line"><span class="cl">  Finished waiting unsuccessfully:
</span></span><span class="line"><span class="cl">    Deployment is not progressing:
</span></span><span class="line"><span class="cl">      ProgressDeadlineExceeded, message:
</span></span><span class="line"><span class="cl">        ReplicaSet &#34;telegraf-7cd76f96&#34; has timed out progressing.
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">NAME                      READY   STATUS              RESTARTS   AGE
</span></span><span class="line"><span class="cl">telegraf-26ppl            0/1     ContainerCreating   0          15m
</span></span><span class="line"><span class="cl">telegraf-7cd76f96-wbthp   0/1     ContainerCreating   0          15m
</span></span><span class="line"><span class="cl">telegraf-g4qlm            0/1     ContainerCreating   0          15m
</span></span><span class="line"><span class="cl">telegraf-h2pcq            0/1     ContainerCreating   0          15m
</span></span></code></pre></div><p>Describe one, and the reason is <code>FailedMount</code>: two secrets don&rsquo;t exist.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">  Warning  FailedMount  5m14s (x13 over 15m)  kubelet            MountVolume.SetUp failed for volume &#34;metrics-proxy-tls-config&#34; : secret &#34;metrics-proxy-tls-config&#34; not found
</span></span><span class="line"><span class="cl">  Warning  FailedMount  3m12s (x14 over 15m)  kubelet            MountVolume.SetUp failed for volume &#34;telegraf-configs&#34; : secret &#34;metrics-proxy-http-config&#34; not found
</span></span></code></pre></div><p>Nothing creates them. The package doesn&rsquo;t. The cluster doesn&rsquo;t. The
PackageInstall reports <code>ReconcileFailed</code> and backs off, and there it
stays, waiting for two secrets that aren&rsquo;t coming.</p>
<h2 id="the-supported-route-first">The supported route first</h2>
<p>On VCF 9.1, VKS add-on management installs Telegraf and Prometheus for
you. Broadcom&rsquo;s page on enabling monitoring for VKS clusters (linked at
the end) puts it plainly: &ldquo;After the prerequisites are met, any new VKS
cluster is automatically monitored.&rdquo; It speaks only of new clusters. The
prerequisites, and how to check each:</p>
<ol>
<li><strong>VCF Operations monitors the vCenter and its Supervisor</strong>:
<em>Enable vSphere Supervisor Collection</em> in the vCenter account&rsquo;s
Advanced Settings, true by default. Look for a vSphere Supervisor
adapter instance and VKS Cluster objects in the inventory.</li>
<li><strong>The Metrics Aggregator runs on the Supervisor.</strong> VCF Automation
installs it by default. Check the Supervisor&rsquo;s services in vCenter.</li>
<li><strong>The add-on repository offers Telegraf and Prometheus</strong>:
<code>AddonRepository</code>, <code>Addon</code> and <code>AddonConfigDefinition</code> resources on
the Supervisor.</li>
</ol>
<p>On 9.0.1 and later, Tom Fojta&rsquo;s <a href="https://fojta.wordpress.com/2026/02/03/monitoring-vks-cluster-in-vcf-automation/">Monitoring VKS Cluster in VCF
Automation</a>
installs the same two packages as add-on tiles in VCF Automation, once
the provider enables the Supervisor Management Proxy. His notes include a
Telegraf failure over an existing <code>metrics-proxy-tls-config</code> secret: the
mirror image of the one above.</p>
<p>f06 took neither route on 23 August. It ran 9.1, but its VCF Automation
arrived the next day. Ops had a vSphere Supervisor adapter instance, the
Metrics Aggregator wasn&rsquo;t registered, and nobody looked for add-ons. By
nobody, I mean me. The lab&rsquo;s catalog item installed the standard package
itself, following Christian Ferber&rsquo;s <a href="https://vrealize.it/2025/10/24/monitoring-vks-clusters-in-vcf-operations/">vrealize.it
write-up</a>
for 9.0.1. That path is the rest of this post.</p>
<p>Two days later, with VCF Automation deployed and the Supervisor rebuilt,
the automatic route left a footprint. A cluster created through VCF
Automation&rsquo;s API had Telegraf and Prometheus namespaces within two
minutes, none of it mine:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">NAME                     STATUS   ROLES           AGE    VERSION
</span></span><span class="line"><span class="cl">vks-demo01-jn7xr-2gkq9   Ready    control-plane   111s   v1.35.5+vmware.1
</span></span><span class="line"><span class="cl">--- namespaces ---
</span></span><span class="line"><span class="cl">NAME                                 STATUS   AGE
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">tanzu-system-monitoring              Active   18s
</span></span><span class="line"><span class="cl">tanzu-system-telegraf                Active   11s
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><p>The rebuilt demo01 had them too, before the catalog item&rsquo;s package stage
ran. So the item&rsquo;s PackageInstalls collided with another kapp app:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">=== telegraf ===
</span></span><span class="line"><span class="cl">kapp: Error: Ownership errors: - Resource &#39;clusterrole/telegraf-kubelet-metric-access (rbac.authorization.k8s.io/v1) cluster&#39; is already associated with a different label &#39;kapp.k14s.io/app=1787648814041772277&#39; - Resource &#39;serviceaccount/telegraf-sa (v1) namespace: tanzu-system-telegraf&#39; is already a
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><p>I didn&rsquo;t inspect what installed them (the Supervisor was serving the
<code>addons.kubernetes.vmware.com</code> API by then). It&rsquo;s a footprint, not a
test. Check for Telegraf before you install it.</p>
<h2 id="what-sits-underneath">What sits underneath</h2>
<p>The two secrets come from the <strong>Supervisor Management Proxy</strong>, a
Supervisor Service on the Supervisor, not in the guest. It runs envoy
behind a load balancer on port 10093.</p>
<p>Each guest cluster gets a headless Service pointing at it
(<code>supervisor-management-proxy</code> in <code>default</code>), plus the two secrets in
<code>kube-system</code>, each with a <code>SecretExport</code>. The package&rsquo;s <code>SecretImport</code>s
copy them into <code>tanzu-system-telegraf</code>, and Telegraf posts to
<code>https://supervisor-management-proxy.default.svc.&lt;serviceDomain&gt;:10093/arc/tkgs/metric</code>.</p>
<p><img alt="Diagram: the secrets, the headless Service and the proxy&rsquo;s load balancer on f06, with the onward path to VCF Operations dashed" loading="lazy" src="/images/telegraf-vks-management-proxy-diagram.svg">
<em>Solid: what f06 showed. Dashed: the 9.1 design, in which the Supervisor also issues Telegraf&rsquo;s certificates and cert-manager rotates them.</em></p>
<p>The manual path needs a Supervisor with a load balancer, registry access
to <code>projects.packages.broadcom.com</code>, vCenter rights to register
Supervisor Services, and cluster-admin in the guest. The cluster is
<code>demo01</code>: one control-plane node and two workers.</p>
<h2 id="1-install-the-supervisor-management-proxy">1. Install the Supervisor Management Proxy</h2>
<p>The lab&rsquo;s catalog item that deploys a Supervisor has an
<code>installMgmtProxy</code> tick box: register the service, then install it.
Every attempt was refused. The definition shows why:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">data.packaging.carvel.dev/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">Package</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">supervisor-management-proxy.vmware.com.0.4.1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">annotations</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">appplatform.vmware.com/use-system-vpc</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;true&#34;</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">appplatform.vmware.com/required_capability</span><span class="p">:</span><span class="w"> </span><span class="l">Load_Balancer_Supported</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">appplatform.vmware.com/vcenter-version-constraints</span><span class="p">:</span><span class="w"> </span><span class="s1">&#39;&gt;=9.0.0&#39;</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">refName</span><span class="p">:</span><span class="w"> </span><span class="l">supervisor-management-proxy.vmware.com</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">version</span><span class="p">:</span><span class="w"> </span><span class="m">0.4.1</span><span class="w">
</span></span></span></code></pre></div><p>The proxy asks for the <strong>system VPC</strong>. vCenter only places services there
that it can verify as published by Broadcom:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">=== install on supervisor domain-c9 ===
</span></span><span class="line"><span class="cl">  install failed: 500
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">      &#34;default_message&#34;: &#34;Service supervisor-management-proxy.vmware.com version 0.4.1 is not a trusted service published by Broadcom and cannot be placed on the system VPC.&#34;,
</span></span><span class="line"><span class="cl">      &#34;id&#34;: &#34;vcenter.wcp.appplatform.signature_verification.system_vpc&#34;
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><p>Registered through the API as Carvel YAML, it never passed, even
byte-identical to Broadcom&rsquo;s download. The other Carvel services here
carry no such annotation and installed fine:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">harbor                       2.14.2+vmware.2-vks.2    no system-vpc annotation  content_type=CARVEL_APPS_YAML
</span></span><span class="line"><span class="cl">cci-ns                       9.1.0-embedded+739b5075  no system-vpc annotation  content_type=CARVEL_APPS_YAML
</span></span><span class="line"><span class="cl">velero                       1.9.0-embedded+25369333  no system-vpc annotation  content_type=CARVEL_APPS_YAML
</span></span></code></pre></div><p>That check is doing its job, and I won&rsquo;t show how the lab got past it.
On a platform you care about, install the proxy the way Broadcom&rsquo;s
proxy page (linked at the end) describes, so vCenter can verify what it
places on the system VPC. However it goes in, <strong>check it</strong> on the
Supervisor:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-gdscript3" data-lang="gdscript3"><span class="line"><span class="cl">  <span class="p">[</span><span class="mi">30</span><span class="n">s</span><span class="p">]</span> <span class="n">service</span><span class="o">=</span><span class="n">CONFIGURED</span>  <span class="n">pod</span><span class="o">=</span><span class="n">Running</span> <span class="n">ready</span><span class="o">=</span><span class="mi">1</span><span class="o">/</span><span class="mi">1</span>
</span></span><span class="line"><span class="cl"><span class="o">---</span> <span class="n">LB</span> <span class="n">services</span> <span class="ow">in</span> <span class="n">proxy</span> <span class="n">ns</span> <span class="o">---</span>
</span></span><span class="line"><span class="cl">  <span class="n">workload</span><span class="o">-</span><span class="n">metrics</span><span class="o">-</span><span class="n">loadbalancer</span><span class="p">:</span> <span class="n">type</span><span class="o">=</span><span class="n">LoadBalancer</span> <span class="n">ip</span><span class="o">=</span><span class="mf">10.26</span><span class="o">.</span><span class="mf">20.21</span> <span class="n">ports</span><span class="o">=</span><span class="mi">10093</span>
</span></span><span class="line"><span class="cl"><span class="o">...</span>
</span></span></code></pre></div><p>Then check it in vCenter. The install record holds only the namespace the
platform chose (the base64 is <code>namespace: svc-supervisor-management-proxy-acqyd</code>):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">--- current install record ---
</span></span><span class="line"><span class="cl">{
</span></span><span class="line"><span class="cl">  &#34;desired_version&#34;: &#34;0.4.1&#34;,
</span></span><span class="line"><span class="cl">  &#34;current_version&#34;: &#34;0.4.1&#34;,
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">        &#34;default_message&#34;: &#34;Reason: ReconcileSucceeded. Message: Reconcile succeeded.&#34;,
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">  &#34;yaml_service_config&#34;: &#34;bmFtZXNwYWNlOiBzdmMtc3VwZXJ2aXNvci1tYW5hZ2VtZW50LXByb3h5LWFjcXlkCg==&#34;,
</span></span><span class="line"><span class="cl">  &#34;service_namespace&#34;: &#34;svc-supervisor-management-proxy-acqyd&#34;,
</span></span><span class="line"><span class="cl">  &#34;display_name&#34;: &#34;Supervisor Management Proxy&#34;,
</span></span><span class="line"><span class="cl">  &#34;config_status&#34;: &#34;CONFIGURED&#34;
</span></span><span class="line"><span class="cl">}
</span></span></code></pre></div><h2 id="2-create-the-cluster-with-servicedomain">2. Create the cluster with serviceDomain</h2>
<p><code>clusterNetwork.serviceDomain</code> can&rsquo;t be added later (step 7). The catalog
item now sets it on every demo cluster. Here&rsquo;s demo01 as rebuilt:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">{
</span></span><span class="line"><span class="cl">  &#34;clusterNetwork&#34;: {
</span></span><span class="line"><span class="cl">    &#34;pods&#34;: {
</span></span><span class="line"><span class="cl">      &#34;cidrBlocks&#34;: [
</span></span><span class="line"><span class="cl">        &#34;192.168.0.0/16&#34;
</span></span><span class="line"><span class="cl">      ]
</span></span><span class="line"><span class="cl">    },
</span></span><span class="line"><span class="cl">    &#34;serviceDomain&#34;: &#34;cluster.local&#34;,
</span></span><span class="line"><span class="cl">    &#34;services&#34;: {
</span></span><span class="line"><span class="cl">      &#34;cidrBlocks&#34;: [
</span></span><span class="line"><span class="cl">        &#34;10.96.0.0/12&#34;
</span></span><span class="line"><span class="cl">      ]
</span></span><span class="line"><span class="cl">    }
</span></span><span class="line"><span class="cl">  },
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><h2 id="3-add-the-package-repository">3. Add the package repository</h2>
<p>Telegraf ships in Broadcom&rsquo;s VKS standard packages. Register the
repository in the guest&rsquo;s <code>tkg-system</code> namespace:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">packaging.carvel.dev/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">PackageRepository</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">broadcom-standard-repo</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">namespace</span><span class="p">:</span><span class="w"> </span><span class="l">tkg-system</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">fetch</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">imgpkgBundle</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">image</span><span class="p">:</span><span class="w"> </span><span class="l">projects.packages.broadcom.com/vsphere/supervisor/packages/2025.8.19/vks-standard-packages:v2025.8.19</span><span class="w">
</span></span></span></code></pre></div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">packagerepository.packaging.carvel.dev/broadcom-standard-repo created
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">NAME                     AGE   DESCRIPTION           PAUSED
</span></span><span class="line"><span class="cl">broadcom-standard-repo   60s   Reconcile succeeded
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">=== packages available ===
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">NAME                                                                  PACKAGEMETADATA NAME                             VERSION                  AGE
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">telegraf.kubernetes.vmware.com.1.34.4+vmware.2-vks.1                  telegraf.kubernetes.vmware.com                   1.34.4+vmware.2-vks.1    54s
</span></span><span class="line"><span class="cl">telegraf.tanzu.vmware.com.1.32.1+vmware.1-tkg.1                       telegraf.tanzu.vmware.com                        1.32.1+vmware.1-tkg.1    54s
</span></span></code></pre></div><h2 id="4-install-telegraf">4. Install Telegraf</h2>
<p>Two values matter, from the package&rsquo;s own schema:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-go" data-lang="go"><span class="line"><span class="cl"><span class="o">---</span><span class="w"> </span><span class="nx">telegraf</span><span class="w"> </span><span class="kn">package</span><span class="w"> </span><span class="nx">valuesSchema</span><span class="w"> </span><span class="nx">keys</span><span class="w"> </span><span class="o">---</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="o">...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nf">domainName</span><span class="w">  </span><span class="p">(</span><span class="k">default</span><span class="p">:</span><span class="w"> </span><span class="nx">cluster</span><span class="p">.</span><span class="nx">local</span><span class="p">)</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="o">...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nf">isMetricProxyConfigured</span><span class="w">  </span><span class="p">(</span><span class="k">default</span><span class="p">:</span><span class="w"> </span><span class="nx">False</span><span class="p">)</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nf">namespace</span><span class="w">  </span><span class="p">(</span><span class="k">default</span><span class="p">:</span><span class="w"> </span><span class="nx">tanzu</span><span class="o">-</span><span class="nx">system</span><span class="o">-</span><span class="nx">telegraf</span><span class="p">)</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="o">...</span><span class="w">
</span></span></span></code></pre></div><p>The catalog item installs it the Carvel way: a values Secret, then a
PackageInstall, in a <code>package-installs</code> namespace whose service account
has cluster-admin.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-js" data-lang="js"><span class="line"><span class="cl">    <span class="kd">function</span> <span class="nx">pkgSecret</span><span class="p">(</span><span class="nx">name</span><span class="p">,</span> <span class="nx">valuesYml</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">        <span class="nx">k8sApply</span><span class="p">(</span><span class="nx">gk</span><span class="p">,</span> <span class="nx">gTok</span><span class="p">,</span> <span class="s2">&#34;/api/v1/namespaces/package-installs/secrets&#34;</span><span class="p">,</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">            <span class="nx">apiVersion</span><span class="o">:</span> <span class="s2">&#34;v1&#34;</span><span class="p">,</span> <span class="nx">kind</span><span class="o">:</span> <span class="s2">&#34;Secret&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="nx">metadata</span><span class="o">:</span> <span class="p">{</span> <span class="nx">name</span><span class="o">:</span> <span class="nx">name</span><span class="p">,</span> <span class="nx">namespace</span><span class="o">:</span> <span class="s2">&#34;package-installs&#34;</span> <span class="p">},</span>
</span></span><span class="line"><span class="cl">            <span class="nx">stringData</span><span class="o">:</span> <span class="p">{</span> <span class="s2">&#34;values.yml&#34;</span><span class="o">:</span> <span class="nx">valuesYml</span> <span class="p">}</span> <span class="p">},</span> <span class="s2">&#34;secret &#34;</span> <span class="o">+</span> <span class="nx">name</span><span class="p">);</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl">    <span class="kd">function</span> <span class="nx">pkgInstall</span><span class="p">(</span><span class="nx">name</span><span class="p">,</span> <span class="nx">refName</span><span class="p">,</span> <span class="nx">ver</span><span class="p">,</span> <span class="nx">valuesSecret</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">        <span class="p">...</span>
</span></span><span class="line"><span class="cl">        <span class="kd">var</span> <span class="nx">pi</span> <span class="o">=</span> <span class="p">{</span> <span class="nx">apiVersion</span><span class="o">:</span> <span class="s2">&#34;packaging.carvel.dev/v1alpha1&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="nx">kind</span><span class="o">:</span> <span class="s2">&#34;PackageInstall&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="nx">metadata</span><span class="o">:</span> <span class="p">{</span> <span class="nx">name</span><span class="o">:</span> <span class="nx">name</span><span class="p">,</span> <span class="nx">namespace</span><span class="o">:</span> <span class="s2">&#34;package-installs&#34;</span> <span class="p">},</span>
</span></span><span class="line"><span class="cl">            <span class="nx">spec</span><span class="o">:</span> <span class="p">{</span> <span class="nx">serviceAccountName</span><span class="o">:</span> <span class="s2">&#34;pkgi-sa&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">                <span class="p">...</span>
</span></span><span class="line"><span class="cl">                <span class="nx">packageRef</span><span class="o">:</span> <span class="p">{</span> <span class="nx">refName</span><span class="o">:</span> <span class="nx">refName</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">                    <span class="nx">versionSelection</span><span class="o">:</span> <span class="p">{</span> <span class="nx">constraints</span><span class="o">:</span> <span class="nx">ver</span> <span class="p">}</span> <span class="p">}</span> <span class="p">}</span> <span class="p">};</span>
</span></span><span class="line"><span class="cl">        <span class="k">if</span> <span class="p">(</span><span class="nx">valuesSecret</span><span class="p">)</span> <span class="p">{</span> <span class="nx">pi</span><span class="p">.</span><span class="nx">spec</span><span class="p">.</span><span class="nx">values</span> <span class="o">=</span> <span class="p">[{</span> <span class="nx">secretRef</span><span class="o">:</span> <span class="p">{</span> <span class="nx">name</span><span class="o">:</span> <span class="nx">valuesSecret</span> <span class="p">}</span> <span class="p">}];</span> <span class="p">}</span>
</span></span><span class="line"><span class="cl">        <span class="p">...</span>
</span></span><span class="line"><span class="cl">    <span class="p">}</span>
</span></span><span class="line"><span class="cl">    <span class="p">...</span>
</span></span><span class="line"><span class="cl">    <span class="nx">pkgSecret</span><span class="p">(</span><span class="s2">&#34;telegraf-values&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;domainName: cluster.local\n&#34;</span> <span class="o">+</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;isMetricProxyConfigured: true\n&#34;</span><span class="p">);</span>
</span></span><span class="line"><span class="cl">    <span class="nx">pkgInstall</span><span class="p">(</span><span class="s2">&#34;telegraf&#34;</span><span class="p">,</span> <span class="s2">&#34;telegraf.kubernetes.vmware.com&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">        <span class="s2">&#34;1.34.4+vmware.2-vks.1&#34;</span><span class="p">,</span> <span class="s2">&#34;telegraf-values&#34;</span><span class="p">);</span>
</span></span></code></pre></div><p>The pods land in <code>tanzu-system-telegraf</code>: one per node from a DaemonSet,
plus one Deployment pod.</p>
<h2 id="5-watch-the-secrets-arrive">5. Watch the secrets arrive</h2>
<p>This is the dependency. Within two minutes of the proxy reporting
<code>CONFIGURED</code>, the secrets and their <code>SecretExport</code>s were in demo01&rsquo;s
<code>kube-system</code>. The <code>SecretImport</code>s had copied them across:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-gdscript3" data-lang="gdscript3"><span class="line"><span class="cl"><span class="o">---</span> <span class="n">secrets</span> <span class="o">+</span> <span class="n">secretimports</span> <span class="ow">in</span> <span class="n">tanzu</span><span class="o">-</span><span class="n">system</span><span class="o">-</span><span class="n">telegraf</span> <span class="o">---</span>
</span></span><span class="line"><span class="cl">  <span class="n">secrets</span><span class="p">:</span> <span class="n">metrics</span><span class="o">-</span><span class="n">proxy</span><span class="o">-</span><span class="n">http</span><span class="o">-</span><span class="n">config</span><span class="p">,</span> <span class="n">metrics</span><span class="o">-</span><span class="n">proxy</span><span class="o">-</span><span class="n">tls</span><span class="o">-</span><span class="n">config</span><span class="p">,</span> <span class="n">telegraf</span><span class="o">-</span><span class="n">registry</span><span class="o">-</span><span class="n">creds</span>
</span></span><span class="line"><span class="cl">  <span class="n">secretimport</span><span class="p">:</span> <span class="n">metrics</span><span class="o">-</span><span class="n">proxy</span><span class="o">-</span><span class="n">http</span><span class="o">-</span><span class="n">config</span> <span class="o">-&gt;</span> <span class="n">ReconcileSucceeded</span><span class="o">=</span><span class="n">True</span>
</span></span><span class="line"><span class="cl">  <span class="n">secretimport</span><span class="p">:</span> <span class="n">metrics</span><span class="o">-</span><span class="n">proxy</span><span class="o">-</span><span class="n">tls</span><span class="o">-</span><span class="n">config</span> <span class="o">-&gt;</span> <span class="n">ReconcileSucceeded</span><span class="o">=</span><span class="n">True</span>
</span></span><span class="line"><span class="cl">  <span class="n">kube</span><span class="o">-</span><span class="n">system</span> <span class="n">secretexport</span><span class="p">:</span> <span class="n">metrics</span><span class="o">-</span><span class="n">proxy</span><span class="o">-</span><span class="n">http</span><span class="o">-</span><span class="n">config</span> <span class="n">toNamespaces</span><span class="o">=</span>
</span></span><span class="line"><span class="cl">  <span class="n">kube</span><span class="o">-</span><span class="n">system</span> <span class="n">secretexport</span><span class="p">:</span> <span class="n">metrics</span><span class="o">-</span><span class="n">proxy</span><span class="o">-</span><span class="n">tls</span><span class="o">-</span><span class="n">config</span> <span class="n">toNamespaces</span><span class="o">=</span>
</span></span><span class="line"><span class="cl"><span class="o">...</span>
</span></span></code></pre></div><p>The pods came up on their own:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">  [30s] telegraf pods ready: 0/4  pkgi: ReconcileFailed=True
</span></span><span class="line"><span class="cl">  [60s] telegraf pods ready: 2/4  pkgi: ReconcileFailed=True
</span></span><span class="line"><span class="cl">  [90s] telegraf pods ready: 4/4  pkgi: ReconcileFailed=True
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">  [480s] telegraf pods ready: 4/4  pkgi: ReconcileFailed=True
</span></span></code></pre></div><p>After an hour and forty minutes of <code>FailedMount</code>, all four were Running
four minutes after the proxy went <code>CONFIGURED</code>. I&rsquo;d quite like that hour
and forty minutes back.</p>
<h2 id="6-clear-the-stale-backoff">6. Clear the stale backoff</h2>
<p>The PackageInstall stayed in its <code>ReconcileFailed</code> backoff with every pod
Running. Bumping an annotation forces a fresh reconcile:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="nv">$K</span><span class="p">=</span><span class="vm">@</span><span class="p">{</span><span class="n">Authorization</span><span class="p">=</span><span class="s2">&#34;Bearer </span><span class="p">$(</span><span class="nv">$lg</span><span class="p">.</span><span class="n">session_id</span><span class="p">)</span><span class="s2">&#34;</span><span class="p">;</span> <span class="s1">&#39;Content-Type&#39;</span><span class="p">=</span><span class="s1">&#39;application/merge-patch+json&#39;</span><span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">...</span>
</span></span><span class="line"><span class="cl"><span class="nv">$u</span><span class="p">=</span><span class="s2">&#34;https://</span><span class="nv">${gk}</span><span class="s2">:6443/apis/packaging.carvel.dev/v1alpha1/namespaces/package-installs/packageinstalls/telegraf&#34;</span>
</span></span><span class="line"><span class="cl"><span class="p">...</span>
</span></span><span class="line"><span class="cl"><span class="c"># force re-reconcile by bumping an annotation</span>
</span></span><span class="line"><span class="cl"><span class="nv">$patch</span><span class="p">=</span><span class="s1">&#39;{&#34;metadata&#34;:{&#34;annotations&#34;:{&#34;lab.kick&#34;:&#34;&#39;</span><span class="p">+(</span><span class="nb">Get-Random</span><span class="p">)+</span><span class="s1">&#39;&#34;}}}&#39;</span>
</span></span><span class="line"><span class="cl"><span class="nb">Invoke-RestMethod</span> <span class="n">-Method</span> <span class="n">Patch</span> <span class="n">-Uri</span> <span class="nv">$u</span> <span class="n">-Headers</span> <span class="nv">$K</span> <span class="n">-Body</span> <span class="nv">$patch</span> <span class="n">-SkipCertificateCheck</span> <span class="p">|</span> <span class="nb">Out-Null</span>
</span></span></code></pre></div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">kicked; polling...
</span></span><span class="line"><span class="cl">  [30s] ReconcileSucceeded=True Reconcile succeeded
</span></span></code></pre></div><h2 id="7-check-the-output-url">7. Check the output URL</h2>
<p>Telegraf still wasn&rsquo;t delivering. Its log, one error a minute:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">  2026-08-23T19:17:56Z I! [agent] Config: Interval:5m0s, Quiet:false, Hostname:&#34;telegraf-7cd76f96-wbthp&#34;, Flush Interval:1m0s
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">  2026-08-23T19:20:03Z E! [agent] Error writing to outputs.http: Post &#34;https://supervisor-management-proxy.default.svc.:10093/arc/tkgs/metric&#34;: dial tcp: lookup supervisor-management-proxy.default.svc. on 10.96.0.10:53: no such host
</span></span><span class="line"><span class="cl">  2026-08-23T19:20:56Z E! [agent] Error writing to outputs.http: Post &#34;https://supervisor-management-proxy.default.svc.:10093/arc/tkgs/metric&#34;: dial tcp: lookup supervisor-management-proxy.default.svc. on 10.96.0.10:53: no such host
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><p>Note the trailing dot, and nothing after <code>svc</code>. The name is that headless
Service:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">--- guest &#39;default&#39; ns services ---
</span></span><span class="line"><span class="cl">  kubernetes: type=ClusterIP extName= clusterIP=10.96.0.1
</span></span><span class="line"><span class="cl">  supervisor: type=ClusterIP extName= clusterIP=None
</span></span><span class="line"><span class="cl">  supervisor-management-proxy: type=ClusterIP extName= clusterIP=None
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">--- endpoints of supervisor-management-proxy (default ns) ---
</span></span><span class="line"><span class="cl">  ips=10.26.20.21 ports=10093
</span></span></code></pre></div><p>The domain part follows the cluster&rsquo;s <code>serviceDomain</code>, which demo01
lacked. The package&rsquo;s own <code>domainName</code> was <code>cluster.local</code> all along, and
made no difference. Adding the field later was refused:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">patch rejected:  { &#34;kind&#34;: &#34;Status&#34;, &#34;apiVersion&#34;: &#34;v1&#34;, &#34;metadata&#34;: {}, &#34;status&#34;: &#34;Failure&#34;, &#34;message&#34;: &#34;admission webhook \u0022capi.validating.tanzukubernetescluster.run.tanzu.vmware.com\u0022 denied the request: spec.clusterNetwork is immutable and cannot be upd
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><p>Two fixes, both applied:</p>
<ul>
<li><strong>Every new cluster:</strong> <code>serviceDomain: cluster.local</code> in the spec
(step 2). Every add-on that builds a service URL assumes it&rsquo;s there.</li>
<li><strong>Live cluster:</strong> a CoreDNS <code>rewrite</code> rule in the guest&rsquo;s <code>coredns</code>
ConfigMap, mapping the domainless name onto the real one. Ugly,
effective, and documented in the cluster&rsquo;s notes. The <code>reload</code> plugin
picked it up:</li>
</ul>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">--- Corefile now ---
</span></span><span class="line"><span class="cl">.:53 {
</span></span><span class="line"><span class="cl">    errors
</span></span><span class="line"><span class="cl">    health {
</span></span><span class="line"><span class="cl">       lameduck 5s
</span></span><span class="line"><span class="cl">    }
</span></span><span class="line"><span class="cl">    ready
</span></span><span class="line"><span class="cl">    rewrite name supervisor-management-proxy.default.svc supervisor-management-proxy.default.svc.cluster.local
</span></span><span class="line"><span class="cl">    kubernetes cluster.local in-addr.arpa ip6.arpa {
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><h2 id="8-leave-the-proxys-values-alone">8. Leave the proxy&rsquo;s values alone</h2>
<p>The error moved on, which counts as progress around here. The name resolves now,
and nothing answers:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">checked at (UTC): 19:35:38
</span></span><span class="line"><span class="cl">telegraf-26ppl: STILL FAILING (4 errors)
</span></span><span class="line"><span class="cl">  last: 2026-08-23T19:35:07Z E! [agent] Error writing to outputs.http: Post &#34;https://supervisor-management-proxy.default.svc.:10093/arc/tkgs/metric&#34;: context deadline exceeded (Client.Timeout exceeded while awaiting headers)
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><p>In 9.1 terms, a prerequisite was missing. The Metrics Aggregator, which
terminates the cluster&rsquo;s TLS in that design, didn&rsquo;t exist on f06 yet. I
tried pointing the proxy at Ops by hand, with a value from its
definition:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="w">        </span><span class="nt">metricsHTTPRemoteEndpoint</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">title</span><span class="p">:</span><span class="w"> </span><span class="l">HTTP remote endpoint configuration for pushing workload metrics</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="l">...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">properties</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">            </span><span class="nt">host</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="l">...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">            </span><span class="nt">port</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="l">...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="nt">default</span><span class="p">:</span><span class="w"> </span><span class="m">443</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="l">...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">            </span><span class="nt">tlsClientSecretName</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="l">string</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="nt">description</span><span class="p">:</span><span class="w"> </span><span class="l">Name of the secret which contains the TLS configuration required to push metrics to remote endpoint.</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="nt">default</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;&#34;</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">            </span><span class="nt">tlsClientSecretNamespace</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="l">...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">              </span><span class="nt">default</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;vmware-system-monitoring&#34;</span><span class="w">
</span></span></span></code></pre></div><p>A <code>PUT</code> on the install record (<code>PATCH</code> returns 404) reconciled:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="nv">$vals</span><span class="p">=</span><span class="s2">&#34;namespace: svc-supervisor-management-proxy-acqyd</span><span class="se">`n</span><span class="s2">metricsHTTPRemoteEndpoint:</span><span class="se">`n</span><span class="s2">  host: </span><span class="se">`&#34;</span><span class="s2">f06-flt-ops01.res.lab</span><span class="se">`&#34;`n</span><span class="s2">  port: 443</span><span class="se">`n</span><span class="s2">&#34;</span>
</span></span><span class="line"><span class="cl"><span class="nv">$b64</span><span class="p">=[</span><span class="no">Convert</span><span class="p">]::</span><span class="n">ToBase64String</span><span class="p">([</span><span class="no">Text.Encoding</span><span class="p">]::</span><span class="n">UTF8</span><span class="p">.</span><span class="py">GetBytes</span><span class="p">(</span><span class="nv">$vals</span><span class="p">))</span>
</span></span><span class="line"><span class="cl"><span class="nv">$body</span><span class="p">=</span><span class="vm">@</span><span class="p">{</span> <span class="n">version</span><span class="p">=</span><span class="s1">&#39;0.4.1&#39;</span><span class="p">;</span> <span class="n">yaml_service_config</span><span class="p">=</span><span class="nv">$b64</span> <span class="p">}</span> <span class="p">|</span> <span class="nb">ConvertTo-Json</span>
</span></span><span class="line"><span class="cl"><span class="p">...</span>
</span></span><span class="line"><span class="cl">  <span class="nv">$r</span><span class="p">=</span><span class="nb">Invoke-WebRequest</span> <span class="n">-Method</span> <span class="n">Put</span> <span class="n">-Uri</span> <span class="s2">&#34;https://</span><span class="nv">$vc</span><span class="s2">/api/vcenter/namespace-management/clusters/domain-c9/supervisor-services/</span><span class="nv">$svc</span><span class="s2">&#34;</span> <span class="n">-Headers</span> <span class="nv">$H</span> <span class="n">-ContentType</span> <span class="s1">&#39;application/json&#39;</span> <span class="n">-Body</span> <span class="nv">$body</span> <span class="n">-SkipCertificateCheck</span> <span class="n">-TimeoutSec</span> <span class="mf">120</span>
</span></span></code></pre></div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">PUT -&gt; 204
</span></span><span class="line"><span class="cl">  [30s] CONFIGURED
</span></span><span class="line"><span class="cl">  [60s] CONFIGURED
</span></span><span class="line"><span class="cl">  [90s] CONFIGURED
</span></span></code></pre></div><p>Telegraf timed out as before. Broadcom&rsquo;s proxy page is blunt: &ldquo;No
additional configuration values for the Supervisor Management Proxy
service are required in this case.&rdquo;</p>
<p>Set at install time, the value did harm. With no <code>tlsClientSecretName</code>,
the package renders a nameless <code>SecretImport</code>, and kapp refuses it:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">=== supervisor-management-proxy.vmware.com ===
</span></span><span class="line"><span class="cl">Reason: ReconcileFailed. Message: kapp: Error: Validation errors:
</span></span><span class="line"><span class="cl">- Expected &#39;metadata.name&#39; on resource &#39;secretimport/ (secretgen.carvel.dev/v1alpha1) namespace: svc-supervisor-management-proxy-htarg&#39; to be non-empty (stdin doc 5).
</span></span></code></pre></div><p>The next day, the Supervisor was rebuilt onto the lab&rsquo;s planned address
range. The catalog item installed the proxy and the Metrics Aggregator
with no values, and both went <code>CONFIGURED</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-gdscript3" data-lang="gdscript3"><span class="line"><span class="cl"><span class="o">---</span> <span class="n">ALL</span> <span class="n">LoadBalancer</span> <span class="n">VIPs</span> <span class="n">on</span> <span class="n">the</span> <span class="n">rebuilt</span> <span class="n">supervisor</span> <span class="o">---</span>
</span></span><span class="line"><span class="cl"><span class="o">...</span>
</span></span><span class="line"><span class="cl">  <span class="n">svc</span><span class="o">-</span><span class="n">metrics</span><span class="o">-</span><span class="n">aggregator</span><span class="o">-</span><span class="mi">5</span><span class="n">q0of</span>             <span class="n">workload</span><span class="o">-</span><span class="n">metrics</span><span class="o">-</span><span class="n">loadbalancer</span> <span class="mf">192.168</span><span class="o">.</span><span class="mf">144.7</span>
</span></span><span class="line"><span class="cl">  <span class="n">svc</span><span class="o">-</span><span class="n">supervisor</span><span class="o">-</span><span class="n">management</span><span class="o">-</span><span class="n">proxy</span><span class="o">-</span><span class="n">htarg</span>    <span class="n">workload</span><span class="o">-</span><span class="n">metrics</span><span class="o">-</span><span class="n">loadbalancer</span> <span class="mf">192.168</span><span class="o">.</span><span class="mf">144.6</span>
</span></span></code></pre></div><p>demo01 was recreated on that Supervisor with <code>serviceDomain</code> set. This
check read a Telegraf pod&rsquo;s last two minutes of log. The broken state had
logged an error a minute:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">=== observability packages ===
</span></span><span class="line"><span class="cl">  cert-manager   ReconcileSucceeded=True
</span></span><span class="line"><span class="cl">  contour        Reconciling=True
</span></span><span class="line"><span class="cl">  fluent-bit     ReconcileSucceeded=True
</span></span><span class="line"><span class="cl">  prometheus     ReconcileFailed=True
</span></span><span class="line"><span class="cl">  telegraf       ReconcileFailed=True
</span></span><span class="line"><span class="cl">=== telegraf delivery (last 2 min) ===
</span></span><span class="line"><span class="cl">  CLEAN - metrics flowing to mgmt proxy
</span></span></code></pre></div><p>&ldquo;Metrics flowing&rdquo; is my script&rsquo;s wording for no failed writes. And the
<code>telegraf</code> PackageInstall is the one kapp refused, so the clean Telegraf
wasn&rsquo;t mine.</p>
<h2 id="9-vcf-operations">9. VCF Operations</h2>
<p>Broadcom&rsquo;s consumption docs show the result in VCF Automation (Manage
and Govern &gt; Kubernetes Management &gt; Clusters) and in a Workload
Management Activated Cluster Summary tab in VCF Operations. The 9.0.1
write-up adds a switch on the VKS Cluster object, which the catalog item
still names at the end of each run:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">... | Ops manual step: set &#39;Pod And Container Monitoring Enabled&#39; on VKS Cluster demo01 in VCF Operations inventory; ...
</span></span></code></pre></div><p>This record has neither. The switch was never set, and my one capture
of a VKS Cluster object (vks-demo01, 1 September, in the <a href="/posts/vks-kubectl-vs-vcfa-all-apps/">two-ways
post</a>) counts 0 pods, deployments
and DaemonSets. The last hop I can show is Telegraf writing without
errors. Not the grand finale I&rsquo;d hoped for, but an honest one.</p>
<h2 id="clean-up-and-repeatability">Clean-up and repeatability</h2>
<ul>
<li><strong>One owner per add-on.</strong> Where add-on management installs Telegraf,
don&rsquo;t add your own PackageInstall. To run it yourself, first set the
cluster label <code>addons.kubernetes.vmware.com/automated-monitoring</code> to
<code>disabled</code>. The catalog item&rsquo;s answer to the ownership errors,
<code>--dangerous-override-ownership-of-existing-resources=true</code>, only hands
the objects to a second owner.</li>
<li><strong>Supervisor Service API on 9.1:</strong> service
<code>carvel_spec.version_spec.content</code>, version <code>carvel_spec.content</code>,
install <code>{supervisor_service, version}</code>, values
<code>PUT {version, yaml_service_config}</code>, removal
<code>PATCH …?action=deactivate</code>, then <code>DELETE</code>. New definition bytes need a
new service record.</li>
<li><strong>A definition outlives the Supervisor.</strong> What you register lives on
vCenter&rsquo;s service record, so a rebuilt Supervisor installs the same
definition again. Replacing it means deactivating and deleting that
record first.</li>
<li><strong>Services break later.</strong> In September, the proxy, the aggregator and
four other Supervisor Services went to <code>ERROR</code> when the platform
replaced their kapp service accounts. Recreating the old ones fixed all
six within two minutes.</li>
</ul>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>On VCF 9.1, Kubernetes metrics in VCF Operations are meant to be a
platform setting, not a job per cluster. Get three prerequisites right,
and new clusters arrive monitored.</p>
<p>The layer underneath still matters, because every failure here had the
same shape: a generic symptom (<code>FailedMount</code>, a name that won&rsquo;t resolve,
a timeout) caused by a decision made somewhere else.</p>
<p>For a customer, that means a Supervisor built for observability before
any cluster asks, a cluster baseline with the fields add-ons assume, and
one owner per add-on. That&rsquo;s exactly what a standard cluster class and a
Supervisor build checklist exist to encode.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>On 9.1, start with the supported route: Ops collecting the Supervisor,
the Metrics Aggregator and the add-on repository. Then check for
Telegraf before installing it.</li>
<li>On VKS, the Telegraf package <strong>hard-depends on the Supervisor
Management Proxy</strong> whenever the metric proxy flag is set. <code>FailedMount</code>
on two <code>metrics-proxy-*</code> secrets is the tell.</li>
<li>The dependency heals retroactively: install the proxy, bump the
PackageInstall annotation, wait.</li>
<li>Set <code>serviceDomain</code> at cluster create. It&rsquo;s immutable, and every add-on
that builds a service URL will assume it.</li>
<li>A name that resolves isn&rsquo;t a metric delivered. Read the output log for
an error-free window, and leave the proxy&rsquo;s values alone.</li>
<li>When a package &ldquo;does nothing&rdquo;, describe the pod, not the package.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/connect-to-data-sources/vsphere-supervisor-monitoring/steps-to-monitor-vsphere-supervisor-clusters-and-resources/prerequisites-for-vsphere-supervisor-monitoring.html">Enabling Monitoring for VKS Clusters</a>:
the 9.1 route and its prerequisites.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/connect-to-data-sources/vsphere-supervisor-monitoring/steps-to-monitor-vsphere-supervisor-clusters-and-resources/enabling-the-vsphere-supervisor-collection.html">Enabling Monitoring for vSphere Supervisor</a>:
the Supervisor collection setting on the vCenter account.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-consumption/latest/managing-vsphere-kuberenetes-service-clusters-and-workloads/operating-tkg-service-clusters/monitoring-vks-clusters-using-vcf-operations.html">Monitoring VKS Clusters Using VCF Operations</a>:
the automated add-ons, the metrics-aggregator service, and where
metrics appear.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-consumption/latest/managing-vsphere-kuberenetes-service-clusters-and-workloads/operating-tkg-service-clusters/disable-automated-monitoring-on-vks-clusters-using-vcf-a.html">Disable Automated Monitoring on VKS Clusters Using VCF Automation</a>:
the per-cluster label, for clusters where you run Telegraf yourself.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-service-administration-and-development/9-0/using-supervisor-services/install-the-supervisor-management-proxy-service.html">Using the Supervisor Management Proxy Service</a>:
what the proxy is for; no extra values for VCF Operations.</li>
</ul>
<p><em>Previously: <a href="/posts/telegraf-windows-2025/">Telegraf on Windows Server 2025</a>.
More in the <a href="/series/observability-on-vcf/">Observability on VCF</a> series.</em></p>
<hr>
<p><em>Lab environment; opinions my own.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Telegraf on Windows Server 2025: unsupported, works anyway</title>
      <link>https://thenestedlab.com/posts/telegraf-windows-2025/</link>
      <pubDate>Wed, 16 Sep 2026 12:20:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/telegraf-windows-2025/</guid>
      <description>The VCF Operations agent matrix doesn&amp;rsquo;t list Windows Server 2025. Add one missing Windows component, WMIC, and the normal install just works. What &amp;lsquo;unsupported&amp;rsquo; really means, and what to watch.</description>
      <content:encoded><![CDATA[<p>Two facts, both true, and a little awkward in the same room:</p>
<ol>
<li>The VCF Operations application-monitoring agent (Telegraf, packaged by
Broadcom) does not list Windows Server 2025 as a supported OS.</li>
<li>It installs from the VCF Operations UI, runs, and reports on Windows
Server 2025, once one missing Windows component is put back.</li>
</ol>
<p>This post is about the gap between those two. &ldquo;Unsupported&rdquo; is a
statement about <em>who fixes it when it breaks</em>, not about whether it
works. And there is a right way to run unsupported software in
production. It starts with knowing exactly what you&rsquo;re relying on.</p>
<h2 id="what-unsupported-means-here">What &ldquo;unsupported&rdquo; means here</h2>
<p>The matrix is a promise. Broadcom has tested this combination and will
take a support case on it. Server 2025 wasn&rsquo;t in the test set at release.</p>
<p>Nothing in the agent is gated on the OS version. It&rsquo;s Telegraf with
Windows inputs (<code>win_perf_counters</code>, <code>win_services</code>, <code>win_eventlog</code>) and
an output to Ops. The Windows APIs those inputs use haven&rsquo;t changed in a
decade.</p>
<p>So it works. You just own it.</p>
<h2 id="the-one-prerequisite-put-wmic-back">The one prerequisite: put WMIC back</h2>
<p>Windows Server 2025 no longer ships the WMI command-line tool, <code>wmic</code>.
It&rsquo;s been deprecated for years, and it&rsquo;s now a Feature on Demand rather
than part of the base install. The agent&rsquo;s install bootstrap still calls
it. So on a stock Server 2025 build, the install from Ops does not
complete.</p>
<p>Add the capability first, then install:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="n">DISM</span> <span class="p">/</span><span class="n">Online</span> <span class="p">/</span><span class="nb">Add-Capability</span> <span class="p">/</span><span class="n">CapabilityName</span><span class="err">:</span><span class="n">WMIC</span><span class="p">~~~~</span>
</span></span><span class="line"><span class="cl"><span class="n">wmic</span> <span class="n">os</span> <span class="n">get</span> <span class="n">caption</span>   <span class="c"># should now answer</span>
</span></span></code></pre></div><p>(Settings &gt; System &gt; Optional features &gt; Add &gt; &ldquo;WMIC&rdquo; does the same thing
through the UI.) That&rsquo;s the whole workaround. Not much of a plot twist, I know.</p>
<p>Note it wherever you keep build standards for Server 2025. It&rsquo;s a one-line
prerequisite for the agent, and it will stay one until the bootstrap stops
needing it.</p>
<h2 id="the-normal-install-from-the-ui">The normal install, from the UI</h2>
<p>With WMIC present, the install path is exactly the supported-OS one. No
scripts, no manual <code>telegraf.conf</code>, no heroics:</p>
<p><em>Operate &gt; Workloads &gt; Applications &gt; Manage Telegraf Agents</em>, tick the
VM, <em>Agent Actions &gt; Install</em>, supply guest credentials, wait. On the
<a href="/series/the-windows-build-pipeline/">pipeline-built W2025 server</a> it
registered as a <strong>Product Managed Agent</strong>, version 9.1.0.0.3033, with
<code>Last Operation Status</code> reading <code>Install Success</code>.</p>
<p><img alt="Manage Telegraf Agents: Test-2025 - Agent Running, Product Managed Agent, Install Success, version 9.1.0.0.3033, both collection ticks green, with a Ping Check receiving data under Custom Monitoring" loading="lazy" src="/images/ui/o3-ops-manage-telegraf-agents-w2025.jpg">
<em>The agent list with the Windows Server 2025 row expanded. Agent Running, product-managed, Install Success, both collection ticks green. Under Custom Monitoring sits a Ping Check, configured from the same screen and already receiving data, with no config file touched.</em></p>
<p><img alt="The Windows OS on Windows 2025 object: one object, Normal, no alerts, with Custom Script, Ping Check and Services children and live CPU/memory properties" loading="lazy" src="/images/ui/o9-ops-w2025-windows-os-summary.jpg">
<em>The object the agent created, as Ops sees it: green, no alerts, CPU and memory properties populated. This is the picture that matters, not the install dialog.</em></p>
<p><img alt="Ping Check metrics for the W2025 agent: Availability flat at 100 and Average Response Time in a steady band across a three-hour window" loading="lazy" src="/images/ui/o10-ops-w2025-ping-check-availability.jpg">
<em>And the proof it&rsquo;s doing work rather than just existing: a Ping Check run from the Server 2025 guest, availability flat at 100 across the morning, response time steady at a couple of milliseconds.</em></p>
<p><img alt="The VM object in Ops: Microsoft Windows Server 2025 (64-bit), tools running" loading="lazy" src="/images/ui/o2-ops-w2025-vm-summary.jpg"></p>
<p><img alt="Windows OS on Windows 2025: AgentManagedType = Product Managed, Tags|source = Windows_2025" loading="lazy" src="/images/ui/o1-ops-w2025-agent-metrics.jpg">
<em>The &ldquo;Windows OS on Windows 2025&rdquo; child object, with <code>Telegraf Availability</code> in the metric tree and <code>AgentManagedType</code> reading Product Managed.</em></p>
<h2 id="then-make-it-yours">Then make it yours</h2>
<p>Two small things turn &ldquo;it happens to work&rdquo; into something you can run:</p>
<p><strong>Record exactly what you&rsquo;re running.</strong> Agent build, Telegraf version, OS
build and the WMIC prerequisite, in whatever you use for a CMDB. When the
support matrix catches up, you&rsquo;ll want to know whether you&rsquo;re on the
version they tested.</p>
<p><strong>Give it a check that will go flat first.</strong> The Ping Check above is
configured from the agent row in Ops (<em>Custom Monitoring</em>). HTTP, TCP and
UDP checks live there too, and a <em>Custom Script</em> entry on the same screen
can run anything on the box. Point one at something trivially
OS-dependent, and alert on its <em>absence</em>. If a Windows update changes an
API under the agent, that line stops before anything else does.</p>
<h2 id="what-to-watch-because-its-unsupported">What to watch, <em>because</em> it&rsquo;s unsupported</h2>
<ul>
<li><strong>Agent upgrades from Ops.</strong> The upgrade path is tested on supported OSes.
Take a snapshot before pushing an agent upgrade to the W2025 fleet, and
upgrade one first. Optimism is not a rollback plan.</li>
<li><strong>Windows cumulative updates.</strong> Performance counter names are stable.
Provider GUIDs occasionally aren&rsquo;t. Watch your canary check after Patch
Tuesday.</li>
<li><strong>WMIC on new builds.</strong> Any image or pipeline that produces Server 2025
needs the capability added. Without it, the next install will fail the
way the first one did.</li>
<li><strong>Service account and UAC.</strong> WMIC was the only install blocker we hit.
The documented Windows catch is
<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/workload-monitoring-and-observability/os-and-application-monitoring/steps-to-monitor-your-applications/install-and-uninstall-an-agent/install-an-agent-from-the-ui.html">UAC</a>.
With UAC on and a non-administrator account in Administrators, the agent
stays at &ldquo;Not Started&rdquo; and the last operation reads &ldquo;Download Success&rdquo;.
The bootstrap then has to be run by hand.</li>
<li><strong>Don&rsquo;t file cases on it.</strong> Reproduce on a supported OS first. That&rsquo;s
the deal you made.</li>
</ul>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>The practical lesson for customers is about <strong>how</strong> to adopt something the
vendor hasn&rsquo;t blessed yet. New operating systems arrive before support
matrices catch up, and &ldquo;wait&rdquo; is often not an option. The approach here is
how an operations team gets Windows Server 2025 monitored on day one,
without taking on hidden risk:</p>
<ul>
<li>Find the real blocker (one missing Windows component, not the agent).</li>
<li>Use the standard install path.</li>
<li>Record exactly what you&rsquo;re running.</li>
<li>Add a check that detects breakage early.</li>
<li>Upgrade one node first.</li>
</ul>
<p>The same discipline applies to any unsupported-but-working combination.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>&ldquo;Unsupported&rdquo; means <em>you</em> own the fix path. Decide that consciously,
then record versions, add a canary check and upgrade one node first.</li>
<li>Server 2025 ships without WMIC, and the agent bootstrap needs it. Add the
capability first, and the ordinary UI install works.</li>
<li>The Windows inputs aren&rsquo;t gated on the OS version, and W2025 runs the
agent fine. Alert on metric <strong>absence</strong>, not just thresholds.</li>
<li>Configure checks from the agent row in Ops. Leave <code>telegraf.conf</code> alone,
so upgrades from Ops stay clean.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/workload-monitoring-and-observability/os-and-application-monitoring/supported-platforms.html">Supported Platforms</a>: the operating systems the product-managed agent supports; Windows stops at Server 2022</li>
<li><a href="https://knowledge.broadcom.com/external/article/397548/telegraf-agent-support-for-windows-serve.html">Windows Server 2025 support for Telegraf Agent in VCF and Aria Operations</a>: Server 2025 not yet supported, and the install that stays &ldquo;in progress&rdquo;</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/workload-monitoring-and-observability/os-and-application-monitoring/steps-to-monitor-your-applications/install-and-uninstall-an-agent/install-an-agent-from-the-ui.html">Install an Agent from the UI</a>: the Manage Telegraf Agents install, and what UAC changes about it</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/workload-monitoring-and-observability/os-and-application-monitoring/prerequisites/communication-with-cloud-proxy-and-vcenter-server.html">Communication with Cloud Proxy and vCenter</a>: the vCenter guest-operation privileges the UI install needs</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/workload-monitoring-and-observability/os-and-application-monitoring/steps-to-monitor-your-applications/additional-operations-from-the-manage-agents-tab/monitor-remote-checks.html">Activate Remote Checks</a>: ICMP, UDP, TCP and HTTP checks under Custom Monitoring</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/workload-monitoring-and-observability/os-and-application-monitoring/steps-to-monitor-your-applications/additional-operations-from-the-manage-agents-tab/custom-script.html">Custom Script</a>: scripts the agent runs on the box, each returning a single integer</li>
</ul>
<p><em>Previously: <a href="/posts/fluent-bit-two-ways/">fluent-bit two ways</a>. Next in
the <a href="/series/observability-on-vcf/">Observability on VCF</a> series: the
same Telegraf on VKS, and the dependency that isn&rsquo;t in its README.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Support status as observed at time of
writing — check the current matrix.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>fluent-bit two ways: VKS add-on and Windows agent, one log endpoint</title>
      <link>https://thenestedlab.com/posts/fluent-bit-two-ways/</link>
      <pubDate>Wed, 16 Sep 2026 12:00:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/fluent-bit-two-ways/</guid>
      <description>One log pipeline, two very different worlds: fluent-bit as a VKS package and as a Windows Server 2025 service, both shipping to the same VCF Operations for Logs. One backend, two configs, lessons from both.</description>
      <content:encoded><![CDATA[<p>Logs from a Kubernetes cluster and logs from a Windows server end up in
the same place: VCF Operations for Logs. The roads there couldn&rsquo;t look
more different.</p>
<p>On VKS, fluent-bit is a <em>package</em>: declare a values secret, reconcile,
done. On Windows it&rsquo;s a <em>service</em>: install it, write a config by hand,
restart, tail the debug log. Same binary, same output plugin, same
endpoint. This post puts the two side by side, so the shared shape is
easy to see.</p>
<h2 id="the-endpoint">The endpoint</h2>
<p>VCF Operations for Logs ingests over its <strong>CFAPI</strong> on port 9543:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">https://&lt;ops-logs&gt;:9543/api/v2/events
</span></span></code></pre></div><p>Auth is <a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/log-analysis/overview-of-log-management-agents/agent-authentication-for-log-ingestion.html">enabled but optional by default</a>.
The endpoint accepts events with or without a token, unless strict
authentication is turned on in Global Settings. With that on, it drops any
request without a valid Bearer token. Our shippers sent no token, and the lab ran
the default.</p>
<p>What matters to fluent-bit is the URI, the port, TLS and a JSON body
shaped as <code>{&quot;events&quot;:[...]}</code>. Both shippers below produce exactly that.
Eventually.</p>
<h2 id="way-1-vks-package">Way 1: VKS package</h2>
<p>VKS ships fluent-bit in its standard package repository. On a cluster with
the repo registered:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">kubectl ... get packages -n tkg-system 2&gt;&amp;1 | Select-String -Pattern &#39;NAME|telegraf|prometheus|fluent|cert-manager|contour&#39; | Out-String
</span></span><span class="line"><span class="cl">NAME                                                                  PACKAGEMETADATA NAME                             VERSION                  AGE
</span></span><span class="line"><span class="cl">...
</span></span><span class="line"><span class="cl">fluent-bit.kubernetes.vmware.com.4.0.2+vmware.1-vks.1                 fluent-bit.kubernetes.vmware.com                 4.0.2+vmware.1-vks.1     55s
</span></span><span class="line"><span class="cl">fluent-bit.kubernetes.vmware.com.4.0.5+vmware.1-vks.1                 fluent-bit.kubernetes.vmware.com                 4.0.5+vmware.1-vks.1     55s
</span></span><span class="line"><span class="cl">fluent-bit.tanzu.vmware.com.2.2.3+vmware.1-tkg.2                      fluent-bit.tanzu.vmware.com                      2.2.3+vmware.1-tkg.2     56s
</span></span><span class="line"><span class="cl">fluent-bit.tanzu.vmware.com.3.1.9+vmware.1-tkg.1                      fluent-bit.tanzu.vmware.com                      3.1.9+vmware.1-tkg.1     55s
</span></span><span class="line"><span class="cl">fluent-bit.tanzu.vmware.com.3.2.7+vmware.1-tkg.1                      fluent-bit.tanzu.vmware.com                      3.2.7+vmware.1-tkg.1     55s
</span></span><span class="line"><span class="cl">...
</span></span></code></pre></div><p>A values secret configures the output. The package&rsquo;s own default output
is already shaped for CFAPI, so the values are short:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">namespace</span><span class="p">:</span><span class="w"> </span><span class="l">tanzu-system-logging</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">fluent_bit</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">config</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">outputs</span><span class="p">:</span><span class="w"> </span><span class="p">|</span><span class="sd">
</span></span></span><span class="line"><span class="cl"><span class="sd">      [OUTPUT]
</span></span></span><span class="line"><span class="cl"><span class="sd">        Name          http
</span></span></span><span class="line"><span class="cl"><span class="sd">        Match         *
</span></span></span><span class="line"><span class="cl"><span class="sd">        Host          f06-flt-log01.res.lab
</span></span></span><span class="line"><span class="cl"><span class="sd">        Port          9543
</span></span></span><span class="line"><span class="cl"><span class="sd">        URI           api/v2/events
</span></span></span><span class="line"><span class="cl"><span class="sd">        Format        json
</span></span></span><span class="line"><span class="cl"><span class="sd">        tls.debug     4
</span></span></span><span class="line"><span class="cl"><span class="sd">        tls           on
</span></span></span><span class="line"><span class="cl"><span class="sd">        tls.verify    off
</span></span></span><span class="line"><span class="cl"><span class="sd">        json_date_key timestamp</span><span class="w">
</span></span></span></code></pre></div><p>Then a <code>PackageInstall</code> that points at it:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">packaging.carvel.dev/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">PackageInstall</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w"> </span>{<span class="w"> </span><span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="nt">fluent-bit, namespace</span><span class="p">:</span><span class="w"> </span><span class="l">package-installs }</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">serviceAccountName</span><span class="p">:</span><span class="w"> </span><span class="l">pkgi-sa</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">packageRef</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">refName</span><span class="p">:</span><span class="w"> </span><span class="l">fluent-bit.kubernetes.vmware.com</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">versionSelection</span><span class="p">:</span><span class="w"> </span>{<span class="w"> </span><span class="nt">constraints</span><span class="p">:</span><span class="w"> </span><span class="m">4.0.5</span><span class="l">+vmware.1-vks.1 }</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">values</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span>- <span class="nt">secretRef</span><span class="p">:</span><span class="w"> </span>{<span class="w"> </span><span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">fluent-bit-values }</span><span class="w">
</span></span></span></code></pre></div><p><code>Reconcile succeeded</code>. The proof is in the pod logs, not the Ops UI. Ops
for Logs federates its API authentication through the identity broker, so
a script has a far easier time checking from the shipper&rsquo;s side:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[ info] [output:http:http.0] f06-flt-log01.res.lab:9543, HTTP status=200
</span></span><span class="line"><span class="cl">[ info] [output:http:http.0] f06-flt-log01.res.lab:9543, HTTP status=200
</span></span></code></pre></div><p>One 200 per batch. That&rsquo;s the whole VKS side, and it almost felt like
cheating. It&rsquo;s also why the packaged path is the one to recommend. The
inputs (container logs, kubelet, systemd) come pre-wired, the parsers are
right, and the only decision you make is the output.</p>
<p><strong>One trap</strong> (shared with Telegraf, <a href="/series/observability-on-vcf/">next post</a>).
A cluster created without <code>clusterNetwork.serviceDomain</code> gives in-cluster
service names of the form <code>...svc.</code>, with no domain. Some add-ons build
those into URLs that don&rsquo;t resolve. The field is immutable after create,
so set it explicitly on every new cluster.</p>
<p><img alt="VCF Operations — Logs: text contains vks-demo01, last 24 h, 1.74K events" loading="lazy" src="/images/ui/l2-ops-logs-vks-demo01-query.jpg">
<em>The receiving end. One filter (<code>text contains vks-demo01</code>) and a day of the cluster&rsquo;s logs, one bar per hour.</em></p>
<p><img alt="The stream: container logs arriving with Kubernetes metadata — app, cluster, container, kubernetes_namespace, node, pod" loading="lazy" src="/images/ui/l1-ops-logs-vks-demo01-24h.jpg">
<em>Every event carries the fields the package&rsquo;s filters add: <code>cluster</code>, <code>kubernetes_namespace</code>, <code>pod</code> and <code>container</code>. That&rsquo;s what lets Windows and VKS sit side by side in one explorer later.</em></p>
<h2 id="way-2-windows-server-2025-service">Way 2: Windows Server 2025 service</h2>
<p>fluent-bit ships a Windows build with a <code>winevtlog</code> input. On the
<a href="/series/the-windows-build-pipeline/">pipeline-built W2025 server</a> I used the
portable zip rather than the MSI. No installer, just a folder under <code>C:\</code>
and a service registered with <code>sc.exe</code>.</p>
<p>A small script writes the config below (<code>$LogsHost</code> is the appliance&rsquo;s
FQDN, <code>$LogsPort</code> is 9543). Four of its lines turned out to matter:</p>
<details class="nl-fold">
<summary>The config the script writes: input, filter and output</summary>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-ini" data-lang="ini"><span class="line"><span class="cl"><span class="na">...</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">[INPUT]</span>
</span></span><span class="line"><span class="cl">    <span class="na">Name          winevtlog</span>
</span></span><span class="line"><span class="cl">    <span class="na">Channels      System,Application,Security</span>
</span></span><span class="line"><span class="cl">    <span class="na">Interval_Sec  5</span>
</span></span><span class="line"><span class="cl">    <span class="na">DB            $root\winevt.db</span>
</span></span><span class="line"><span class="cl">    <span class="na">String_Inserts On</span>
</span></span><span class="line"><span class="cl">    <span class="na">Tag           winevt</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">[FILTER]</span>
</span></span><span class="line"><span class="cl">    <span class="na">Name    modify</span>
</span></span><span class="line"><span class="cl">    <span class="na">Match   *</span>
</span></span><span class="line"><span class="cl">    <span class="na">Rename  Message  text</span>
</span></span><span class="line"><span class="cl">    <span class="na">Add     hostname $hostname</span>
</span></span><span class="line"><span class="cl">    <span class="na">Add     appname  v-windows</span>
</span></span><span class="line"><span class="cl">    <span class="na">Add     source   windows-server-2025</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">[OUTPUT]</span>
</span></span><span class="line"><span class="cl">    <span class="na">Name    http</span>
</span></span><span class="line"><span class="cl">    <span class="na">Match   *</span>
</span></span><span class="line"><span class="cl">    <span class="na">Host    $LogsHost</span>
</span></span><span class="line"><span class="cl">    <span class="na">Port    $LogsPort</span>
</span></span><span class="line"><span class="cl">    <span class="na">URI     /api/v2/events</span>
</span></span><span class="line"><span class="cl">    <span class="na">Format  json</span>
</span></span><span class="line"><span class="cl">    <span class="na">Header  Content-Type application/json</span>
</span></span><span class="line"><span class="cl">    <span class="na">tls     On</span>
</span></span><span class="line"><span class="cl">    <span class="na">tls.verify Off</span>
</span></span><span class="line"><span class="cl">    <span class="na">json_date_key    timestamp</span>
</span></span><span class="line"><span class="cl">    <span class="c1"># ISO-8601 timestamps are rejected (400: not a valid Long); the API reads the number as epoch MILLISECONDS</span>
</span></span><span class="line"><span class="cl">    <span class="na">json_date_format epoch_ms</span>
</span></span><span class="line"><span class="cl">    <span class="na">net.dns.resolver LEGACY</span>
</span></span><span class="line"><span class="cl">    <span class="na">net.connect_timeout 20</span>
</span></span><span class="line"><span class="cl">    <span class="na">Retry_Limit      5</span>
</span></span></code></pre></div>
</details>

<p>The service came up first time. Getting a single event to <em>land</em> took four
attempts, and each failure taught me something the documentation doesn&rsquo;t
say:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[error] [output:http:http.0] no upstream connections available to f06-flt-log01.res.lab:9543
</span></span></code></pre></div><p><strong>1. fluent-bit couldn&rsquo;t resolve a name Windows could.</strong> <code>Resolve-DnsName</code>
worked. <code>Test-NetConnection … -Port 9543</code> worked. fluent-bit&rsquo;s own async
resolver didn&rsquo;t. <code>net.dns.resolver LEGACY</code> hands DNS back to the OS.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[error] [output:http:http.0] 10.26.5.25:9543, HTTP status=404
</span></span></code></pre></div><p><strong>2. The appliance routes on the Host header.</strong> Dodging DNS by using the
IP gets a 404 for <em>every</em> path. Address it by FQDN.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">{&#34;errorMessage&#34;:&#34;Cannot deserialize value of type `java.lang.Long` from String \&#34;2026-09-16T10:55:00.000000Z\&#34;&#34;}
</span></span></code></pre></div><p><strong>3. Timestamps must be numeric.</strong> fluent-bit&rsquo;s <code>iso8601</code> output is a
string, and the ingest API wants a Long. Worse, it reads that number as
<strong>milliseconds</strong>. So on the Windows agent the default <code>double</code> (epoch
<em>seconds</em>) is accepted with a 200, and your events are filed in January
1970. That&rsquo;s a long way to scroll back. The fix is <code>epoch_ms</code>. Our VKS
values set no format at all, and those events land on time.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">{&#34;received&#34;:0,&#34;message&#34;:&#34;events ingested&#34;,&#34;status&#34;:&#34;ok&#34;}
</span></span></code></pre></div><p><strong>4. The message field must be called <code>text</code>.</strong> Anything else (<code>Message</code>
as <code>winevtlog</code> emits it, <code>message</code>, <code>log</code>) returns 200 and
<code>received: 0</code>. Silently dropped. Hence the <code>Rename</code> filter. Broadcom&rsquo;s own
<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/log-analysis/overview-of-log-management-agents/install-fluent-bit-on-windows-server-for-vcf-operations-for-logs.html">reference config for Windows</a> carries exactly this line. I found it the hard
way first.</p>
<p>Then, finally:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[ info] [output:http:http.0] f06-flt-log01.res.lab:9543, HTTP status=200
</span></span><span class="line"><span class="cl">[ info] [output:http:http.0] f06-flt-log01.res.lab:9543, HTTP status=200
</span></span></code></pre></div><p><img alt="Ops for Logs: hostname contains TEST-W2025 — Windows System, Application and Security events, with channel, computer, eventid and appname as fields" loading="lazy" src="/images/ui/w3-ops-logs-test-w2025-events.jpg">
<em>Fifty events in the first five minutes: service state changes, a &ldquo;system time was changed&rdquo; security event with its full subject block, all searchable with the same fields as everything else.</em></p>
<p><img alt="The same explorer with both hostnames in one filter: TEST-W2025 and vks-demo01" loading="lazy" src="/images/ui/w4-ops-logs-windows-and-vks-filter.jpg">
<em>One filter, two worlds. A Windows server and a Kubernetes cluster, in the same query, in the same store.</em></p>
<p>One more Windows note. A bare <code>fluent-bit.exe</code> registered with <code>sc.exe</code>
isn&rsquo;t a proper service binary. It ignores the stop signal, and Windows
sits at &ldquo;waiting for service to stop&rdquo; until you kill the process. The MSI
installs a real service wrapper, so use it for anything that isn&rsquo;t a lab.</p>
<h2 id="the-shape-they-share">The shape they share</h2>
<table>
	<thead>
			<tr>
					<th></th>
					<th>VKS package</th>
					<th>Windows service</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Install</td>
					<td><code>PackageInstall</code></td>
					<td>MSI / <code>sc.exe create</code></td>
			</tr>
			<tr>
					<td>Config</td>
					<td>values Secret</td>
					<td><code>fluent-bit.conf</code></td>
			</tr>
			<tr>
					<td>Inputs</td>
					<td>pre-wired (containers, kubelet, systemd)</td>
					<td><code>winevtlog</code> channels you choose</td>
			</tr>
			<tr>
					<td>Output</td>
					<td><code>http</code> → CFAPI :9543 <code>api/v2/events</code></td>
					<td><code>http</code> → CFAPI :9543 <code>/api/v2/events</code>, <code>epoch_ms</code>, resolver and header lines</td>
			</tr>
			<tr>
					<td>Verify</td>
					<td>pod log <code>HTTP status=200</code></td>
					<td>service log <code>HTTP status=200</code> — and check the <em>date</em> on what arrived</td>
			</tr>
			<tr>
					<td>Restart safety</td>
					<td>Kubernetes</td>
					<td><code>DB</code> bookmark file</td>
			</tr>
	</tbody>
</table>
<p>Both outputs post JSON events to the same :9543 endpoint. They differ in
the URI form, the date handling (fluent-bit&rsquo;s default against <code>epoch_ms</code>)
and the resolver and header lines that only Windows carries. That&rsquo;s the
lesson. Standardise the <em>sink</em>, and let each platform own its <em>source</em>.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>The business outcome is one place to look. Windows servers, Kubernetes
clusters and the platform itself all ship logs to the same VCF Operations
for Logs, with the same fields, searchable in one query. So an incident
that spans a Windows service and a container gets investigated on one
screen, not across three tools.</p>
<p>Standardising the <em>destination</em>, while each platform keeps its native
shipper, is also what keeps the estate maintainable. There&rsquo;s one endpoint
to secure and retain, and no bespoke agent per team.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Ops for Logs ingestion is CFAPI on <code>:9543/api/v2/events</code>, in JSON. One
endpoint serves every shipper.</li>
<li>On VKS, use the <strong>package</strong> and decide only the output. Verify from the
pod log, because the Ops API is awkward to script against.</li>
<li>Set <code>serviceDomain</code> on every new VKS cluster. It&rsquo;s immutable.</li>
<li>On Windows: the FQDN, not the IP (Host-header routing),
<code>net.dns.resolver LEGACY</code>, <code>Rename Message text</code> and
<code>json_date_format epoch_ms</code>. Each one fails differently, and two of them
fail <em>silently</em>.</li>
<li>A <code>200</code> is not proof. <code>received: 0</code> is a drop, and a seconds timestamp
from the Windows agent is a 200 filed in 1970. Look for the event in the
explorer before you call it done, however sincere the 200 looks.</li>
<li>Prove &ldquo;one endpoint&rdquo; with one explorer view showing both sources.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-service-administration-and-development/9-0/managing-vsphere-kuberenetes-service-clusters-and-workloads/installing-standard-packages-on-tkg-service-clusters/installing-standard-packages-on-tkg-cluster-using-tkr-for-vsphere-8-x/install-fluent-bit/install-fluent-bit-package.html">Install Fluent Bit Package</a>: the VKS standard package, <code>fluent-bit.kubernetes.vmware.com</code> from 4.0.x, installed with a data values file</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-service-administration-and-development/9-0/managing-vsphere-kuberenetes-service-clusters-and-workloads/installing-standard-packages-on-tkg-service-clusters/standard-package-reference/fluent-bit-package-reference.html">Fluent Bit Package Reference</a>: the package&rsquo;s values, with an <code>http</code> output to VCF Operations on port 9543</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/log-analysis/overview-of-log-management-agents/install-fluent-bit-on-windows-server-for-vcf-operations-for-logs.html">Set up the Windows system to collect logs in VCF Operations</a>: Broadcom&rsquo;s fluent-bit MSI and <code>winevtlog</code> config for Windows Server 2022 and 2025, <code>Rename Message text</code> and <code>epoch_ms</code> included</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/infrastructure-operations/log-analysis/overview-of-log-management-agents/agent-authentication-for-log-ingestion.html">Agent Authentication for Log Ingestion</a>: optional Bearer-token authentication on HTTP ingestion, fluent-bit included</li>
</ul>
<p><em>Next: <a href="/series/observability-on-vcf/">Telegraf on Windows Server 2025: unsupported, works anyway</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
