<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Vpc on The Nested Lab</title>
    <link>https://thenestedlab.com/tags/vpc/</link>
    <description>Recent content in Vpc on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Wed, 16 Sep 2026 08:30:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/tags/vpc/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>What&#39;s a VPC? Let Pac-Man explain</title>
      <link>https://thenestedlab.com/posts/whats-a-vpc-with-pacman/</link>
      <pubDate>Wed, 16 Sep 2026 08:30:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/whats-a-vpc-with-pacman/</guid>
      <description>Part 0 of the Pod Papers: an NSX VPC explained with a running game. Private by default, one deliberate door out — and a self-inflicted outage that taught me five green layers can hide one wrong integer.</description>
      <content:encoded><![CDATA[<p>Before this series gets into trunk subnets and binding maps, it&rsquo;s worth
spending ten minutes on the thing everything else stands on: <strong>what an NSX
VPC actually is</strong> from the tenant&rsquo;s chair. No slides. A game of Pac-Man.</p>
<p>The lab has Pac-Man running twice on VCF 9.1 — once on a VKS cluster I
built by hand with <code>kubectl</code>, once on a cluster deployed through VCF
Automation&rsquo;s catalog. Both live inside VPCs. Both were reachable when I
started:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">http://192.168.144.15/  -&gt;  &lt;title&gt;Pacman in HTML 5 Canvas
</span></span><span class="line"><span class="cl">http://192.168.144.23/  -&gt;  &lt;title&gt;Pacman in HTML 5 Canvas
</span></span></code></pre></div><figure class="nl-video">
  <video autoplay loop muted playsinline controls preload="metadata" style="aspect-ratio:710 / 610">
    <source src="/images/pacman-vip-23.mp4" type="video/mp4">
  </video>
  <figcaption>Pac-Man, live at <code>192.168.144.23</code> — a VIP on the VPC load balancer. The only door in.</figcaption>
</figure>

<h2 id="a-vpc-is-a-private-universe-with-a-door-policy">A VPC is a private universe with a door policy</h2>
<p>Think of an NSX VPC as a tenant&rsquo;s own routed network space: its own
subnets, its own gateway, its own address plan — carved out by the tenant,
not filed as a ticket with the network team. Three rules define it:</p>
<ol>
<li><strong>Private by default.</strong> A <code>Private</code> subnet is reachable only from inside
the same VPC. Nobody outside can route to it — not other tenants, not
other VPCs in the same org, not the corporate network.</li>
<li><strong>Your addresses are your business.</strong> Because private subnets aren&rsquo;t
advertised anywhere, two VPCs can use <em>identical</em> CIDRs. (This is the
superpower the rest of the series is built on.)</li>
<li><strong>Every door out is deliberate.</strong> Traffic leaves via the transit
gateway — SNAT&rsquo;d — or arrives via a <strong>LoadBalancer VIP</strong> from an
external block the provider allocated. Nothing is exposed by accident.</li>
</ol>
<p>Pac-Man&rsquo;s pods sit on a private subnet. The only reason <code>192.168.144.23</code>
answers is a Kubernetes <code>Service</code> of type <code>LoadBalancer</code>, which NSX turns
into a VIP on the VPC&rsquo;s load balancer. So let&rsquo;s remove the door.</p>
<h2 id="before-close-the-door">Before: close the door</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">$ kubectl patch svc pacman -n pacman -p &#39;{&#34;spec&#34;:{&#34;type&#34;:&#34;ClusterIP&#34;}}&#39;
</span></span><span class="line"><span class="cl">service/pacman patched
</span></span><span class="line"><span class="cl">NAME     TYPE        CLUSTER-IP       EXTERNAL-IP   PORT(S)   AGE
</span></span><span class="line"><span class="cl">pacman   ClusterIP   10.106.219.213   &lt;none&gt;        80/TCP    4d13h
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">$ curl -m 5 http://192.168.144.23/
</span></span><span class="line"><span class="cl">curl: timed out / unreachable
</span></span></code></pre></div><p>The pods are running. The service exists. The game is fine — <em>for anything
inside the VPC</em>. From my desk it&rsquo;s simply gone. That&rsquo;s the whole VPC model
in one <code>curl</code>: the boundary isn&rsquo;t a firewall rule somebody wrote, it&rsquo;s the
absence of a route.</p>
<h2 id="after-open-it-again">After: open it again</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">$ kubectl patch svc pacman -n pacman -p &#39;{&#34;spec&#34;:{&#34;type&#34;:&#34;LoadBalancer&#34;}}&#39;
</span></span><span class="line"><span class="cl">service/pacman patched
</span></span><span class="line"><span class="cl">NAME     TYPE           CLUSTER-IP       EXTERNAL-IP      PORT(S)        AGE
</span></span><span class="line"><span class="cl">pacman   LoadBalancer   10.106.219.213   192.168.144.23   80:31467/TCP   4d13h
</span></span></code></pre></div><p>Same VIP handed straight back. NSX programmed a virtual server and pool on
the VPC LB; the supervisor stitched it to the cluster&rsquo;s NodePort. Door open.</p>
<p>And then the game <em>didn&rsquo;t load</em>.</p>
<h2 id="the-outage-i-gave-myself-this-is-the-useful-bit">The outage I gave myself (this is the useful bit)</h2>
<p>Everything was green:</p>
<ul>
<li><code>kubectl get svc</code> — LoadBalancer, VIP assigned</li>
<li><code>kubectl get endpoints</code> — pod IPs present</li>
<li>NSX — virtual server up, pool members healthy</li>
<li><code>iptables</code> on the node — NodePort rules identical to a working
neighbour service</li>
</ul>
<p>Five layers, all green, and <code>curl</code> hung. The control experiment was the
manually-built cluster&rsquo;s Pac-Man at <code>.15</code>, untouched throughout, still
playing.</p>
<p>The cause: my &ldquo;harmless&rdquo; <code>ClusterIP</code> patch earlier had included a <code>ports</code>
list. <code>kubectl patch</code> with a merge patch <strong>replaces arrays</strong>, it doesn&rsquo;t
merge them — and my array said <code>targetPort: 80</code>. Pac-Man listens on
<strong>8080</strong>. Every layer above was faithfully forwarding traffic to a port
nothing was listening on, and every layer reported success because <em>its</em>
job was done.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">$ kubectl patch svc pacman -n pacman --type=json \
</span></span><span class="line"><span class="cl">    -p &#39;[{&#34;op&#34;:&#34;replace&#34;,&#34;path&#34;:&#34;/spec/ports/0/targetPort&#34;,&#34;value&#34;:8080}]&#39;
</span></span><span class="line"><span class="cl">service/pacman patched
</span></span><span class="line"><span class="cl">$ curl -s http://192.168.144.23/ | grep -o &#39;&lt;title&gt;.*&lt;/title&gt;&#39;
</span></span><span class="line"><span class="cl">&lt;title&gt;Pacman in HTML 5 Canvas&lt;/title&gt;
</span></span></code></pre></div><p>Instant recovery. The diagnosis walked the entire paravirtual chain —
VIP → supervisor <code>VirtualMachineService</code> → NSX VS/pool → NodePort
<code>iptables</code> → pod — and it&rsquo;s exactly the walk you&rsquo;ll need one day:</p>
<table>
	<thead>
			<tr>
					<th>Layer</th>
					<th>Check</th>
					<th>What &ldquo;green&rdquo; hides</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>VIP</td>
					<td><code>kubectl get svc</code> EXTERNAL-IP</td>
					<td>nothing about the backend</td>
			</tr>
			<tr>
					<td>NSX LB</td>
					<td>virtual server + pool status</td>
					<td>pool health is TCP to the <em>NodePort</em>, not the pod</td>
			</tr>
			<tr>
					<td>Endpoints</td>
					<td><code>kubectl get endpoints</code></td>
					<td>it lists pod IP:<strong>targetPort</strong> — read the number</td>
			</tr>
			<tr>
					<td>Node</td>
					<td><code>iptables -t nat -L KUBE-SERVICES</code></td>
					<td>rules can be perfect and point at the wrong port</td>
			</tr>
			<tr>
					<td>Pod</td>
					<td><code>kubectl exec ... ss -ltn</code></td>
					<td>the only place the truth lives</td>
			</tr>
	</tbody>
</table>
<p>Bonus find on the way: kube-proxy <em>and</em> Antrea on that cluster had dropped
their API watches days earlier (<code>http2: client connection lost</code>) and never
re-established informers until restarted. It didn&rsquo;t cause this outage, but
it&rsquo;s the kind of thing you only find when you&rsquo;re forced to look.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>If you run a platform for more than one team, this is the feature you&rsquo;ve
been asking the network team for. A VPC gives each team, project or customer
its own private network space — created by them, in minutes, with nothing
reachable from outside until they publish it. Security teams like it for the
same reason developers do: exposure is a deliberate, auditable act, not a
side effect of plugging something in.</p>
<p>What organisations do with it once they have it:</p>
<ul>
<li><strong>Per-team sandboxes</strong> that can&rsquo;t see each other, provisioned without a
ticket.</li>
<li><strong>Partner or supplier environments</strong> isolated from the corporate estate
but hosted on the same platform.</li>
<li><strong>Multi-tenant hosting</strong> — service providers and internal IT alike — with
isolation enforced by topology rather than a growing pile of firewall rules.</li>
</ul>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>A VPC&rsquo;s boundary is <strong>the absence of a route</strong>, not a rule. <code>Private</code>
subnets are unreachable from outside by construction — which is also why
identical CIDRs across VPCs just work.</li>
<li>A <code>LoadBalancer</code> service is the <em>deliberate</em> door: NSX VIP from the
external block, programmed per service. Flip the type and the door
closes with nothing else to clean up.</li>
<li><code>kubectl patch</code> (merge) <strong>replaces <code>spec.ports</code></strong>, it doesn&rsquo;t merge it.
Patch a single field with <code>--type=json</code>, or don&rsquo;t include the array.</li>
<li>Five green layers can hide one wrong integer. Keep a <em>working control</em>
(here: the untouched <code>.15</code> instance) and compare layer by layer.</li>
<li>Read <code>kubectl get endpoints</code> as <code>IP:targetPort</code> — the port is the part
people skim.</li>
</ul>
<p><em>Next in the Pod Papers: <a href="/posts/nested-esxi-nsx-vpc/">nested ESXi inside a VPC</a> —
where &ldquo;private by default&rdquo; meets a host that fakes its own MAC address.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Output captured live, trimmed for length,
never edited for outcome — including the outage.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Nested ESXi inside an NSX VPC: the trunk-subnet design</title>
      <link>https://thenestedlab.com/posts/nested-esxi-nsx-vpc/</link>
      <pubDate>Wed, 16 Sep 2026 08:20:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/nested-esxi-nsx-vpc/</guid>
      <description>Plain VPC subnets silently blackhole a nested ESXi host. Here&amp;rsquo;s why — and the trunk subnet + binding map design that makes nested labs work as an ordinary NSX VPC tenant, verified end to end.</description>
      <content:encoded><![CDATA[<p>The host booted clean. Management IP configured, services up, DCUI happy.
And every single packet it sent — ARP included — died silently.</p>
<p>That&rsquo;s how my first attempt at running nested ESXi inside an NSX VPC ended,
and the failure mode is nasty precisely because nothing <em>looks</em> wrong. If
you&rsquo;re trying to build nested vSphere labs on VCF 9 with VPC networking,
this post is the map of the minefield — and the design that gets you across
it, verified live.</p>
<h2 id="the-setup">The setup</h2>
<p>VCF 9.1, vSphere Supervisor with NSX VPC networking. The goal: deploy nested
ESXi hosts as ordinary VM Service VMs inside a tenant&rsquo;s VPC — no physical
fabric changes, no provider tickets, no special treatment. The kind of thing
you want for training pods, cert-study labs, or reproducing customer issues.</p>
<p>Nested ESXi needs what physical ESXi needs: a management network, vMotion,
vSAN — traditionally VLANs trunked to every host. But a VPC is an overlay
world. There are no VLANs to trunk. So what happens if you just attach the
nested host&rsquo;s vNIC to a normal VPC subnet?</p>
<h2 id="failure-1-the-silent-blackhole">Failure #1: the silent blackhole</h2>
<p>Here&rsquo;s the trap. A standard VPC subnet port gets <strong>address bindings</strong>: NSX
pins the exact IP + MAC it allocated to that vNIC, and SpoofGuard drops
everything else.</p>
<p>ESXi&rsquo;s vmk0 doesn&rsquo;t use the vNIC&rsquo;s MAC. It synthesises its own:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vmk0
</span></span><span class="line"><span class="cl">   MAC Address: 00:50:ac:1e:00:8c     &lt;- NOT the vNIC MAC (04:50:56:...)
</span></span></code></pre></div><p>So every frame the management interface sends carries a MAC the port doesn&rsquo;t
own. NSX drops it all — ARP, ping, everything — while the host itself boots
green and reports healthy. There is no error anywhere. You just can&rsquo;t reach
it, ever.</p>
<p><img alt="Standard VPC subnet port: SpoofGuard pins one IP+MAC; vmk0&rsquo;s synthesised MAC loses, silently" loading="lazy" src="/images/post1-blackhole.svg"></p>
<p>(There&rsquo;s a second trap stacked on top: VPC subnets run with DHCP deactivated,
so the appliance also sits at &ldquo;waiting for DHCP&rdquo; unless you inject static
addressing via OVF <code>guestinfo.*</code> properties. More on that below.)</p>
<h2 id="the-design-that-works-a-trunk-subnet--binding-maps">The design that works: a trunk subnet + binding maps</h2>
<p>The fix isn&rsquo;t a hack — it&rsquo;s a first-class NSX VPC construct that&rsquo;s barely
documented in the wild: <strong><code>SubnetConnectionBindingMap</code></strong>.</p>
<p>The idea:</p>
<ol>
<li>Create one ordinary VPC subnet to act as a <strong>trunk</strong> (<code>sn-trunk</code>). The
nested host&rsquo;s vNICs attach <em>only</em> here.</li>
<li>Create a normal VPC subnet per traditional network — <code>sn-mgmt</code>,
<code>sn-vmotion</code>, <code>sn-vsan</code>.</li>
<li>Bind each of those to the trunk with a <strong>binding map carrying a VLAN tag</strong>.
The nested host&rsquo;s vSwitch tags frames exactly as it would on metal; the
binding map strips the tag and delivers the frame into the right subnet.</li>
</ol>
<p>Pure L2 demultiplexing. One vNIC carries N VLANs, the VPC never routes on a
tag, and the physical fabric never sees any of it (the 802.1Q header rides
inside the Geneve overlay).</p>
<p>All of it is tenant-creatable through the supervisor as Kubernetes objects:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="c"># sn-trunk and sn-mgmt are ordinary Private Subnets; the interesting object:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">crd.nsx.vmware.com/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">SubnetConnectionBindingMap</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w"> </span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">bm-mgmt}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">subnetName</span><span class="p">:</span><span class="w"> </span><span class="l">sn-mgmt         </span><span class="w"> </span><span class="c"># the map is a child of the VLAN subnet...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">targetSubnetName</span><span class="p">:</span><span class="w"> </span><span class="l">sn-trunk  </span><span class="w"> </span><span class="c"># ...and points AT the trunk</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vlanTrafficTag</span><span class="p">:</span><span class="w"> </span><span class="m">1610</span><span class="w">
</span></span></span></code></pre></div><p>That direction is easy to invert, so it&rsquo;s worth saying twice: <strong>the binding
map belongs to the VLAN subnet and points at the trunk</strong>, not the other way
round.</p>
<p><img alt="NSX: sn-trunk realized once per VPC, binding maps hanging off the VLAN subnets" loading="lazy" src="/images/ui/u11b-nsx-sntrunk-per-vpc.jpg"></p>
<p>On the nested host, nothing exotic — plain VST, like physical:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Name                Virtual Switch  Active Clients  VLAN ID
</span></span><span class="line"><span class="cl">------------------  --------------  --------------  -------
</span></span><span class="line"><span class="cl">Management Network  vSwitch0                     1     1610
</span></span><span class="line"><span class="cl">vMotion             vSwitch0                     1     1611
</span></span><span class="line"><span class="cl">vSAN                vSwitch0                     1     1612
</span></span></code></pre></div><p><img alt="Host Client: port groups on VLANs 1610 / 1611 / 1612" loading="lazy" src="/images/ui/u12a-hostclient-portgroups-vlans.jpg">
<em>The same three VLANs as the nested host sees them.</em></p>
<p>And because there&rsquo;s no DHCP in a VPC subnet, the nested-ESXi appliance gets
its identity through OVF properties in the VM Service spec:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vAppConfig</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">properties</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.ipaddress, value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.40&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.netmask,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;255.255.255.224&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.gateway,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.33&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.vlan,     value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;1610&#34;</span>}}<span class="w">
</span></span></span></code></pre></div><h2 id="does-it-actually-work-the-receipts">Does it actually work? The receipts</h2>
<p>Two nested hosts, vNICs on <code>sn-trunk</code>, three VLANs. From host one:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[root@esx01:~] vmkping -c2 172.30.0.41            # mgmt, VLAN 1610
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">[root@esx01:~] vmkping -I vmk1 -c3 172.30.0.71    # vMotion, VLAN 1611
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">[root@esx01:~] vmkping -I vmk2 -c3 172.30.0.101   # vSAN, VLAN 1612
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span></code></pre></div><p><img alt="Live capture: vmnic0 down, vMotion and vSAN VLANs still passing at 0% loss" loading="lazy" src="/images/demo-c6-nic-failover.jpg">
<em>The transcript that matters: fail the first NIC, and every VLAN keeps flowing on the second — captured live.</em></p>
<p>Two more results worth knowing before you design around this:</p>
<p><strong>Untagged frames are dropped.</strong> I put a probe vmk on the untagged
portgroup using the address NSX itself had allocated to the trunk port:
100% loss, empty ARP table, while tagged traffic flowed happily beside it.
Every network your nested host uses needs a VLAN and a binding map — there
is no untagged fallback.</p>
<p><strong>Failover behaves like real hardware.</strong> With two vNICs on the trunk teamed
active/active, <code>esxcli network nic down -n vmnic0</code> moved every VLAN onto
vmnic1 with zero loss — and the SSH session I was watching from never
dropped. The vmk MAC migrating between trunk ports mid-flow is exactly the
scenario that MAC-pinned standard ports would blackhole; the trunk carries
it fine.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Running whole vSphere environments <em>inside</em> a VPC turns the platform into
something most customers never had: a way to stand up complete, isolated
copies of infrastructure on demand, without a physical fabric change and
without waiting for anyone. That&rsquo;s what makes it commercially interesting:</p>
<ul>
<li><strong>Training and certification labs</strong> where every learner gets a real
vSphere environment, not a shared one.</li>
<li><strong>Reproducing a customer problem</strong> on a like-for-like copy instead of on
the customer&rsquo;s estate.</li>
<li><strong>Rehearsing upgrades and migrations</strong> end to end before the change
window, then throwing the copy away.</li>
<li><strong>Vendor and feature evaluations</strong> with real behaviour, at zero risk to
production.</li>
</ul>
<p>This is the design Comms-care uses to give every consultant a dedicated
environment, and the same pattern scales to a classroom or a proof-of-concept
factory.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>A nested ESXi vNIC on a <strong>standard</strong> VPC subnet is dead on arrival:
vmk0&rsquo;s synthesised MAC loses to SpoofGuard, silently.</li>
<li>Attach nested-host vNICs <strong>only to a trunk subnet</strong>; one binding map per
VLAN; the map lives under the VLAN subnet and points at the trunk.</li>
<li><strong>No DHCP in VPC subnets</strong> — bootstrap addressing via <code>guestinfo.*</code>
(appliances) or cloud-init (Linux). Static IP plans are a feature in a
lab anyway.</li>
<li>ESXi&rsquo;s default TCP/IP stack has <strong>one</strong> gateway — set per-vmk override
gateways (<code>esxcli ... ipv4 set -g</code>) so vMotion/vSAN carry their own
subnet&rsquo;s gateway.</li>
<li>Recreating a VM <strong>reallocates</strong> its NSX addresses. Pin what you depend on.</li>
<li>MTU: everything here ran at 1500. Raise the trunk and the nested vDS
before you do vSAN at any real scale.</li>
</ul>
<p>Next in this series: what happens when you want <em>ten</em> of these labs — with
byte-identical IP plans, firewalled from each other by construction. That&rsquo;s
where NSX VPCs go from &ldquo;workaround&rdquo; to genuinely better than physical.</p>
<hr>
<p><em>Lab environment; opinions my own. Everything above was captured from a live
VCF 9.1 environment — output trimmed for length, never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>The load balancer that must exist before the namespace</title>
      <link>https://thenestedlab.com/posts/the-lb-that-must-exist-first/</link>
      <pubDate>Wed, 16 Sep 2026 08:10:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/the-lb-that-must-exist-first/</guid>
      <description>VIPs pending forever, a retryable error that never stops retrying, and an ordering rule the docs don&amp;rsquo;t tell you: in a self-service NSX VPC, the LBService must exist before the namespace that will use it.</description>
      <content:encoded><![CDATA[<p>Everything was green. The VPC: realized. The namespace: ready. The VMs:
powered on, endpoints populated, ports listening. And the LoadBalancer
services sat at <code>&lt;pending&gt;</code> — for an hour.</p>
<p>This is the story of the least helpful error message in my recent memory,
what it actually means, and the one-line ordering rule that would have saved
an afternoon. If you&rsquo;re doing self-service NSX VPCs on VCF 9 with the
vSphere Supervisor, you will hit this. Bookmark accordingly.</p>
<h2 id="the-setup">The setup</h2>
<p>Tenant-created VPC (via the VCF Automation CCI API), a supervisor namespace
pinned to it, and a couple of <code>VirtualMachineService</code> objects of type
<code>LoadBalancer</code> to publish SSH and HTTPS for the workloads inside. Standard
stuff — the exact pattern that works out of the box in the org&rsquo;s default VPC.</p>
<p>The k8s side looked perfect:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">$ kubectl get endpoints -n pod-a
</span></span><span class="line"><span class="cl">NAME           ENDPOINTS                        AGE
</span></span><span class="line"><span class="cl">esx01-access   172.30.0.40:443,172.30.0.40:22   6m36s
</span></span><span class="line"><span class="cl">esx02-access   172.30.0.41:443,172.30.0.41:22   6m35s
</span></span></code></pre></div><p>Endpoints resolved. VIPs: nothing. The only clue, a recurring event:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Warning  FailedRealizeNSXResource  service/esx01-access
</span></span><span class="line"><span class="cl">Generic error occurred during realizing network for Service
</span></span></code></pre></div><p>&ldquo;Generic error.&rdquo; Wonderful.</p>
<h2 id="digging-what-ncp-actually-wants">Digging: what NCP actually wants</h2>
<p>The supervisor&rsquo;s network container plugin (NCP) logs told the real story:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">nsx_ujo.ncp.nsx.policy.lb_layer4_service Lb Service not Found for Namespace pod-a
</span></span><span class="line"><span class="cl">NCP00270 Failed to process virtual ip for service ...: Lbs pod-a is not found
</span></span><span class="line"><span class="cl">Encountered retryable error ... : Lbs pod-a is not found
</span></span></code></pre></div><p>NCP wants an NSX <strong>LBService</strong> in the namespace&rsquo;s VPC. In the org&rsquo;s
<em>default</em> VPC, one exists — the platform created it when the VPC was born.
In my self-service VPC? Nobody had created one. Fair enough — that&rsquo;s
actually documented behaviour once you know where to look: a fresh VPC needs
a <code>LoadBalancer</code> object (and before that, a <code>VPCAttachment</code> to a
connectivity profile with the service gateway enabled, or the LB creation
itself fails with a much better error message).</p>
<p>So I created the attachment, then the LBService. NSX: <code>Realized=True</code>.
Problem solved?</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Warning  FailedRealizeNSXResource  service/esx01-access
</span></span><span class="line"><span class="cl">Generic error occurred during realizing network for Service
</span></span></code></pre></div><p>No.</p>
<h2 id="the-actual-bug-shaped-behaviour-a-snapshot-not-a-lookup">The actual bug-shaped behaviour: a snapshot, not a lookup</h2>
<p>Here&rsquo;s the part that costs you the afternoon. That &ldquo;retryable error&rdquo; retries
the <em>lookup in NCP&rsquo;s cache</em> — not the discovery. <strong>NCP snapshots the VPC&rsquo;s
LB inventory when the namespace is created.</strong> An LBService that appears
afterwards is never discovered, no matter how long you wait:</p>
<ul>
<li>Recreating the k8s services: no effect.</li>
<li>Tagging the LBService with the <code>nsx-op/*</code> ownership tags the working ones
carry: no effect — the cache doesn&rsquo;t re-read NSX.</li>
<li>Restarting NCP would force a full resync — but supervisor system pods are
protected; even <code>Administrator@vsphere.local</code> gets a Forbidden.</li>
<li>Mutating the namespace to nudge a re-sync: also blocked, by the
supervisor&rsquo;s namespace validation webhook.</li>
</ul>
<p>As a tenant, there is exactly one fix: <strong>delete and recreate the namespace</strong>,
now that its VPC has an LB. Fifteen minutes of rebuild for want of one
ordering rule.</p>
<p>And the control experiment proves the rule: a namespace created <em>after</em> its
VPC already had an LBService got its VIPs assigned without any drama —</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">esx01-access   VIP=192.168.144.34   22 OPEN · 443 OPEN
</span></span><span class="line"><span class="cl">esx02-access   VIP=192.168.144.35   22 OPEN · 443 OPEN
</span></span></code></pre></div><p>And once the ordering is right, this is what &ldquo;working&rdquo; looks like — the
pod&rsquo;s state a couple of minutes after a correctly-ordered deployment:</p>
<p><img alt="Live replay: catalog-deployed pod with both VMs powered on and VIPs assigned" loading="lazy" src="/images/c2-catalog-pod.gif"></p>
<p><img alt="VCFA deployment topology: namespace, subnets, hosts, two VIPs" loading="lazy" src="/images/ui/u4-deployment-topology.jpg">
<em>What the requester sees once the order is right.</em></p>
<h2 id="the-ordering-rule">The ordering rule</h2>
<p>For every self-service VPC that will publish LoadBalancer services, create —
in this order, <em>before</em> the namespace:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">1. VPC                                   (vpc.nsx.vmware.com/v1alpha1)
</span></span><span class="line"><span class="cl">2. VPCAttachment                         (connectivity profile w/ service gateway
</span></span><span class="line"><span class="cl">                                          — LB creation errors without it)
</span></span><span class="line"><span class="cl">3. LoadBalancer   {regionName, vpcName}  (the step everyone misses)
</span></span><span class="line"><span class="cl">4. ...and only THEN the Supervisor Namespace
</span></span></code></pre></div><p>Encode it in whatever provisions your VPCs — a script, a pipeline, an
operator. It&rsquo;s four API calls and it turns a silent, undiagnosable
<code>&lt;pending&gt;</code> into a platform that just works.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Nobody buys a platform for its ordering rules — but this is exactly the kind
of edge that decides whether self-service provisioning feels reliable or
flaky to the people using it. In a customer deployment the answer isn&rsquo;t a
blog post; it&rsquo;s that the provisioning automation already does the four
steps in the right order, every time, so a tenant never sees a VIP stuck at
<code>&lt;pending&gt;</code>. Knowing where the sharp edges are — because you&rsquo;ve been cut by
them in a lab — is most of what an experienced delivery partner is for.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>In a self-service NSX VPC, <strong>the LBService must predate the namespace</strong>.
NCP discovers LBs at namespace-add and never again.</li>
<li><code>FailedRealizeNSXResource: Generic error</code> on a Service = go read the NCP
logs; the real message (<code>Lbs &lt;ns&gt; is not found</code>, NCP00270) is there.</li>
<li><code>VPCAttachment</code> (service gateway) is the prerequisite for the LB itself —
that one at least fails loudly.</li>
<li>Retro-tagging NSX objects to look &ldquo;owned&rdquo; doesn&rsquo;t help a cache that never
re-reads. Recreating the namespace is the only tenant-level fix.</li>
<li>While you&rsquo;re at it: new namespaces also reject VM creation until image
<code>status.disks</code> syncs (~1–3 minutes after content library attach). Build
the wait into your automation and both sharp edges disappear.</li>
</ul>
<p><em>Previously in this series: <a href="/posts/nested-esxi-nsx-vpc/">nested ESXi inside an NSX VPC</a>.
Next: three datacenters, one IP plan — identical isolated pods.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Output captured live, trimmed for length,
never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Three datacenters, one IP plan: identical isolated pods with NSX VPCs</title>
      <link>https://thenestedlab.com/posts/three-datacenters-one-ip-plan/</link>
      <pubDate>Wed, 16 Sep 2026 08:00:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/three-datacenters-one-ip-plan/</guid>
      <description>Three nested-ESXi pods, byte-identical addressing — same subnets, same VLANs, same host IPs, even the same MACs — with zero reachability between them. How overlapping VPC CIDRs and deterministic subnet realization turn cookie-cutter environments into a first-class feature.</description>
      <content:encoded><![CDATA[<p>Here are three hosts, all answering to <code>vmk0 = 172.30.0.40</code>:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">ssh root@192.168.144.30  -&gt;  [root@esx01-a:~]  vmk0  172.30.0.40
</span></span><span class="line"><span class="cl">ssh root@192.168.144.32  -&gt;  [root@esx01-b:~]  vmk0  172.30.0.40
</span></span><span class="line"><span class="cl">ssh root@192.168.144.34  -&gt;  [root@esx01-c:~]  vmk0  172.30.0.40
</span></span></code></pre></div><p>Same IP. Same VLAN. Same gateway. Same <em>MAC address</em>, as it turns out. And
none of them can reach any of the others. This is the post where NSX VPCs
stop being a workaround for nested labs and become genuinely better than
the physical alternative.</p>
<p><img alt="Three pods, identical IP plans, no route between them" loading="lazy" src="/images/product-01-hook.jpg"></p>
<h2 id="why-identical-addressing-matters">Why identical addressing matters</h2>
<p>If you&rsquo;ve ever built training pods, cert-study labs, or per-team
reproduction environments, you know the pain: every copy needs a unique
address plan, so every runbook, every screenshot, every &ldquo;type this exact
command&rdquo; has to be parameterised per pod. Students in seat 7 see different
numbers from the slides. Reproductions drift from the original.</p>
<p>The fix is obvious and normally impossible: <strong>give every pod the same
addresses</strong>. On a physical fabric that means VRFs, per-pod NAT, and a
network team that stops answering your emails. In an NSX VPC it&rsquo;s the
default behaviour.</p>
<h2 id="the-mechanism-overlapping-privateips">The mechanism: overlapping privateIPs</h2>
<p>Each pod gets its own VPC, and every VPC declares the same private range:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">vpc.nsx.vmware.com/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">VPC</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w"> </span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">nested-vpc-a}     </span><span class="w"> </span><span class="c"># then -b, then -c</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">privateIPs</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="s2">&#34;172.30.0.0/16&#34;</span><span class="p">]</span><span class="w">      </span><span class="c"># identical in all three</span><span class="w">
</span></span></span></code></pre></div><p>A <code>Private</code> subnet is never advertised beyond its VPC, so NSX has no
objection to three VPCs carving up the same /16. The pods aren&rsquo;t
&ldquo;firewalled from each other&rdquo; — there is simply no route between them.
Isolation by construction, not by policy.</p>
<p><img alt="NSX: four VPCs, four sn-mgmt subnets, same CIDR" loading="lazy" src="/images/ui/u11a-nsx-snmgmt-four-vpcs.jpg">
<em>NSX&rsquo;s subnet view filtered to <code>sn-mgmt</code>: four rows, four VPCs, one CIDR.</em></p>
<h2 id="the-trick-deterministic-realization">The trick: deterministic realization</h2>
<p>Identical <em>ranges</em> aren&rsquo;t enough — I want identical <em>subnets</em>, so the
management gateway is <code>.33</code> and the hosts are <code>.40</code>/<code>.41</code> in every pod.
NSX allocates subnets from <code>privateIPs</code> in creation order, and a fresh VPC
allocates deterministically. So the topology is applied in a <strong>fixed
order</strong> — trunk, mgmt, vMotion, vSAN — and every pod realizes the same
map:</p>
<table>
	<thead>
			<tr>
					<th>Subnet</th>
					<th>Realized</th>
					<th>VLAN</th>
					<th>Hosts</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>sn-trunk</td>
					<td>172.30.0.0/27</td>
					<td>—</td>
					<td>(carries the tags)</td>
			</tr>
			<tr>
					<td>sn-mgmt</td>
					<td>172.30.0.32/27</td>
					<td>1610</td>
					<td>.40 / .41, gw .33</td>
			</tr>
			<tr>
					<td>sn-vmotion</td>
					<td>172.30.0.64/27</td>
					<td>1611</td>
					<td>.70 / .71, gw .65</td>
			</tr>
			<tr>
					<td>sn-vsan</td>
					<td>172.30.0.96/27</td>
					<td>1612</td>
					<td>.100 / .101, gw .97</td>
			</tr>
	</tbody>
</table>
<p>In the catalog blueprint that order is enforced with <code>dependsOn</code> between
the subnet resources — the one place a declarative tool needs to be told
about sequence. Skip it and two pods can come out with mgmt and vMotion
swapped, which works perfectly and confuses everyone.</p>
<p><img alt="NSX: nested-vpc-a expanded, the /16 private block" loading="lazy" src="/images/ui/u11-nsx-vpc-a-cidr.jpg"></p>
<h2 id="the-door-one-vip-per-host">The door: one VIP per host</h2>
<p>Each pod is unreachable from outside by design, so each host gets a
<code>VirtualMachineService</code> of type <code>LoadBalancer</code> publishing SSH and HTTPS.
The VIPs come from the org&rsquo;s <em>external</em> block, and they&rsquo;re the only
addresses that differ between pods:</p>
<table>
	<thead>
			<tr>
					<th>Pod</th>
					<th>VPC</th>
					<th>esx01 VIP</th>
					<th>esx02 VIP</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>a</td>
					<td>nested-vpc-a</td>
					<td>192.168.144.30</td>
					<td>.31</td>
			</tr>
			<tr>
					<td>b</td>
					<td>nested-vpc-b</td>
					<td>192.168.144.32</td>
					<td>.33</td>
			</tr>
			<tr>
					<td>c</td>
					<td>nested-vpc-c</td>
					<td>192.168.144.34</td>
					<td>.35</td>
			</tr>
	</tbody>
</table>
<p>Which is how the opening transcript works: three VIPs, three hosts, one
inside address.</p>
<h2 id="proving-the-isolation">Proving the isolation</h2>
<p>Claims are cheap. The test matrix, from a VM in a <em>fourth</em> VPC (the org
default):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">ping 172.30.0.140 (own VPC)....... REACHABLE
</span></span><span class="line"><span class="cl">ping 172.30.0.40  (pod space)..... unreachable
</span></span><span class="line"><span class="cl">curl http://172.31.0.2/ (shared).. shared-svc repo01
</span></span></code></pre></div><p><img alt="Isolation matrix: own-VPC reachable, pod space unreachable, shared service reachable" loading="lazy" src="/images/demo-c7-isolation.jpg"></p>
<p>Its own VPC&rsquo;s <code>172.30.0.140</code>: reachable. <code>172.30.0.40</code> — an address that
exists in three other VPCs simultaneously: unreachable, because from here
there is no such route. (The third line is the shared-services VPC, which
is <a href="/series/the-vpc-pod-papers/">the next post</a>.)</p>
<p>Inside each pod, east-west is normal: <code>esx01 → esx02</code> vmkping passes on
all three VLANs, in all three pods. And the detail I didn&rsquo;t expect: the
nested-ESXi appliance derives vmk0&rsquo;s MAC deterministically from its
config, so <strong>the three hosts share a MAC as well as an IP</strong>. Harmless — each
VPC is its own L2 domain — but a nice demonstration of how complete the
separation is.</p>
<h2 id="what-this-replaces">What this replaces</h2>
<table>
	<thead>
			<tr>
					<th></th>
					<th>Physical / VLAN-based pods</th>
					<th>VPC pods</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Identical addressing</td>
					<td>VRF per pod + NAT, fabric change per pod</td>
					<td>default behaviour</td>
			</tr>
			<tr>
					<td>Adding a pod</td>
					<td>switch config, IPAM, firewall rules</td>
					<td>one API call for the VPC, one blueprint request</td>
			</tr>
			<tr>
					<td>Isolation guarantee</td>
					<td>policy (auditable, breakable)</td>
					<td>topology (no route exists)</td>
			</tr>
			<tr>
					<td>Tenant self-service</td>
					<td>no</td>
					<td>yes — the VPC is a tenant object</td>
			</tr>
	</tbody>
</table>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>&ldquo;Identical environments&rdquo; sounds like a lab nicety. It&rsquo;s actually one of the
most requested things in enterprise IT, usually asked for in other words:</p>
<ul>
<li><strong>Training at scale</strong> — every seat in the room sees the same addresses as
the slides, so material is written once and never parameterised per pod.</li>
<li><strong>Per-engineer or per-team replicas</strong> of a reference environment, for
development and testing that behaves exactly like the original.</li>
<li><strong>Regulatory or business-unit separation</strong> on shared infrastructure
without VRF sprawl or a bespoke firewall estate — isolation is a property
of the topology, which is the easiest kind to evidence to an auditor.</li>
<li><strong>Blue/green copies</strong> of an environment for change rehearsal, then
cut-over or discard.</li>
</ul>
<p>On a physical network each of these is a project. On VCF with NSX VPCs it&rsquo;s
a template.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Overlapping <code>privateIPs</code> across VPCs is <strong>supported and intentional</strong>.
Identical pods are a feature, not a hack.</li>
<li>Fresh VPCs realize subnets <strong>deterministically in creation order</strong> —
fix the order (<code>dependsOn</code> in a blueprint) and every pod gets the same
map.</li>
<li>Pods are unreachable from outside by construction; publish exactly what
you mean to via <code>LoadBalancer</code> VIPs from the external block.</li>
<li>Prove isolation from a <em>different</em> VPC, with a positive control (own
VPC reachable) beside the negative.</li>
<li>Expect duplicate MACs across pods from appliance images. It&rsquo;s fine.</li>
</ul>
<p><em>Previously: <a href="/posts/the-lb-that-must-exist-first/">the LB that must exist first</a>.
Next: one WSUS for pods that can&rsquo;t see each other.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Output captured live, trimmed for length,
never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Shared services for isolated tenants: PrivateTGW subnets</title>
      <link>https://thenestedlab.com/posts/shared-services-for-isolated-tenants/</link>
      <pubDate>Wed, 16 Sep 2026 07:50:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/shared-services-for-isolated-tenants/</guid>
      <description>Three pods with identical private addressing all need the same WSUS, repo and AD. One shared-services VPC with a PrivateTGW subnet serves all of them over the transit gateway — and can&amp;rsquo;t reach back into any of them. The directional test, and why SNAT is what makes it work.</description>
      <content:encoded><![CDATA[<p>The pods from <a href="/posts/three-datacenters-one-ip-plan/">the last post</a> are
perfectly isolated. That&rsquo;s the requirement — and immediately the problem.
Every one of them needs Windows updates, a package repo, DNS, maybe a
domain controller. Do I really run a WSUS <em>per pod</em>?</p>
<p>No. There&rsquo;s a third subnet access mode for exactly this, and the design it
enables is hub-and-spoke with a very specific property: <strong>spokes reach the
hub; the hub cannot reach the spokes; spokes never reach each other.</strong></p>
<h2 id="the-three-access-modes">The three access modes</h2>
<p>Everything in this series comes down to one field on a VPC subnet:</p>
<table>
	<thead>
			<tr>
					<th><code>accessMode</code></th>
					<th>Advertised to</th>
					<th>Use</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td><code>Private</code></td>
					<td>nobody outside the VPC</td>
					<td>workloads — isolation <em>and</em> overlapping CIDRs</td>
			</tr>
			<tr>
					<td><code>PrivateTGW</code></td>
					<td>every VPC attached to the org&rsquo;s transit gateway</td>
					<td>shared services</td>
			</tr>
			<tr>
					<td><code>Public</code></td>
					<td>the external network</td>
					<td>internet/corp-facing endpoints</td>
			</tr>
	</tbody>
</table>
<p><code>PrivateTGW</code> subnets draw their addresses from a <strong>separate transit
block</strong> (here <code>172.31.0.0/…</code>), not from the VPC&rsquo;s own <code>privateIPs</code>. That&rsquo;s
the key: the shared range can&rsquo;t collide with the pods&rsquo; <code>172.30.0.0/16</code>
because it comes from a different pool that all VPCs agree on.</p>
<h2 id="the-build">The build</h2>
<p>One more VPC, <code>shared-svc</code>, with the same ordering rules as any other
(VPC → VPCAttachment → LoadBalancer → namespace), then a single subnet:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">crd.nsx.vmware.com/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">Subnet</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w"> </span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">sn-services}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w"> </span>{<span class="nt">accessMode</span><span class="p">:</span><span class="w"> </span><span class="nt">PrivateTGW, ipv4SubnetSize</span><span class="p">:</span><span class="w"> </span><span class="m">32</span>}<span class="w">
</span></span></span></code></pre></div><p>It realized as <code>172.31.0.0/27</code>, and a VM on it — <code>svc-repo01</code>,
<code>172.31.0.2</code>, serving HTTP — became the shared repo.</p>
<p><img alt="Architecture: VPC per pod, transit gateway, shared-services VPC" loading="lazy" src="/images/product-02-architecture.jpg"></p>
<h2 id="the-test-that-matters-is-directional">The test that matters is directional</h2>
<p>Reachability <em>to</em> the service is the easy claim. From <code>esx01</code> in each of
the three pods (ESXi ships python3, so <code>urllib</code> is the test client):</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">pod-a esx01 -&gt; http://172.31.0.2/   200  shared-svc repo01
</span></span><span class="line"><span class="cl">pod-b esx01 -&gt; http://172.31.0.2/   200  shared-svc repo01
</span></span><span class="line"><span class="cl">pod-c esx01 -&gt; http://172.31.0.2/   200  shared-svc repo01
</span></span><span class="line"><span class="cl">default-vpc  -&gt; http://172.31.0.2/   200  shared-svc repo01
</span></span></code></pre></div><p>Three pods with <strong>identical source addresses</strong> (<code>172.30.0.40</code>) all hit
one service and all get answers. How does the reply find its way back to
the right pod when three of them claim <code>.40</code>? Because pod traffic crosses
the transit gateway <strong>SNAT&rsquo;d to a per-VPC transit address</strong>. The service
never sees <code>172.30.0.40</code>; it sees three distinct transit IPs. Ambiguity
never arises.</p>
<p>Now the other direction — from <code>svc-repo01</code> back toward a pod:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">svc-repo01 -&gt; 172.30.0.40   unreachable
</span></span><span class="line"><span class="cl">svc-repo01 -&gt; 172.30.0.41   unreachable
</span></span></code></pre></div><p>Not &ldquo;blocked&rdquo; — <em>unroutable</em>. Pod subnets are <code>Private</code>, so they were never
advertised to the transit gateway, and even if they had been, <code>172.30.0.40</code>
would be ambiguous across three VPCs. The hub literally cannot initiate
into a spoke. For a shared service that will one day be compromised,
that&rsquo;s the property you want.</p>
<h2 id="what-goes-in-the-hub">What goes in the hub</h2>
<p>Anything that&rsquo;s <em>consumed</em> by pods and <em>stateless about which pod is
asking</em>: WSUS/patch mirrors, OS and package repos, container registries,
NTP, DNS forwarders, license servers. Domain controllers work too, with
the usual caveat that identical hostnames across pods need per-pod domains
or a naming scheme.</p>
<p>What does <strong>not</strong> go in the hub: anything that needs to <em>reach into</em> a pod
(monitoring pollers, backup agents pulling, jump hosts). Those either live
in the pod, or the pod publishes a <code>LoadBalancer</code> VIP for them — the
deliberate door from <a href="/posts/whats-a-vpc-with-pacman/">part 0</a>.</p>
<h2 id="tightening-further">Tightening further</h2>
<p>The transit gateway gives you reachability; policy gives you precision. A
<code>VPCGatewayFirewallPolicy</code> on <code>shared-svc</code> can restrict inbound to
<code>tcp/80,443</code> from the transit range and nothing else, so the repo is a
repo and not a foothold. I left it open for the test; you shouldn&rsquo;t.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>This is the pattern that makes isolated tenants <em>affordable</em>. Without it,
every isolated environment needs its own patch server, repository, DNS and
directory — cost and drift that quietly kill the idea. With it, a customer
runs one set of shared services for dozens of tenants, keeps them patched in
one place, and can still show a security reviewer that the shared service
has no path back into any tenant.</p>
<p>The same hub serves well beyond patching: central logging and monitoring
collectors, licence servers, artifact registries, build agents — anything
tenants consume but shouldn&rsquo;t be able to be reached <em>by</em>.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li><code>PrivateTGW</code> is the shared-services mode: addresses from the <strong>transit
block</strong>, advertised to every attached VPC, no collision with pod space.</li>
<li>Access is <strong>one-way by construction</strong>: pods → service works (SNAT&rsquo;d
per VPC), service → pod has no route. Test both directions and write
down both results.</li>
<li>Identical pod addressing and shared services coexist <em>because</em> of the
SNAT — the hub sees per-VPC transit addresses, never the overlapping
private ones.</li>
<li>Same ordering rules apply to the hub VPC (VPC → attachment → LB →
namespace → subnets → wait for image sync → workloads).</li>
<li>Add a gateway firewall policy on the hub. Reachability is not
authorisation.</li>
</ul>
<p><em>Previously: <a href="/posts/three-datacenters-one-ip-plan/">three datacenters, one IP plan</a>.
Next: the whole pod as a single catalog item.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Output captured live, trimmed for length,
never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>A datacenter in a catalog tile: nested ESXi pods via VCF Automation All Apps</title>
      <link>https://thenestedlab.com/posts/nested-esxi-via-vcfa-all-apps/</link>
      <pubDate>Wed, 16 Sep 2026 07:40:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/nested-esxi-via-vcfa-all-apps/</guid>
      <description>The whole isolated pod — namespace, trunk subnets, binding maps, two dual-NIC nested ESXi hosts with an ISO attached, SSH/HTTPS VIPs — as one VCF Automation blueprint, published to the catalog. Anatomy of the blueprint, the ordering it enforces, and the three things it can&amp;rsquo;t express.</description>
      <content:encoded><![CDATA[<p>Everything in this series so far was built with <code>kubectl</code> and API calls.
That proves the platform. It doesn&rsquo;t make a <em>product</em>. This post turns the
pod into a <strong>catalog item</strong>: fill in a name, pick a VPC, click Request, and
a few minutes later there&rsquo;s a datacenter-in-miniature with two SSH prompts
waiting.</p>
<p><img alt="VCFA catalog: the nested-esxi-pod tile" loading="lazy" src="/images/ui/u1-catalog-tile.jpg"></p>
<h2 id="all-apps-in-one-paragraph">All Apps in one paragraph</h2>
<p>VCF Automation 9.1 has two provisioning models side by side. <strong>VM Apps</strong> is
the classic Aria Automation path — cloud templates through an IaaS engine
that drives vCenter. <strong>All Apps</strong> is the supervisor-native path: the
blueprint composes Kubernetes objects (a Supervisor Namespace, VM Service
VMs, NSX subnets, VKS clusters) and the vSphere Supervisor&rsquo;s controllers
reconcile them. A blueprint is <code>formatVersion: 2</code>; its resources are
<code>CCI.Supervisor.Namespace</code> and <code>CCI.Supervisor.Resource</code> — the latter is
literally &ldquo;here&rsquo;s a manifest, apply it in that namespace.&rdquo;</p>
<p>That makes the blueprint a <em>composition</em> of the manifests from the earlier
posts, with two additions: inputs, and <code>dependsOn</code>.</p>
<h2 id="the-blueprint-section-by-section">The blueprint, section by section</h2>
<h3 id="inputs--the-form">Inputs — the form</h3>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">inputs</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">podName</span><span class="p">:</span><span class="w">  </span>{<span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="nt">string, default</span><span class="p">:</span><span class="w"> </span><span class="nt">nested-pod, pattern</span><span class="p">:</span><span class="w"> </span><span class="s1">&#39;^[a-z0-9]([-a-z0-9]*[a-z0-9])?$&#39;</span>}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vpcName</span><span class="p">:</span><span class="w">  </span>{<span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="nt">string, description</span><span class="p">:</span><span class="w"> </span><span class="l">Must exist and be Realized before deploying.}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">esxOva</span><span class="p">:</span><span class="w">   </span>{<span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="nt">string, default</span><span class="p">:</span><span class="w"> </span><span class="l">vmi-61bb062ddfc506b79}  </span><span class="w"> </span><span class="c"># Nested ESXi 9.1 appliance</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">isoImage</span><span class="p">:</span><span class="w"> </span>{<span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="nt">string, default</span><span class="p">:</span><span class="w"> </span><span class="l">vmi-39f562e2ae9e9c501}  </span><span class="w"> </span><span class="c"># the ISO to attach</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vmClass</span><span class="p">:</span><span class="w">  </span>{<span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="nt">string, default</span><span class="p">:</span><span class="w"> </span><span class="nt">best-effort-large, enum</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="l">best-effort-large, best-effort-xlarge, best-effort-2xlarge]}</span><span class="w">
</span></span></span></code></pre></div><p><img alt="The request form" loading="lazy" src="/images/ui/u2-request-form.jpg"></p>
<h3 id="the-namespace--with-libraries-attached">The namespace — with libraries attached</h3>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">namespace</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="l">CCI.Supervisor.Namespace</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">properties</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">generateName</span><span class="p">:</span><span class="w"> </span><span class="l">${input.podName}-       </span><span class="w"> </span><span class="c"># NOT name — new namespaces get a suffix</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">className</span><span class="p">:</span><span class="w"> </span><span class="l">large</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">regionName</span><span class="p">:</span><span class="w"> </span><span class="l">f06</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">vpcName</span><span class="p">:</span><span class="w"> </span><span class="l">${input.vpcName}             </span><span class="w"> </span><span class="c"># pins the namespace to the pod&#39;s VPC</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">storageClasses</span><span class="p">:</span><span class="w"> </span><span class="p">[</span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="nt">vSAN Default Storage Policy, limit</span><span class="p">:</span><span class="w"> </span><span class="l">400000Mi}]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">zones</span><span class="p">:</span><span class="w"> </span><span class="p">[</span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="nt">domain-c9, cpuLimit</span><span class="p">:</span><span class="w"> </span><span class="nt">40000M, memoryLimit</span><span class="p">:</span><span class="w"> </span><span class="l">64000Mi, ...}]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">contentSources</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="nt">ISO, type</span><span class="p">:</span><span class="w"> </span><span class="l">ContentLibrary}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="nt">f06-vks-lib01, type</span><span class="p">:</span><span class="w"> </span><span class="l">ContentLibrary}</span><span class="w">
</span></span></span></code></pre></div><p><code>contentSources</code> is the line that closes the gap a lot of first attempts
hit: a VCFA-created namespace has <strong>no content library</strong>, so there are no
<code>VirtualMachineImage</code>s and nothing can be deployed. Declaring the libraries
here attaches them at creation.</p>
<h3 id="the-topology--ordered-on-purpose">The topology — ordered on purpose</h3>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">snTrunk</span><span class="p">:</span><span class="w">   </span>{<span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="nt">CCI.Supervisor.Resource, properties</span><span class="p">:</span><span class="w"> </span>{<span class="nt">context</span><span class="p">:</span><span class="w"> </span><span class="l">${resource.namespace.id}, manifest: &lt;Subnet sn-trunk&gt;}}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">snMgmt</span><span class="p">:</span><span class="w">    </span>{<span class="nt">dependsOn</span><span class="p">:</span><span class="w"> </span><span class="nt">[snTrunk], ...  manifest</span><span class="p">:</span><span class="w"> </span><span class="l">&lt;Subnet sn-mgmt&gt;}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">snVmotion</span><span class="p">:</span><span class="w"> </span>{<span class="nt">dependsOn</span><span class="p">:</span><span class="w"> </span><span class="nt">[snMgmt],  ...  manifest</span><span class="p">:</span><span class="w"> </span><span class="l">&lt;Subnet sn-vmotion&gt;}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">bmMgmt</span><span class="p">:</span><span class="w">    </span>{<span class="nt">... manifest</span><span class="p">:</span><span class="w"> </span><span class="l">&lt;SubnetConnectionBindingMap sn-mgmt -&gt; sn-trunk, vlan 1610&gt;}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">bmVmotion</span><span class="p">:</span><span class="w"> </span>{<span class="nt">... manifest</span><span class="p">:</span><span class="w"> </span><span class="l">&lt;SubnetConnectionBindingMap sn-vmotion -&gt; sn-trunk, vlan 1611&gt;}</span><span class="w">
</span></span></span></code></pre></div><p>The <code>dependsOn</code> chain is the whole reason <a href="/posts/three-datacenters-one-ip-plan/">every pod has identical
CIDRs</a>: fresh VPCs realize subnets
in creation order, and the blueprint fixes that order.</p>
<p><img alt="Blueprint canvas and YAML side by side" loading="lazy" src="/images/ui/u3-blueprint-canvas-yaml.jpg"></p>
<h3 id="the-hosts--dual-nic-iso-attached-bootstrapped-by-ovf">The hosts — dual-NIC, ISO attached, bootstrapped by OVF</h3>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">esx01</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">type</span><span class="p">:</span><span class="w"> </span><span class="l">CCI.Supervisor.Resource</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">dependsOn</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="l">bmMgmt]                   </span><span class="w"> </span><span class="c"># no point booting before VLAN 1610 exists</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">properties</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">manifest</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">VirtualMachine               </span><span class="w"> </span><span class="c"># vmoperator.vmware.com/v1alpha5</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">hardware</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">cdrom</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w"> </span><span class="l">... the ISO, declared, connected ... ]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">network</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">interfaces</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w"> </span><span class="l">eth0 -&gt; sn-trunk, eth1 -&gt; sn-trunk ]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">vAppConfig</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w"> </span><span class="l">guestinfo.hostname / ipaddress / vlan / ... ]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">wait</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">fields</span><span class="p">:</span><span class="w"> </span><span class="p">[</span>{<span class="nt">path</span><span class="p">:</span><span class="w"> </span><span class="nt">status.powerState, value</span><span class="p">:</span><span class="w"> </span><span class="l">PoweredOn}]</span><span class="w">
</span></span></span></code></pre></div><p>(Abridged — the full resource carries the image references, VM class,
guest ID and the complete <code>guestinfo</code> set.)</p>
<p>Two vNICs, both on the trunk — <a href="/series/the-vpc-pod-papers/">the nested equivalent of a VCF host&rsquo;s two
pNICs</a>. The ISO rides along as a declarative
CD-ROM. And the <code>wait</code> block makes the deployment&rsquo;s <em>completion</em> mean
something: the request doesn&rsquo;t finish until the host is powered on.</p>
<h3 id="the-doors--one-vip-per-host">The doors — one VIP per host</h3>
<p>A <code>VirtualMachineService</code> of type <code>LoadBalancer</code> per host, selecting it by
label and publishing 22 and 443, and a blueprint <strong>output</strong> that reads the
VIP back out of the service&rsquo;s status:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">outputs</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">esx01Ssh</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;ssh root@${resource.esx01Access.object.status.loadBalancer.ingress[0].ip}&#34;</span>}<span class="w">
</span></span></span></code></pre></div><p>The outputs surface in the deployment view — the requester gets the SSH
command, not a scavenger hunt.</p>
<p><img alt="Deployment topology after a successful request" loading="lazy" src="/images/ui/u4-deployment-topology.jpg"></p>
<p><img alt="Request → deployment in progress → complete" loading="lazy" src="/images/u7-catalog-request-flow.gif">
<em>The request flow, end to end.</em></p>
<h2 id="what-the-blueprint-cannot-express-yet">What the blueprint cannot express (yet)</h2>
<p>Three cluster-scoped objects have <strong>no blueprint resource type</strong>, and they
must exist <em>before</em> the request — <a href="/posts/the-lb-that-must-exist-first/">in this order</a>:</p>
<ol>
<li><code>VPC</code> — <code>privateIPs: 172.30.0.0/16</code>, same in every pod</li>
<li><code>VPCAttachment</code> — connectivity profile with the service gateway; the LB
creation fails loudly without it</li>
<li><code>LoadBalancer</code> — silently, permanently required before the namespace</li>
</ol>
<p>Today that&rsquo;s a short script or a runbook step per pod. The honest framing:
the blueprint is the <em>pod</em>; the VPC is the <em>tenancy</em>, and tenancy is
still created one layer up. I&rsquo;d expect that layer to become blueprintable;
until then, keep the three calls next to the blueprint in version control.</p>
<h2 id="publishing-one-version-at-a-time">Publishing: one version at a time</h2>
<p>Blueprint → <code>BlueprintVersion</code> → release. Validation happens at <em>version</em>
time, not create time, and the result lives in <code>status.validationMessages</code>
rather than the HTTP code — a 200 with <code>ContentValid: False</code> is a thing.
And only <strong>one</strong> version can be published: unrelease 1.0.0 before releasing
1.1.0, or you get a 409. (The full list of sharp edges is
<a href="/series/the-vpc-pod-papers/">its own post</a>.)</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>This is where platform engineering turns into a service. The difference
between &ldquo;we can build you an environment&rdquo; and &ldquo;request one from the
catalog&rdquo; is the difference between days and minutes — and between a
bespoke build and one that is consistent, quota-controlled and recorded
every time. For an organisation that means:</p>
<ul>
<li><strong>Time-to-environment</strong> measured in minutes, requested by the people who
need it, without a queue.</li>
<li><strong>Consistency by construction</strong> — every environment comes from the same
definition, so support, training material and runbooks all match.</li>
<li><strong>Governance built in</strong> — quotas, ownership, history and clean teardown
are properties of the deployment record, not a spreadsheet.</li>
</ul>
<p>The nested-ESXi pod is one catalog item. The same approach delivers any
environment shape: application stacks for developers, sandboxes for a
proof of concept, demo kits for a sales team, isolated builds for a
partner.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>All Apps blueprints are <strong>compositions of manifests</strong>: <code>CCI.Supervisor.Namespace</code>
plus <code>CCI.Supervisor.Resource</code> per object. If it works with <code>kubectl</code>, it
works in a blueprint.</li>
<li><code>generateName</code>, not <code>name</code>, for the namespace; <code>contentSources</code> to attach
libraries at creation; <code>zones</code>/<code>storageClasses</code> flat, not wrapped.</li>
<li><code>dependsOn</code> is how you get <strong>deterministic CIDRs</strong> — order the subnets.</li>
<li><code>wait.fields</code> turns &ldquo;request complete&rdquo; into &ldquo;host is powered on&rdquo;.</li>
<li>VPC / VPCAttachment / LoadBalancer are <strong>prerequisites outside the
blueprint</strong>, in that order, before every request.</li>
<li>One published version per blueprint; validation in <code>status</code>, not the
HTTP response.</li>
</ul>
<p><em>Previously: <a href="/posts/shared-services-for-isolated-tenants/">shared services for isolated tenants</a>.
This closes the Pod Papers&rsquo; core arc — the companion posts on
<a href="/series/the-vpc-pod-papers/">dual-NIC</a>, <a href="/series/the-vpc-pod-papers/">no-DHCP bootstrap</a>
and <a href="/series/the-vpc-pod-papers/">blueprint gotchas</a> fill in the details.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Blueprint <code>nested-esxi-pod</code> 1.1.0 is
live in the lab catalog; YAML above trimmed for length.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Dual-NIC nested hosts: what redundancy means when the fabric is virtual</title>
      <link>https://thenestedlab.com/posts/dual-nic-nested-hosts/</link>
      <pubDate>Wed, 16 Sep 2026 07:30:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/dual-nic-nested-hosts/</guid>
      <description>VCF wants two pNICs per host. In a nested lab the second vNIC adds no physical redundancy — so why add it? Because bringup validation and uplink teaming expect it, and because the failover test tells you something real about the trunk. vmnic0 down, 0% loss, and the SSH session watching it never dropped.</description>
      <content:encoded><![CDATA[<p>&ldquo;Naturally, a VCF host has at least two NICs. Are we testing that, or have
you virtualised it away?&rdquo;</p>
<p>Fair question, and the honest answer has two halves. In a nested lab the
<em>physical</em> redundancy is provided by the outer host — its vDS, its NSX
uplinks — and a second vNIC on the nested VM adds precisely none. But VCF
doesn&rsquo;t know it&rsquo;s nested. Bringup&rsquo;s host validation and the vDS uplink
teaming it configures <strong>expect two vmnics</strong>, and a host with one gets
flagged. So the nested hosts get two vNICs, both on the trunk subnet, and
the question becomes: does failover between them actually work inside a
VPC?</p>
<h2 id="the-setup">The setup</h2>
<p>Both vNICs attach to the same <code>sn-trunk</code> subnet — <a href="/posts/nested-esxi-nsx-vpc/">the trunk from part
1</a> — and ESXi sees them as two 10G vmnics:</p>
<p><img alt="Host Client: vmnic0 and vmnic1, both 10 Gbit/s on vSwitch0" loading="lazy" src="/images/ui/u13-hostclient-dual-nics.jpg"></p>
<p>vSwitch0 teams them active/active with the default originating-port-ID
policy; every portgroup (Management 1610, vMotion 1611, vSAN 1612) inherits
it. Nothing you wouldn&rsquo;t do on metal.</p>
<h2 id="the-test-pull-a-nic-while-watching-from-inside">The test: pull a NIC while watching from inside</h2>
<p>The interesting bit isn&rsquo;t whether pings continue — it&rsquo;s <em>which</em> session
I&rsquo;m watching from. I&rsquo;m SSH&rsquo;d to <code>esx01</code> <strong>through its public VIP</strong>, which
means my session traverses the NSX LB → the VPC → the trunk port → whichever
vmnic happens to carry vmk0. If failover breaks anything, it breaks the
terminal I&rsquo;m typing in.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[root@esx01-a:~] esxcli network nic list
</span></span><span class="line"><span class="cl">Name    ...  Admin Status  Link Status  Speed  MAC Address
</span></span><span class="line"><span class="cl">vmnic0  ...  Up            Up           10000  04:50:56:00:5c:03
</span></span><span class="line"><span class="cl">vmnic1  ...  Up            Up           10000  04:50:56:00:68:00
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] esxcli network nic down -n vmnic0    # FAIL THE FIRST NIC
</span></span><span class="line"><span class="cl">vmnic0  ...  Down          Down             0  04:50:56:00:5c:03
</span></span><span class="line"><span class="cl">vmnic1  ...  Up            Up           10000  04:50:56:00:68:00
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] vmkping -I vmk1 172.30.0.71          # vMotion VLAN, now over vmnic1
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] vmkping -I vmk2 172.30.0.101         # vSAN VLAN, now over vmnic1
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] esxcli network nic up -n vmnic0      # restore
</span></span><span class="line"><span class="cl"># session never dropped.
</span></span></code></pre></div><p><img alt="Before / after: vmnic0 down, every VLAN still passing" loading="lazy" src="/images/c6-nic-failover-beforeafter.gif"></p>
<p><img alt="Full failover transcript" loading="lazy" src="/images/demo-c6-nic-failover.jpg"></p>
<p>Every VLAN moved to vmnic1. Zero loss on vMotion and vSAN. And the
management session — the one <em>most</em> likely to notice — never blinked.</p>
<h2 id="why-this-is-a-real-result-not-a-party-trick">Why this is a real result, not a party trick</h2>
<p>Think about what just happened at the NSX layer. vmk0&rsquo;s MAC — a MAC ESXi
synthesised, not the vNIC&rsquo;s — was being learned on trunk port A. When
vmnic0 went down, the same MAC appeared on trunk port B mid-flow, with an
established TCP session riding on it.</p>
<p>On a <strong>standard</strong> VPC subnet port that is exactly the scenario SpoofGuard
exists to stop: the port&rsquo;s address bindings pin one MAC, and a frame from a
different MAC — or the <em>same</em> MAC arriving on a different port — is dropped.
Part 1 showed that killing the host on a standard subnet before it ever
spoke. This test shows the trunk subnet tolerating the live migration of a
foreign MAC between two of its ports, which is the property nested vSphere
(and anything else with a vSwitch inside a VM) fundamentally needs.</p>
<p>So the second vNIC buys three things, none of them physical redundancy:</p>
<ol>
<li><strong>Bringup and vLCM stop complaining</strong> about a single-uplink host.</li>
<li><strong>The teaming policy you&rsquo;ll configure in production gets exercised</strong> —
uplink failover, active/standby for vSAN, whatever you&rsquo;re rehearsing.</li>
<li><strong>A live proof that the trunk carries MAC mobility</strong>, which is the
real assurance that the design isn&rsquo;t relying on a quiet network.</li>
</ol>
<h2 id="what-it-does-not-buy-and-how-to-say-so">What it does <em>not</em> buy, and how to say so</h2>
<p>If someone asks &ldquo;is this host redundant?&rdquo;, the answer is &ldquo;the nested host
believes it is; actual redundancy lives one layer down.&rdquo; In a training pod
that&rsquo;s the correct and useful answer — students configure and test failover
exactly as they would on metal, and the outer platform does the real work.
In a reproduction lab for a customer NIC-teaming issue, it&rsquo;s usually enough
too: most teaming bugs are in ESXi&rsquo;s policy handling, not in the copper.</p>
<p>Where it&rsquo;s genuinely insufficient: anything about physical link behaviour —
LACP negotiation, LLDP, flapping, MTU mismatch on one uplink. The virtual
fabric never fails asymmetrically, so it can&rsquo;t reproduce those.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>The practical value here is knowing <em>what a nested environment can and
can&rsquo;t prove</em> — which is what lets you decide when a virtual lab is enough
and when it isn&rsquo;t. For training, upgrade rehearsals, configuration and
policy testing and the vast majority of &ldquo;how does it behave when…&rdquo;
questions, nested is enough and dramatically cheaper. For physical link
behaviour — LACP, optics, asymmetric faults — you still want metal. Being
able to make that call confidently is worth more than the test itself.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Give nested VCF hosts <strong>two vNICs on the same trunk subnet</strong>. Bringup,
vLCM and vDS teaming expect ≥ 2 vmnics; humouring them costs nothing.</li>
<li>Test failover <strong>from a session that depends on it</strong> (SSH via the VIP).
Pings passing while your terminal dies is not success.</li>
<li>The trunk subnet tolerates a <strong>vmk MAC moving between ports mid-flow</strong>
— that&rsquo;s the property standard subnets lack and nested vSphere needs.</li>
<li>Be precise in the write-up: nested dual-NIC gives <em>policy</em> realism, not
<em>physical</em> redundancy. Physical link faults can&rsquo;t be reproduced here.</li>
<li>Rebuilds re-run the vmk config: the appliance creates vmk0 only;
vmk1/vmk2, their VLANs and override gateways are applied post-boot.</li>
</ul>
<p><em>Companion to <a href="/posts/nested-esxi-nsx-vpc/">nested ESXi inside an NSX VPC</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Output captured live, trimmed for length,
never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>VPC subnets have no DHCP — and that&#39;s fine</title>
      <link>https://thenestedlab.com/posts/vpc-subnets-have-no-dhcp/</link>
      <pubDate>Wed, 16 Sep 2026 07:20:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/vpc-subnets-have-no-dhcp/</guid>
      <description>The nested-ESXi appliance sat at &amp;lsquo;waiting for DHCP&amp;rsquo; forever. VPC subnets don&amp;rsquo;t hand out addresses — the VM Service does, through bootstrap providers. cloud-init for Linux, sysprep for Windows, OVF guestinfo for appliances, and the per-vmk gateway detail that makes the Host Client tell the truth.</description>
      <content:encoded><![CDATA[<p>The second trap from <a href="/posts/nested-esxi-nsx-vpc/">part 1</a> deserves its
own short post, because it catches everything, not just ESXi: <strong>a VPC
subnet has DHCP deactivated.</strong> Drop a stock appliance onto one and it will
boot, sit at &ldquo;waiting for DHCP&rdquo;, and wait politely until the heat death of
the universe.</p>
<p>This isn&rsquo;t a gap. It&rsquo;s the model: NSX allocates the address at the <em>port</em>
and pins it there with address bindings; the <em>guest</em> has to be told what
it was given. The VM Service does that telling through <strong>bootstrap
providers</strong>, and once you know the three of them, static addressing stops
being a chore and starts being a feature.</p>
<h2 id="three-providers-three-guest-types">Three providers, three guest types</h2>
<table>
	<thead>
			<tr>
					<th>Guest</th>
					<th>Provider</th>
					<th>Carries</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Linux</td>
					<td><code>cloudInit</code></td>
					<td>user-data (users, <code>write_files</code>, <code>runcmd</code>) + network config</td>
			</tr>
			<tr>
					<td>Windows</td>
					<td><code>sysprep</code></td>
					<td>unattend XML / sysprep spec, identity, network</td>
			</tr>
			<tr>
					<td>Appliances (OVF)</td>
					<td><code>vAppConfig</code></td>
					<td>OVF properties (<code>guestinfo.*</code>) the appliance reads on boot</td>
			</tr>
	</tbody>
</table>
<p>All three are <strong>typed fields on the <code>VirtualMachine</code> object</strong>, not bolt-on
customisation specs. The network side is already known to the platform —
the VM Service knows which subnet each interface landed on and what NSX
allocated — so for cloud-init and sysprep the addressing is injected for
you. Appliances are the exception, because each one has its own idea of
which properties it wants.</p>
<h2 id="appliances-vappconfig">Appliances: vAppConfig</h2>
<p>The nested-ESXi appliance reads <code>guestinfo.*</code> OVF properties. In the VM
Service spec that&rsquo;s:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vAppConfig</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">properties</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.hostname,  value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;esx01.pod-a.res.lab&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.ipaddress, value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.40&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.netmask,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;255.255.255.224&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.gateway,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.33&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.vlan,      value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;1610&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.dns,       value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;10.20.52.1&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.ssh,       value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;True&#34;</span>}}<span class="w">
</span></span></span></code></pre></div><p>The one that trips people: <strong>the address you give must be the one NSX
allocated to the port.</strong> In a standard subnet that&rsquo;s enforced by
SpoofGuard; on a trunk subnet the VLAN subnets have their own allocations
and you&rsquo;re choosing addresses within them. Either way, pick from the
realized range — and remember <a href="/posts/nested-esxi-nsx-vpc/">recreating a VM reallocates its
addresses</a>, so the fixed <code>.40</code>/<code>.41</code> in a
<a href="/posts/three-datacenters-one-ip-plan/">deterministic pod</a> is a design
choice, not luck.</p>
<h2 id="linux-cloud-init">Linux: cloud-init</h2>
<p>A Secret holding user-data, referenced from the VM:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">cloudInit</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">cloudConfig</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">users</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w"> </span><span class="l">... a local user with a key ... ]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">write_files</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span>- <span class="nt">path</span><span class="p">:</span><span class="w"> </span><span class="l">/var/www/html/index.html</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">content</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;shared-svc repo01\n&#34;</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">runcmd</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span>- <span class="p">[</span><span class="l">systemctl, enable, --now, nginx]</span><span class="w">
</span></span></span></code></pre></div><p>Networking arrives via the platform&rsquo;s own network-config; you don&rsquo;t
write it. The <code>svc-repo01</code> VM from <a href="/posts/shared-services-for-isolated-tenants/">the shared-services
post</a> was exactly this —
<code>write_files</code> + <code>runcmd</code>, web page verified from three pods.</p>
<h2 id="windows-sysprep">Windows: sysprep</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">sysprep</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">sysprep</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">guiUnattended</span><span class="p">:</span><span class="w"> </span>{<span class="nt">autoLogon</span><span class="p">:</span><span class="w"> </span><span class="nt">true, autoLogonCount</span><span class="p">:</span><span class="w"> </span><span class="nt">1, timeZone</span><span class="p">:</span><span class="w"> </span><span class="m">85</span>}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">identification</span><span class="p">:</span><span class="w"> </span>{<span class="nt">joinWorkgroup</span><span class="p">:</span><span class="w"> </span><span class="l">WORKGROUP}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">userData</span><span class="p">:</span><span class="w"> </span>{<span class="nt">fullName</span><span class="p">:</span><span class="w"> </span><span class="nt">Lab, orgName</span><span class="p">:</span><span class="w"> </span><span class="nt">Lab, computerName</span><span class="p">:</span><span class="w"> </span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">win01}}</span><span class="w">
</span></span></span></code></pre></div><p>Or <code>rawSysprep</code> with an unattend XML in a Secret if you already have one.
The ISO from <a href="/posts/nested-esxi-via-vcfa-all-apps/">the blueprint post</a>
can ride along as a declarative <code>hardware.cdrom</code> — handy for tools and
agents on first boot.</p>
<h2 id="the-esxi-footnote-one-gateway-many-vmks">The ESXi footnote: one gateway, many vmks</h2>
<p>Once the appliance is up and you add vMotion and vSAN vmks on their own
subnets, you hit a detail that makes the Host Client <em>look</em> wrong:</p>
<p><img alt="Host Client: vmk0/1/2, one service each" loading="lazy" src="/images/ui/u12-hostclient-vmk-adapters.jpg"></p>
<p>ESXi&rsquo;s default TCP/IP stack has <strong>one</strong> default gateway — vmk0&rsquo;s <code>.33</code> —
and the UI repeats it on every vmk row. Same-subnet vMotion never uses a
gateway so nothing breaks, but each NSX subnet <em>does</em> have its own
gateway, and cross-subnet traffic from vmk1/vmk2 would take the wrong exit.
Set per-vmk override gateways:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">esxcli network ip interface ipv4 set -i vmk1 -t static -I 172.30.0.70  -N 255.255.255.224 -g 172.30.0.65
</span></span><span class="line"><span class="cl">esxcli network ip interface ipv4 set -i vmk2 -t static -I 172.30.0.100 -N 255.255.255.224 -g 172.30.0.97
</span></span></code></pre></div><p>Now the display is truthful and the routing is correct. (The full-realism
alternative is a dedicated <code>vmotion</code> netstack for vmk1; I kept the default
stack so the service tags stay visible in the Host Client.)</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>For the business, &ldquo;no DHCP&rdquo; translates into something security and
operations teams both want: <strong>predictable addressing</strong>. Every environment
has a known address plan, firewall rules can be written once, and nothing
turns up on the network with an address nobody expected. Bootstrap
providers deliver the second benefit — images stay generic and
configuration is injected at deploy time, so there are fewer golden images
to maintain and far less drift between environments.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li><strong>No DHCP in VPC subnets</strong> — by design. NSX allocates at the port; the
guest is told via a bootstrap provider.</li>
<li><code>cloudInit</code> (Linux), <code>sysprep</code> (Windows), <code>vAppConfig</code> (appliances) —
typed fields on the VM, not customisation specs.</li>
<li>Appliance addresses must match the <strong>realized</strong> subnet; fix the order
of subnet creation if you want fixed addresses across pods.</li>
<li>ESXi has <strong>one</strong> default gateway per stack. Set <code>-g</code> per vmk, or the
Host Client lies to you and cross-subnet traffic exits wrong.</li>
<li>A static IP plan is a feature in a lab: it&rsquo;s what makes screenshots,
runbooks and pods identical.</li>
</ul>
<p><em>Companion to <a href="/posts/nested-esxi-nsx-vpc/">nested ESXi inside an NSX VPC</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
