<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>vSphere and vSAN on The Nested Lab</title>
    <link>https://thenestedlab.com/products/vsphere-and-vsan/</link>
    <description>Recent content in vSphere and vSAN on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/products/vsphere-and-vsan/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Nested ESXi inside an NSX VPC: the trunk-subnet design</title>
      <link>https://thenestedlab.com/posts/nested-esxi-nsx-vpc/</link>
      <pubDate>Wed, 16 Sep 2026 08:20:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/nested-esxi-nsx-vpc/</guid>
      <description>Plain VPC subnets silently blackhole a nested ESXi host. Why, and the trunk subnet and binding map design that makes nested labs work as an ordinary NSX VPC tenant.</description>
      <content:encoded><![CDATA[<p>The host booted clean. Management IP configured, services up, DCUI happy.
And every single packet it sent, ARP included, died silently. The host was
having a lovely time; it just couldn&rsquo;t tell anyone.</p>
<p>That&rsquo;s how my first attempt at running nested ESXi inside an NSX VPC ended.
The failure is nasty precisely because nothing <em>looks</em> wrong. If you&rsquo;re
building nested vSphere labs on VCF 9 with VPC networking, this post is the
map of the minefield. It&rsquo;s also the design that gets you across it, verified
live.</p>
<h2 id="the-setup">The setup</h2>
<p>VCF 9.1, with the vSphere Supervisor on NSX VPC networking. The goal was to
deploy nested ESXi hosts as ordinary VM Service VMs inside a tenant&rsquo;s VPC.
No physical fabric changes, no provider tickets, no special treatment. It&rsquo;s
the kind of thing you want for training pods, cert-study labs, or
reproducing customer issues.</p>
<p>Nested ESXi needs what physical ESXi needs: a management network, vMotion
and vSAN. Traditionally, those are VLANs trunked to every host. But a VPC is
an overlay world, and there are no VLANs to trunk. So what happens if you
just attach the nested host&rsquo;s vNIC to a normal VPC subnet?</p>
<h2 id="failure-1-the-silent-blackhole">Failure #1: the silent blackhole</h2>
<p>Here&rsquo;s the trap. A standard VPC subnet port gets <strong>address bindings</strong>. NSX
pins the exact IP and MAC it allocated to that vNIC, and SpoofGuard drops
everything else.</p>
<p>ESXi&rsquo;s vmk0 doesn&rsquo;t use the vNIC&rsquo;s MAC. It makes up its own:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">vmk0
</span></span><span class="line"><span class="cl">   MAC Address: 00:50:ac:1e:00:8c     &lt;- NOT the vNIC MAC (04:50:56:...)
</span></span></code></pre></div><p>So every frame the management interface sends carries a MAC the port doesn&rsquo;t
own. NSX drops all of it (ARP, ping, everything), while the host itself boots
green and reports healthy. There&rsquo;s no error anywhere, which is somehow worse
than a bad one. You just can&rsquo;t reach it, ever.</p>
<p><img alt="Standard VPC subnet port: SpoofGuard pins one IP+MAC; vmk0&rsquo;s synthesised MAC loses, silently" loading="lazy" src="/images/post1-blackhole.svg"></p>
<p>(There&rsquo;s a second trap stacked on top. Our VPC subnets run with DHCP
deactivated, so the appliance also sits at &ldquo;waiting for DHCP&rdquo; unless you
inject static addressing through OVF <code>guestinfo.*</code> properties. NSX can give a
VPC subnet a DHCP server or relay (<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/virtual-private-cloud-in-nsx/virtual-private-clouds-overview/add-a-subnet-for-the-vpc.html">Add a Subnet to a VPC</a>);
ours had neither. More on that below.)</p>
<h2 id="the-design-that-works-a-trunk-subnet--binding-maps">The design that works: a trunk subnet + binding maps</h2>
<p>The fix isn&rsquo;t a hack. It&rsquo;s a first-class NSX VPC construct that&rsquo;s barely
documented in the wild: <strong><code>SubnetConnectionBindingMap</code></strong>.</p>
<p>The idea:</p>
<ol>
<li>Create one ordinary VPC subnet to act as a <strong>trunk</strong> (<code>sn-trunk</code>). The
nested host&rsquo;s vNICs attach <em>only</em> here.</li>
<li>Create a normal VPC subnet for each traditional network: <code>sn-mgmt</code>,
<code>sn-vmotion</code> and <code>sn-vsan</code>.</li>
<li>Bind each of those to the trunk with a <strong>binding map carrying a VLAN tag</strong>.
The nested host&rsquo;s vSwitch tags frames exactly as it would on metal. The
binding map strips the tag and delivers the frame into the right subnet.</li>
</ol>
<p>It&rsquo;s pure L2 demultiplexing. One vNIC carries N VLANs, and the VPC never
routes on a tag. The physical fabric never sees any of it, because the
802.1Q header rides inside the Geneve overlay.</p>
<p>A tenant can create all of it through the supervisor, as Kubernetes objects:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="c"># sn-trunk and sn-mgmt are ordinary Private Subnets; the interesting object:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">apiVersion</span><span class="p">:</span><span class="w"> </span><span class="l">crd.nsx.vmware.com/v1alpha1</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">kind</span><span class="p">:</span><span class="w"> </span><span class="l">SubnetConnectionBindingMap</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">metadata</span><span class="p">:</span><span class="w"> </span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">bm-mgmt}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="nt">spec</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">subnetName</span><span class="p">:</span><span class="w"> </span><span class="l">sn-mgmt         </span><span class="w"> </span><span class="c"># the map is a child of the VLAN subnet...</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">targetSubnetName</span><span class="p">:</span><span class="w"> </span><span class="l">sn-trunk  </span><span class="w"> </span><span class="c"># ...and points AT the trunk</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vlanTrafficTag</span><span class="p">:</span><span class="w"> </span><span class="m">1610</span><span class="w">
</span></span></span></code></pre></div><p>That direction is easy to get backwards, so it&rsquo;s worth saying twice: <strong>the
binding map belongs to the VLAN subnet and points at the trunk</strong>, not the
other way round.</p>
<p><img alt="NSX: sn-trunk realized once per VPC, binding maps hanging off the VLAN subnets" loading="lazy" src="/images/ui/u11b-nsx-sntrunk-per-vpc.jpg"></p>
<p>On the nested host, nothing exotic: plain virtual switch tagging (VST), just
like physical. It never suspects a thing.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Name                Virtual Switch  Active Clients  VLAN ID
</span></span><span class="line"><span class="cl">------------------  --------------  --------------  -------
</span></span><span class="line"><span class="cl">Management Network  vSwitch0                     1     1610
</span></span><span class="line"><span class="cl">vMotion             vSwitch0                     1     1611
</span></span><span class="line"><span class="cl">vSAN                vSwitch0                     1     1612
</span></span></code></pre></div><p><img alt="Host Client: port groups on VLANs 1610 / 1611 / 1612" loading="lazy" src="/images/ui/u12a-hostclient-portgroups-vlans.jpg">
<em>The same three VLANs as the nested host sees them.</em></p>
<p>And because there&rsquo;s no DHCP in our VPC subnets, the nested-ESXi appliance
gets its identity through OVF properties in the VM Service spec:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vAppConfig</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">properties</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.ipaddress, value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.40&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.netmask,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;255.255.255.224&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.gateway,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.33&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.vlan,     value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;1610&#34;</span>}}<span class="w">
</span></span></span></code></pre></div><h2 id="does-it-actually-work-the-receipts">Does it actually work? The receipts</h2>
<p>Two nested hosts, vNICs on <code>sn-trunk</code>, three VLANs. From host one:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[root@esx01:~] vmkping -c2 172.30.0.41            # mgmt, VLAN 1610
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">[root@esx01:~] vmkping -I vmk1 -c3 172.30.0.71    # vMotion, VLAN 1611
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">[root@esx01:~] vmkping -I vmk2 -c3 172.30.0.101   # vSAN, VLAN 1612
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span></code></pre></div><p><img alt="Live capture: vmnic0 down, vMotion and vSAN VLANs still passing at 0% loss" loading="lazy" src="/images/demo-c6-nic-failover.jpg">
<em>The transcript that matters: fail the first NIC, and every VLAN keeps flowing on the second — captured live.</em></p>
<p>Two more results are worth knowing before you design around this.</p>
<p><strong>Untagged frames are dropped.</strong> I put a probe vmk on the untagged port
group, using the address NSX itself had allocated to the trunk port. The
result was 100% loss and an empty ARP table, while tagged traffic flowed
happily beside it, as if to make a point. Every network your nested host
uses needs a VLAN and a binding map. There is no untagged fallback.</p>
<p><strong>Failover behaves like real hardware.</strong> With two vNICs on the trunk, teamed
active/active, <code>esxcli network nic down -n vmnic0</code> moved every VLAN onto
vmnic1 with zero loss. The SSH session I was watching from never dropped.
A vmk MAC moving between trunk ports mid-flow is exactly what MAC-pinned
standard ports would blackhole. The trunk carries it fine.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Running whole vSphere environments <em>inside</em> a VPC turns the platform into
something most customers never had. It&rsquo;s a way to stand up complete,
isolated copies of infrastructure on demand, with no physical fabric change
and no waiting for anyone. That&rsquo;s what makes it commercially interesting:</p>
<ul>
<li><strong>Training and certification labs</strong> where every learner gets a real
vSphere environment, not a shared one.</li>
<li><strong>Reproducing a customer problem</strong> on a like-for-like copy, rather than on
their own estate.</li>
<li><strong>Rehearsing upgrades and migrations</strong> end to end before the change
window, then throwing the copy away.</li>
<li><strong>Vendor and feature evaluations</strong> with real behaviour, at zero risk to
production.</li>
</ul>
<p>This is the design Comms-care uses to give every consultant a dedicated
environment. The same pattern scales to a classroom or a proof-of-concept
factory.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>A nested ESXi vNIC on a <strong>standard</strong> VPC subnet is dead on arrival:
vmk0&rsquo;s made-up MAC loses to SpoofGuard, silently.</li>
<li>Attach nested-host vNICs <strong>only to a trunk subnet</strong>, with one binding map
per VLAN. The map lives under the VLAN subnet and points at the trunk.</li>
<li><strong>No DHCP in our VPC subnets</strong>, so bootstrap addressing goes through
<code>guestinfo.*</code> (appliances) or cloud-init (Linux). Static IP plans are a
feature in a lab anyway.</li>
<li>ESXi&rsquo;s default TCP/IP stack has <strong>one</strong> gateway. Set per-vmk override
gateways (<code>esxcli ... ipv4 set -g</code>) so vMotion and vSAN use their own
subnet&rsquo;s gateway.</li>
<li>Recreating a VM <strong>reallocates</strong> its NSX addresses. Pin what you depend on.</li>
<li>MTU: everything here ran at 1500. Raise the trunk and the nested vDS
before you do vSAN at any real scale.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-service-administration-and-development/9-0/managing-vsphere-kuberenetes-service-clusters-and-workloads/managing-networking-for-tkg-service-clusters/enable-antrea-egress-separate-subnet-on-a-tkg-cluster-with-nsx-vpc/create-a-subnetconnectionbindingmap-cr-on-the-supervisor.html">Create a SubnetConnectionBindingMap CR on the Supervisor</a>: the binding map CR and its <code>vlanTrafficTag</code>, shown for VKS egress subnets.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/segments/creating-a-child-segment.html">Creating a Child Segment</a>: the NSX mechanism underneath: a binding map on the child segment, pointing at its parent with a VLAN ID.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/segments/segment-profiles/understanding-spoofguard-segment-profile.html">Understanding SpoofGuard Segment Profile</a>: port address bindings, and traffic dropped when its MAC or IP does not match them.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/virtual-private-cloud-in-nsx/virtual-private-clouds-overview/add-a-vpc-service-profile.html">Add a VPC Service Profile</a>: the DHCP settings and segment profiles, SpoofGuard included, that a VPC&rsquo;s subnets inherit.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-consumption/latest/vm-service/deploy-vms-with-configurable-ovf-properties-vsphere-iaas-control-plane.html">Deploy VMs with Configurable OVF Properties in vSphere Supervisor</a>: OVF properties set through the VM Service&rsquo;s vAppConfig transport.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vsphere/vsphere/9-0/vsphere-networking/setting-up-vmkernel-networking/configure-the-vmkernel-adapter-gateway-by-using-esxcli.html">Configure the VMkernel Adapter Gateway by Using esxcli Commands</a>: a gateway per VMkernel adapter, set with esxcli.</li>
</ul>
<p>Next in this series: what happens when you want <em>ten</em> of these labs, with
byte-identical IP plans, firewalled from each other by construction. That&rsquo;s
where NSX VPCs go from &ldquo;workaround&rdquo; to genuinely better than physical.</p>
<hr>
<p><em>Lab environment; opinions my own. Everything above was captured from a live
VCF 9.1 environment — output trimmed for length, never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>The load balancer that must exist before the namespace</title>
      <link>https://thenestedlab.com/posts/the-lb-that-must-exist-first/</link>
      <pubDate>Wed, 16 Sep 2026 08:10:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/the-lb-that-must-exist-first/</guid>
      <description>Addresses pending for ever, a retryable error that never stops retrying, and an ordering rule the docs don&amp;rsquo;t tell you: in a self-service NSX VPC, the load balancer must exist before the namespace that uses it.</description>
      <content:encoded><![CDATA[<p>Everything was green. The VPC: realized. The namespace: ready. The VMs:
powered on, endpoints populated, ports listening. And the LoadBalancer
services sat at <code>&lt;pending&gt;</code>. For an hour.</p>
<p>This is the story of the least helpful error message in my recent memory,
what it actually means, and the one-line ordering rule that would have saved
an afternoon. If you&rsquo;re doing self-service NSX VPCs on VCF 9 with the
vSphere Supervisor, you will hit this. Bookmark accordingly.</p>
<h2 id="the-setup">The setup</h2>
<p>The VPC was tenant-created, through VCF Automation&rsquo;s Cloud Consumption
Interface (CCI) API. A supervisor namespace was pinned to it, with a couple
of <code>VirtualMachineService</code> objects of type <code>LoadBalancer</code> to publish SSH and
HTTPS for the workloads inside. Standard stuff: the exact pattern that works
out of the box in the org&rsquo;s default VPC.</p>
<p>The Kubernetes side looked perfect:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">$ kubectl get endpoints -n pod-a
</span></span><span class="line"><span class="cl">NAME           ENDPOINTS                        AGE
</span></span><span class="line"><span class="cl">esx01-access   172.30.0.40:443,172.30.0.40:22   6m36s
</span></span><span class="line"><span class="cl">esx02-access   172.30.0.41:443,172.30.0.41:22   6m35s
</span></span></code></pre></div><p>Endpoints resolved. VIPs: nothing. The only clue was a recurring event:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Warning  FailedRealizeNSXResource  service/esx01-access
</span></span><span class="line"><span class="cl">Generic error occurred during realizing network for Service
</span></span></code></pre></div><p>&ldquo;Generic error.&rdquo; Wonderful.</p>
<h2 id="digging-what-ncp-actually-wants">Digging: what NCP actually wants</h2>
<p>The supervisor&rsquo;s network container plugin (NCP) told the real story in its
logs:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">nsx_ujo.ncp.nsx.policy.lb_layer4_service Lb Service not Found for Namespace pod-a
</span></span><span class="line"><span class="cl">NCP00270 Failed to process virtual ip for service ...: Lbs pod-a is not found
</span></span><span class="line"><span class="cl">Encountered retryable error ... : Lbs pod-a is not found
</span></span></code></pre></div><p>NCP wants an NSX <strong>LBService</strong> in the namespace&rsquo;s VPC. In the org&rsquo;s
<em>default</em> VPC, one exists, because the platform created it when the VPC was
born. In my self-service VPC? Nobody had created one.</p>
<p>Fair enough. That&rsquo;s actually documented behaviour, once you know where to
look. A fresh VPC needs a <code>LoadBalancer</code> object. Before that, it needs a
<code>VPCAttachment</code> to a connectivity profile with the service gateway enabled.
Without the attachment, the load balancer creation itself fails, with a much
better error message.</p>
<p>So I created the attachment, then the LBService. NSX: <code>Realized=True</code>.
Problem solved?</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Warning  FailedRealizeNSXResource  service/esx01-access
</span></span><span class="line"><span class="cl">Generic error occurred during realizing network for Service
</span></span></code></pre></div><p>No.</p>
<h2 id="the-actual-bug-shaped-behaviour-a-snapshot-not-a-lookup">The actual bug-shaped behaviour: a snapshot, not a lookup</h2>
<p>Here&rsquo;s the part that costs you the afternoon. That &ldquo;retryable error&rdquo; retries
the <em>lookup in NCP&rsquo;s cache</em>, not the discovery. <strong>NCP snapshots the VPC&rsquo;s
load balancer inventory when the namespace is created.</strong> An LBService that
appears afterwards is never discovered, however long you wait:</p>
<ul>
<li>Recreating the Kubernetes services: no effect.</li>
<li>Tagging the LBService with the <code>nsx-op/*</code> ownership tags the working ones
carry: no effect, because the cache doesn&rsquo;t re-read NSX.</li>
<li>Restarting NCP would force a full resync, but supervisor system pods are
protected. Even <code>Administrator@vsphere.local</code> gets a Forbidden, which is a
humbling thing to read.</li>
<li>Mutating the namespace to nudge a resync: also blocked, by the
supervisor&rsquo;s namespace validation webhook.</li>
</ul>
<p>As a tenant, there is exactly one fix: <strong>delete and recreate the namespace</strong>,
now that its VPC has a load balancer. Fifteen minutes of rebuild, for want of
one ordering rule.</p>
<p>And the control experiment proves the rule. A namespace created <em>after</em> its
VPC already had an LBService got its VIPs assigned without any drama:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">esx01-access   VIP=192.168.144.34   22 OPEN · 443 OPEN
</span></span><span class="line"><span class="cl">esx02-access   VIP=192.168.144.35   22 OPEN · 443 OPEN
</span></span></code></pre></div><p>Once the ordering is right, this is what &ldquo;working&rdquo; looks like. It&rsquo;s the
pod&rsquo;s state a couple of minutes after a correctly ordered deployment:</p>
<p><img alt="Live replay: catalog-deployed pod with both VMs powered on and VIPs assigned" loading="lazy" src="/images/c2-catalog-pod.gif"></p>
<p><img alt="VCFA deployment topology: namespace, subnets, hosts, two VIPs" loading="lazy" src="/images/ui/u4-deployment-topology.jpg">
<em>What the requester sees once the order is right.</em></p>
<h2 id="the-ordering-rule">The ordering rule</h2>
<p>For every self-service VPC that will publish LoadBalancer services, create
these in order, <em>before</em> the namespace:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">1. VPC                                   (vpc.nsx.vmware.com/v1alpha1)
</span></span><span class="line"><span class="cl">2. VPCAttachment                         (connectivity profile w/ service gateway
</span></span><span class="line"><span class="cl">                                          — LB creation errors without it)
</span></span><span class="line"><span class="cl">3. LoadBalancer   {regionName, vpcName}  (the step everyone misses)
</span></span><span class="line"><span class="cl">4. ...and only THEN the Supervisor Namespace
</span></span></code></pre></div><p>Encode it in whatever provisions your VPCs: a script, a pipeline, an
operator. It&rsquo;s four API calls, which is a good deal cheaper than an
afternoon. And it turns a silent, undiagnosable <code>&lt;pending&gt;</code> into a platform
that just works.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Nobody buys a platform for its ordering rules. But this is exactly the kind
of edge that decides whether self-service provisioning feels reliable or
flaky to the people using it.</p>
<p>In a customer deployment, the answer isn&rsquo;t a blog post. It&rsquo;s provisioning
automation that already does the four steps in the right order, every time,
so a tenant never sees a VIP stuck at <code>&lt;pending&gt;</code>. Knowing where the sharp
edges are, because you&rsquo;ve been cut by them in a lab, is most of what an
experienced delivery partner is for.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>In a self-service NSX VPC, <strong>the LBService must predate the namespace</strong>.
NCP discovers load balancers when the namespace is added, and never again.</li>
<li><code>FailedRealizeNSXResource: Generic error</code> on a Service means: go and read
the NCP logs. The real message (<code>Lbs &lt;ns&gt; is not found</code>, NCP00270) is
there.</li>
<li><code>VPCAttachment</code> (service gateway) is the prerequisite for the load
balancer itself. That one, at least, fails loudly.</li>
<li>Retro-tagging NSX objects to look &ldquo;owned&rdquo; doesn&rsquo;t help a cache that never
re-reads. Recreating the namespace is the only tenant-level fix.</li>
<li>While you&rsquo;re at it: new namespaces also reject VM creation until image
<code>status.disks</code> syncs (~1–3 minutes after content library attach). Build
the wait into your automation and both sharp edges disappear.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/vsphere-supervisor-installation-and-configuration/supervisor-networking-with-virtual-private-clouds.html">Deploying Supervisor with VCF Networking with VPC</a>: what NCP creates for a namespace given no VPC: a VPC with its load balancer and SNAT IP.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/vsphere-supervisor-installation-and-configuration/configuring-and-managing-vsphere-namespaces/managing-vsphere-namespaces-on-a-supervisor-with-nsx-vpc/create-and-configure-a-vsphere-namespace-on-a-supervisor-with-vpc/create-namespaces-with-vpc-nosnat-nolb.html">Create vSphere Namespaces on VPCs without SNAT and Load Balancer</a>: without the VPC&rsquo;s load balancer, LoadBalancer services and VirtualMachineServices cannot be deployed.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/virtual-private-cloud-in-nsx/virtual-private-clouds-overview/add-a-vpc-connectivity-profile.html">Add a VPC Connectivity Profile</a>: the transit gateway, the service gateway and default outbound NAT.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/vsphere-supervisor-installation-and-configuration/configuring-and-managing-vsphere-namespaces/managing-vsphere-namespaces-on-a-supervisor-with-nsx-vpc/create-and-configure-a-vsphere-namespace-on-a-supervisor-with-vpc.html">Create and Configure a vSphere Namespace on a Supervisor with NSX VPC</a>: placing a new namespace in an existing VPC.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/adding-and-managing-virtual-private-clouds/add-a-vpc.html">Create a Virtual Private Cloud in VCF Automation</a>: a tenant VPC, its connectivity profile and its load balancing setting.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/managing-blueprints-in-vcf-automation/sample-blueprints-in-vcf-automation-for-all-apps.html">Sample Blueprints in VCF Automation</a>: a VPC, its VPCAttachment and a namespace that depends on them, as blueprint resources.</li>
</ul>
<p><em>Previously in this series: <a href="/posts/nested-esxi-nsx-vpc/">nested ESXi inside an NSX VPC</a>.
Next: three datacenters, one IP plan, with identical isolated pods.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Output captured live, trimmed for length,
never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Dual-NIC nested hosts: what redundancy means when the fabric is virtual</title>
      <link>https://thenestedlab.com/posts/dual-nic-nested-hosts/</link>
      <pubDate>Wed, 16 Sep 2026 07:30:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/dual-nic-nested-hosts/</guid>
      <description>A second vNIC on a nested host adds no physical redundancy, so why add it? We expected bringup to want one, and pulling vmnic0 mid-SSH proved the trunk copes: 0% loss, session intact.</description>
      <content:encoded><![CDATA[<p>&ldquo;Naturally, a VCF host has at least two NICs. Are we testing that, or have
you virtualised it away?&rdquo;</p>
<p>Fair question, and like most fair questions it has an irritating two-part answer. In a nested lab, the
outer host provides the <em>physical</em> redundancy: its vDS and its NSX uplinks.
A second vNIC on the nested VM adds exactly none of that.</p>
<p>But VCF doesn&rsquo;t know it&rsquo;s nested. We expected bringup&rsquo;s host validation, and
the vDS uplink teaming it sets up, to <strong>want two vmnics</strong>. We never tried a
host with one. For the record, Broadcom&rsquo;s docs
<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/building-your-private-cloud-infrastructure/host-management/commission-hosts.html">allow single-pNIC hosts</a>,
and the VCF Installer&rsquo;s
<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/release-notes/vmware-cloud-foundation-9-1-0-0-release-notes/known-issues/vcf-installer-91-known-issues.html">9.1 known issue</a>
only bites single-pNIC hosts that use an NFS datastore.</p>
<p>So our nested hosts get two vNICs, both on the trunk subnet. Which leaves the
fun question (for a given value of fun): does failover between them actually
work inside a VPC?</p>
<p><img alt="The failover test end to end: an SSH session through the NSX load balancer into the trunk subnet and esx01&rsquo;s two vNICs; vmnic0 fails and every VLAN moves to vmnic1" loading="lazy" src="/images/diagrams/dual-nic-failover.svg">
<em>The whole test on one page. The session I&rsquo;m typing in rides the very path I break.</em></p>
<h2 id="the-setup">The setup</h2>
<p>Both vNICs sit on the same <code>sn-trunk</code> subnet
(<a href="/posts/nested-esxi-nsx-vpc/">the trunk from part 1</a>), and ESXi sees two
10G vmnics:</p>
<p><img alt="Host Client, Physical Adapters: vmnic0 and vmnic1, both 10 Gbit/s, both on vSwitch0" loading="lazy" src="/images/ui/u13-hostclient-dual-nics-marked.jpg"></p>
<p>vSwitch0 teams them active/active with the default originating-port-ID
policy. Every port group inherits it: Management 1610, vMotion 1611 and
vSAN 1612. Nothing you wouldn&rsquo;t do on metal. Thrilling stuff, I know.</p>
<h2 id="the-test-pull-a-nic-while-watching-from-inside">The test: pull a NIC while watching from inside</h2>
<p>The interesting part isn&rsquo;t whether pings carry on. It&rsquo;s <em>which</em> session I&rsquo;m
watching from.</p>
<p>I&rsquo;m SSH&rsquo;d into <code>esx01</code> <strong>through its public VIP</strong>. So my session runs through
the NSX load balancer, the VPC, the trunk port and whichever vmnic happens to
carry vmk0. If failover breaks anything, it breaks the terminal I&rsquo;m typing
in: the networking equivalent of sawing off the branch you&rsquo;re sitting on.
Low effort, high stakes: the best kind of test.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">[root@esx01-a:~] esxcli network nic list
</span></span><span class="line"><span class="cl">Name    ...  Admin Status  Link Status  Speed  MAC Address
</span></span><span class="line"><span class="cl">vmnic0  ...  Up            Up           10000  04:50:56:00:5c:03
</span></span><span class="line"><span class="cl">vmnic1  ...  Up            Up           10000  04:50:56:00:68:00
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] esxcli network nic down -n vmnic0    # FAIL THE FIRST NIC
</span></span><span class="line"><span class="cl">vmnic0  ...  Down          Down             0  04:50:56:00:5c:03
</span></span><span class="line"><span class="cl">vmnic1  ...  Up            Up           10000  04:50:56:00:68:00
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] vmkping -I vmk1 172.30.0.71          # vMotion VLAN, now over vmnic1
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] vmkping -I vmk2 172.30.0.101         # vSAN VLAN, now over vmnic1
</span></span><span class="line"><span class="cl">3 packets transmitted, 3 packets received, 0% packet loss
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">[root@esx01-a:~] esxcli network nic up -n vmnic0      # restore
</span></span><span class="line"><span class="cl"># session never dropped.
</span></span></code></pre></div><figure class="nl-video">
  <video autoplay loop muted playsinline controls preload="metadata" style="aspect-ratio:1568 / 604" poster="/images/c6-nic-failover-beforeafter-poster.jpg">
    <source src="/images/c6-nic-failover-beforeafter.mp4" type="video/mp4">
  </video>
  <figcaption>Before and after: vmnic0 down, every VLAN still passing.</figcaption>
</figure>

<details class="nl-fold">
<summary>The full transcript, as captured</summary>
<p><img alt="Full failover transcript" loading="lazy" src="/images/demo-c6-nic-failover.jpg"></p>

</details>

<p>Every VLAN moved to vmnic1. Zero loss on vMotion and vSAN. And the
management session, the one <em>most</em> likely to notice, never blinked. Slightly
disappointing, if I&rsquo;m honest; I&rsquo;d prepared a dramatic paragraph.</p>
<h2 id="why-this-is-a-real-result-not-a-party-trick">Why this is a real result, not a party trick</h2>
<p>Think about what just happened at the NSX layer. vmk0 has a MAC of its own,
one ESXi made up rather than the vNIC&rsquo;s. NSX had learned it on trunk port A.
When vmnic0 went down, the same MAC turned up on trunk port B, mid-flow, with
a live TCP session riding on it.</p>
<p>On a <strong>standard</strong> VPC subnet port, that is exactly what SpoofGuard exists to
stop. The port&rsquo;s address bindings pin one MAC. A frame from a different MAC,
or the <em>same</em> MAC arriving on a different port, gets dropped. Part 1 showed
that rule silencing a host on a standard subnet before it ever spoke.</p>
<p>This test shows the trunk subnet calmly accepting a foreign MAC that moves
between two of its ports. Nested vSphere needs exactly that property, and so
does anything else with a vSwitch inside a VM.</p>
<p>So the second vNIC buys three things, none of them physical redundancy:</p>
<ol>
<li><strong>Bringup and vDS teaming get the two uplinks</strong> we expected them to want.</li>
<li><strong>The teaming policy you&rsquo;ll configure in production gets exercised:</strong>
uplink failover, active/standby for vSAN, whatever you&rsquo;re rehearsing.</li>
<li><strong>A live proof that the trunk carries MAC mobility.</strong> That&rsquo;s the real
reassurance: the design isn&rsquo;t relying on a quiet network.</li>
</ol>
<p>One caution if the host is heading into a VCF Installer bringup. The 9.x
installer&rsquo;s validation wants exactly one physical NIC on vSwitch0, and stops
with &ldquo;has 2 Physical NICs connected to vSphere Standard Switch vSwitch0
(Expecting 1)&rdquo; (<a href="https://knowledge.broadcom.com/external/article/415469">KB 415469</a>).
Give such a host its second vNIC, but leave vmnic1 unclaimed until bringup
takes it. The test above teams both on vSwitch0 because it tests the trunk,
not a bringup.</p>
<h2 id="what-it-does-not-buy-and-how-to-say-so">What it does <em>not</em> buy, and how to say so</h2>
<p>If someone asks &ldquo;is this host redundant?&rdquo;, the honest answer is &ldquo;the nested
host believes it is, bless it; the actual redundancy lives one layer down.&rdquo;</p>
<p>In a training pod, that&rsquo;s the right answer. Students configure and test
failover exactly as they would on metal, while the outer platform does the
real work. It&rsquo;s usually enough to reproduce a NIC-teaming issue from the
field too: most teaming bugs live in ESXi&rsquo;s policy handling, not in the copper.</p>
<p>Where it genuinely falls short: anything about physical link behaviour. LACP
negotiation, LLDP, a flapping link, an MTU mismatch on one uplink. The virtual
fabric never fails lopsidedly, so it can&rsquo;t reproduce any of those.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>The real value is knowing <em>what a nested environment can and can&rsquo;t prove</em>,
so you can tell when a virtual lab is enough. For training, upgrade
rehearsals, configuration and policy testing, and most &ldquo;how does it behave
when…&rdquo; questions, nested is enough and far cheaper. For physical link
behaviour (LACP, optics, lopsided faults) you still want metal. Making that
call with confidence is worth more than the test itself.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Give nested VCF hosts <strong>two vNICs on the same trunk subnet</strong>. We expected
bringup and vDS teaming to want at least two vmnics, and humouring them
costs nothing. Before a VCF Installer bringup, leave vmnic1 off vSwitch0
(KB 415469).</li>
<li>Test failover <strong>from a session that depends on it</strong> (SSH through the VIP).
Pings passing while your terminal dies is not success, however much the
change ticket would like it to be.</li>
<li>The trunk subnet tolerates <strong>a vmk MAC moving between ports mid-flow</strong>.
Standard subnets lack that property, and nested vSphere needs it.</li>
<li>Be precise in the write-up: nested dual-NIC gives <em>policy</em> realism, not
<em>physical</em> redundancy. Physical link faults can&rsquo;t be reproduced here.</li>
<li>Rebuilds re-run the vmk config: the appliance creates vmk0 only, so vmk1,
vmk2, their VLANs and override gateways are applied after boot.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vsphere/vsphere/9-0/vsphere-networking/networking-policies/teaming-and-failover-policy/configure-nic-teaming-and-load-balancing-on-a-standard-switch-or-port-group.html">Configure NIC Teaming, Failover, and Load Balancing on a vSphere Standard Switch or Standard Port Group</a>: the default originating-port policy, failover order, and port groups inheriting the switch&rsquo;s policy.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/segments/segment-profiles/understanding-spoofguard-segment-profile.html">Understanding SpoofGuard Segment Profile</a>: the address bindings a standard port enforces.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/segments/segment-profiles/understanding-mac-discovery-segment-profile.html">Understanding MAC Discovery Segment Profile</a>: MAC learning for nested hypervisors, with many MACs behind one vNIC.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/release-notes/vmware-cloud-foundation-9-1-0-0-release-notes/known-issues/vcf-installer-91-known-issues.html">VCF Installer</a>: the VCF 9.1 known issue where single-pNIC hosts fail NFS datastore validation, and the fix of two or more pNICs.</li>
<li><a href="https://knowledge.broadcom.com/external/article/313547/support-for-running-esxi-as-a-nested-vir.html">Support for running ESXi as a nested virtualization solution</a>: nested ESXi is not supported in production, and is encouraged for learning, training and testing.</li>
<li><a href="https://knowledge.broadcom.com/external/article/415469">VCF 9.0 Installer validation fails at ESX Host Configuration (KB 415469)</a>: one physical NIC on vSwitch0 before deployment, the other left unclaimed.</li>
</ul>
<p><em>Companion to <a href="/posts/nested-esxi-nsx-vpc/">nested ESXi inside an NSX VPC</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Output captured live, trimmed for length,
never edited for outcome.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Our VPC subnets have no DHCP — and that&#39;s fine</title>
      <link>https://thenestedlab.com/posts/vpc-subnets-have-no-dhcp/</link>
      <pubDate>Wed, 16 Sep 2026 07:20:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/vpc-subnets-have-no-dhcp/</guid>
      <description>The nested ESXi appliance sat at &amp;lsquo;waiting for DHCP&amp;rsquo; for ever. VPC subnets don&amp;rsquo;t hand out addresses: the VM Service does, through cloud-init, sysprep or OVF guestinfo, depending on the guest.</description>
      <content:encoded><![CDATA[<p>The second trap from <a href="/posts/nested-esxi-nsx-vpc/">part 1</a> deserves its own
short post, because it catches everything, not just ESXi: <strong>our VPC subnets
have DHCP deactivated</strong>. NSX VPC subnets can run a DHCP server or relay
(<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/virtual-private-cloud-in-nsx/virtual-private-clouds-overview/add-a-subnet-for-the-vpc.html">Add a Subnet to a VPC</a>).
But the subnets the Supervisor made for us, and every <code>Subnet</code> we created
with the defaults, came up <code>DHCP_DEACTIVATED</code> with static IP allocation.</p>
<p>Drop a stock appliance onto one and it will boot, sit at &ldquo;waiting for DHCP&rdquo;,
and wait politely until the heat death of the universe.</p>
<p>This isn&rsquo;t a gap. It&rsquo;s the model. NSX allocates the address at the <em>port</em>
and pins it there with address bindings, and the <em>guest</em> has to be told what
it was given. The VM Service does that telling through <strong>bootstrap
providers</strong>. Once you know the three of them, static addressing stops being
a chore and starts being a feature.</p>
<h2 id="three-providers-three-guest-types">Three providers, three guest types</h2>
<table>
	<thead>
			<tr>
					<th>Guest</th>
					<th>Provider</th>
					<th>Carries</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Linux</td>
					<td><code>cloudInit</code></td>
					<td>user-data (users, <code>write_files</code>, <code>runcmd</code>) + network config</td>
			</tr>
			<tr>
					<td>Windows</td>
					<td><code>sysprep</code></td>
					<td>unattend XML / sysprep spec, identity, network</td>
			</tr>
			<tr>
					<td>Appliances (OVF)</td>
					<td><code>vAppConfig</code></td>
					<td>OVF properties (<code>guestinfo.*</code>) the appliance reads on boot</td>
			</tr>
	</tbody>
</table>
<p>All three are <strong>typed fields on the <code>VirtualMachine</code> object</strong>, not bolt-on
customisation specs. The platform already knows the network side: the VM
Service knows which subnet each interface landed on, and what NSX allocated.
So for cloud-init and sysprep, the addressing is injected for you.
Appliances are the exception, because each one has its own idea of which
properties it wants.</p>
<h2 id="appliances-vappconfig">Appliances: vAppConfig</h2>
<p>The nested-ESXi appliance reads <code>guestinfo.*</code> OVF properties. In the VM
Service spec, that&rsquo;s:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">vAppConfig</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">properties</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.hostname,  value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;esx01.pod-a.res.lab&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.ipaddress, value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.40&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.netmask,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;255.255.255.224&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.gateway,   value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;172.30.0.33&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.vlan,      value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;1610&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.dns,       value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;10.20.52.1&#34;</span>}}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span>- {<span class="nt">key</span><span class="p">:</span><span class="w"> </span><span class="nt">guestinfo.ssh,       value</span><span class="p">:</span><span class="w"> </span>{<span class="nt">value</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;True&#34;</span>}}<span class="w">
</span></span></span></code></pre></div><p>The one that trips people up: <strong>the address you give must be the one NSX
allocated to the port</strong>. In a standard subnet, SpoofGuard enforces that, and
it doesn&rsquo;t negotiate. On a trunk subnet, the VLAN subnets have their own
allocations, and you&rsquo;re choosing addresses within them.</p>
<p>Either way, pick from the realized range. And remember that <a href="/posts/nested-esxi-nsx-vpc/">recreating a VM
reallocates its addresses</a>. So the fixed
<code>.40</code>/<code>.41</code> in a <a href="/posts/three-datacenters-one-ip-plan/">deterministic pod</a>
is a design choice, not luck.</p>
<h2 id="linux-cloud-init">Linux: cloud-init</h2>
<p>A Secret holding user-data, referenced from the VM:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">cloudInit</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">cloudConfig</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">users</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w"> </span><span class="l">... a local user with a key ... ]</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">write_files</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span>- <span class="nt">path</span><span class="p">:</span><span class="w"> </span><span class="l">/var/www/html/index.html</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">          </span><span class="nt">content</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;shared-svc repo01\n&#34;</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">runcmd</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">        </span>- <span class="p">[</span><span class="l">systemctl, enable, --now, nginx]</span><span class="w">
</span></span></span></code></pre></div><p>Networking arrives through the platform&rsquo;s own network-config, so you don&rsquo;t
write it. Every line of YAML I don&rsquo;t write is a line I can&rsquo;t get wrong.</p>
<p>The <code>svc-repo01</code> VM from <a href="/posts/shared-services-for-isolated-tenants/">the shared-services
post</a> was exactly this:
<code>write_files</code> plus <code>runcmd</code>, with its web page verified from three pods.</p>
<h2 id="windows-sysprep">Windows: sysprep</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">bootstrap</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">sysprep</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">sysprep</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">guiUnattended</span><span class="p">:</span><span class="w"> </span>{<span class="nt">autoLogon</span><span class="p">:</span><span class="w"> </span><span class="nt">true, autoLogonCount</span><span class="p">:</span><span class="w"> </span><span class="nt">1, timeZone</span><span class="p">:</span><span class="w"> </span><span class="m">85</span>}<span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">identification</span><span class="p">:</span><span class="w"> </span>{<span class="nt">joinWorkgroup</span><span class="p">:</span><span class="w"> </span><span class="l">WORKGROUP}</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">userData</span><span class="p">:</span><span class="w"> </span>{<span class="nt">fullName</span><span class="p">:</span><span class="w"> </span><span class="nt">Lab, orgName</span><span class="p">:</span><span class="w"> </span><span class="nt">Lab, computerName</span><span class="p">:</span><span class="w"> </span>{<span class="nt">name</span><span class="p">:</span><span class="w"> </span><span class="l">win01}}</span><span class="w">
</span></span></span></code></pre></div><p>Or use <code>rawSysprep</code> with an unattend XML in a Secret, if you already have
one. The ISO from <a href="/posts/nested-esxi-via-vcfa-all-apps/">the blueprint post</a>
can ride along as a declarative <code>hardware.cdrom</code>. That&rsquo;s handy for tools and
agents on first boot.</p>
<h2 id="the-esxi-footnote-one-gateway-many-vmks">The ESXi footnote: one gateway, many vmks</h2>
<p>Once the appliance is up and you add vMotion and vSAN vmks on their own
subnets, you hit a detail that makes the Host Client <em>look</em> wrong:</p>
<p><img alt="Host Client: vmk0/1/2, one service each" loading="lazy" src="/images/ui/u12-hostclient-vmk-adapters.jpg"></p>
<p>ESXi&rsquo;s default TCP/IP stack has <strong>one</strong> default gateway (vmk0&rsquo;s <code>.33</code>), and
the UI repeats it on every vmk row with total confidence. Same-subnet vMotion
never uses a gateway, so nothing breaks. But each NSX subnet <em>does</em> have its
own gateway, and cross-subnet traffic from vmk1 and vmk2 would take the wrong
exit. Set per-vmk override gateways:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">esxcli network ip interface ipv4 set -i vmk1 -t static -I 172.30.0.70  -N 255.255.255.224 -g 172.30.0.65
</span></span><span class="line"><span class="cl">esxcli network ip interface ipv4 set -i vmk2 -t static -I 172.30.0.100 -N 255.255.255.224 -g 172.30.0.97
</span></span></code></pre></div><p>Now the display is truthful and the routing is correct. (The full-realism
alternative is a dedicated <code>vmotion</code> netstack for vmk1. I kept the default
stack so the service tags stay visible in the Host Client.)</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>For the business, &ldquo;no DHCP&rdquo; translates into something security and
operations teams both want: <strong>predictable addressing</strong>. Every environment has
a known address plan, firewall rules can be written once, and nothing turns
up on the network with an address nobody expected.</p>
<p>Bootstrap providers deliver the second benefit. Images stay generic and
configuration is injected at deploy time, so there are fewer golden images to
maintain and far less drift between environments.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li><strong>No DHCP in our VPC subnets</strong>: that&rsquo;s the Supervisor&rsquo;s default, not a VPC
limit. NSX allocates at the port, and the guest is told through a
bootstrap provider.</li>
<li><code>cloudInit</code> (Linux), <code>sysprep</code> (Windows) and <code>vAppConfig</code> (appliances) are
typed fields on the VM, not customisation specs.</li>
<li>Appliance addresses must match the <strong>realized</strong> subnet. Fix the order of
subnet creation if you want fixed addresses across pods.</li>
<li>ESXi has <strong>one</strong> default gateway per stack. Set <code>-g</code> per vmk, or the Host
Client lies to you and cross-subnet traffic exits wrong.</li>
<li>A static IP plan is a feature in a lab: it&rsquo;s what makes screenshots,
runbooks and pods identical.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/virtual-private-cloud-in-nsx/virtual-private-clouds-overview/add-a-subnet-for-the-vpc.html">Add a Subnet to a VPC</a>: a subnet&rsquo;s DHCP setting: none for static addresses, a DHCP server, or DHCP relay.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/advanced-network-management/segments/segment-profiles/understanding-spoofguard-segment-profile.html">Understanding SpoofGuard Segment Profile</a>: the port address bindings SpoofGuard enforces.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-consumption/latest/vm-service/provision-a-vm-using-the-iaas-services-console-in-vcf-automation.html">Provision a VM Using Self-Service</a>: the four bootstrap methods, cloud-init, Sysprep, Linuxprep and vAppConfig, and static IP allocation.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-consumption/latest/vm-service/deploy-vms-with-configurable-ovf-properties-vsphere-iaas-control-plane.html">Deploy VMs with Configurable OVF Properties in vSphere Supervisor</a>: OVF properties set through the VM Service&rsquo;s vAppConfig transport.</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vsphere/vsphere/9-0/vsphere-networking/setting-up-vmkernel-networking/configure-the-vmkernel-adapter-gateway-by-using-esxcli.html">Configure the VMkernel Adapter Gateway by Using esxcli Commands</a>: a gateway per VMkernel adapter, set with esxcli.</li>
</ul>
<p><em>Companion to <a href="/posts/nested-esxi-nsx-vpc/">nested ESXi inside an NSX VPC</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
