<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>VCF Installer on The Nested Lab</title>
    <link>https://thenestedlab.com/products/vcf-installer/</link>
    <description>Recent content in VCF Installer on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/products/vcf-installer/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>One catalog item, one VCF instance: building a lab factory</title>
      <link>https://thenestedlab.com/posts/one-catalog-item-one-vcf-instance/</link>
      <pubDate>Wed, 16 Sep 2026 06:40:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/one-catalog-item-one-vcf-instance/</guid>
      <description>How an interactive PowerShell script grew into a catalog item that builds complete nested VCF 9.1 instances from one request form. The design rules that made it survivable, and the traps behind them.</description>
      <content:encoded><![CDATA[<p>Every nested VCF lab starts the same way: a heroic PowerShell script.
Ours was <code>esxihostdeploy.ps1</code>: ovftool plus PowerCLI, with an interactive
menu asking which environment, which ESX version, which role and how many
hosts.</p>
<p>It worked. It also lived on one person&rsquo;s machine, prompted for credentials,
and knew nothing about anything that comes <em>after</em> the hosts exist. Heroic,
then, but not much of a team player.</p>
<p>This post is about what it became: a set of VCF Automation catalog items
where requesting <strong>one form</strong> produces a complete nested VCF 9.1 instance.
That means ESXi hosts, bringup (vCenter, NSX, SDDC Manager), a vSphere
Supervisor, VCF Automation, Operations and identity. The environment number
is practically the only real input.</p>
<h2 id="the-shape-of-the-factory">The shape of the factory</h2>
<p>Three stages, each a catalog item, plus a wrapper that chains them:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Stage 1  Nested ESX Hosts        VM Apps template + vRO actions
</span></span><span class="line"><span class="cl">         (the old script, reborn declaratively)
</span></span><span class="line"><span class="cl">Stage 2  Deploy VCF 9.1 Instance vRO workflow driving the VCF Installer API
</span></span><span class="line"><span class="cl">         (spec generated, validated, bringup started)
</span></span><span class="line"><span class="cl">Day-N    Supervisor · NSX Edge · VCF Automation · Ops Logs/Networks/RTM ·
</span></span><span class="line"><span class="cl">         Identity (AD)           one catalog item each
</span></span><span class="line"><span class="cl">Wrapper  &#34;Deploy VCF Stack&#34;      one form, checkbox per component
</span></span></code></pre></div><p><img alt="The factory catalog: hosts, bringup, every day-N component, and the wrapper — ten tiles" loading="lazy" src="/images/ui/f1-f00-factory-catalog.jpg"></p>
<p>The wrapper&rsquo;s form has a checkbox per component. Ticking one reveals that
component&rsquo;s tab, with every field already filled in. One lab password feeds
every credential. Tick everything, click request, and go and make a coffee.
Possibly several.</p>
<h2 id="rule-1-derive-everything-from-one-number">Rule 1: derive everything from one number</h2>
<p>Each lab environment is <code>f0X</code>, and <em>everything</em> scales from X by formula:</p>
<table>
	<thead>
			<tr>
					<th>Element</th>
					<th>Pattern</th>
					<th>f03 example</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Names</td>
					<td><code>f0X-m01-*</code></td>
					<td><code>f03-m01-vc01.res.lab</code></td>
			</tr>
			<tr>
					<td>Subnets</td>
					<td><code>10.(20+X).&lt;sub&gt;.0/24</code></td>
					<td><code>10.23.1.0/24</code> (mgmt)</td>
			</tr>
			<tr>
					<td>VLANs</td>
					<td><code>2X0n</code></td>
					<td>2307 (edge TEP)</td>
			</tr>
	</tbody>
</table>
<p>The bringup spec is hundreds of lines of JSON, which the VCF Installer wants
and nobody wants to type. A vRO action generates it from a known-good
reference spec plus X. Nobody edits a deployment spec by hand, which means
nobody typo-breaks a bringup at 2am.</p>
<p>When the old script did this, the formulas lived in string concatenation.
Now they live in one action, with the reference spec beside it.</p>
<p>The same philosophy carried into stage 1:</p>
<ul>
<li>the script&rsquo;s &ldquo;next free esxNN index&rdquo; scan became a vRO action bound to the
request form;</li>
<li>its VLAN/IP arithmetic became template expressions;</li>
<li>its <code>--prop:guestinfo.*</code> flags became <code>ovfProperties</code> in the template.</li>
</ul>
<p>Porting a script isn&rsquo;t rewriting it. It&rsquo;s finding the declarative home for
each behaviour.</p>
<h2 id="rule-2-never-wait-for-anything-you-can-watch-instead">Rule 2: never wait for anything you can watch instead</h2>
<p>One constraint shaped the whole design. A VCF Automation request gets about
<strong>two hours</strong> before the platform gives up on it. Our runs died at exactly
two hours with <code>Delegating token is not service token</code>, which is a roundabout
way of saying &ldquo;time&rsquo;s up&rdquo;.</p>
<p>The project&rsquo;s
<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/vcfa-overview/getting-started-with-organizations-for-vm-apps-in-vcf-automation/map-head-projects-adding-and-managing-projects/projects-how-do-i-add-a-project-for-my-development-team.html">request timeout</a>
also defaults to two hours, and its Provisioning tab can raise it. We never
tried. A full VCF bringup takes longer than that, and it shouldn&rsquo;t hold a
request open for hours anyway. So the wrapper <em>never waits</em>:</p>
<ul>
<li>Bringup is <strong>fire-and-forget</strong>. The workflow authenticates to the
installer, validates the spec, starts the task, and hands back a
<code>watchTaskId</code>. Re-attach any time to check on it.</li>
<li>Fleet deployments (VCF Automation, Ops for Logs/Networks, metrics) are
server-side tasks. The items submit with <code>waitForCompletion=false</code>.</li>
<li>The supervisor item submits enablement and returns; vCenter carries on.</li>
<li>Only fast, deterministic steps (identity configuration, minutes) run to
completion inside the request.</li>
</ul>
<p><img alt="F06-Mgmt-VCF: the wrapper deployment, Create Successful, 13:12 → 14:05" loading="lazy" src="/images/ui/f3-f00-stack-deployment-success.jpg">
<em>A whole VCF instance as one deployment record. The request finished in under an hour, while the build ran on for twelve.</em></p>
<p>The result: a full-stack kick-off <em>completes</em> as a request in well under an
hour (53 minutes on the run pictured). The actual multi-hour build carries
on as server-side tasks you can watch. The request&rsquo;s job isn&rsquo;t to do the
work. It&rsquo;s to <strong>start the work correctly</strong> and tell you where to watch it.</p>
<p>The corollary: a failed component is recorded, and the remaining components
still run. You fix one thing and re-run one item, not the world.</p>
<h2 id="rule-3-plan-mode-for-infrastructure">Rule 3: plan mode for infrastructure</h2>
<p>Every item in the chain supports <code>validateOnly</code>, and the wrapper cascades
it. Tick everything, set validateOnly, and the entire stack is <em>planned</em>
against the live environment with zero changes. Specs are generated,
prerequisites checked, and name and IP collisions caught. A smoke runner
does exactly this on every change to the automation itself.</p>
<p>If you build nothing else into your lab automation, build this. The number
of 2am bringups saved by a five-minute dry run is not small.</p>
<h2 id="the-traps-that-shaped-the-rules">The traps that shaped the rules</h2>
<p>Some of the design above exists because of scars:</p>
<ul>
<li><strong>Hardware validation hates virtual NVMe.</strong> Bringup&rsquo;s hardware
compatibility list (HCL) check will block nested hosts. The spec generator
has to account for it, or you discover it two hours in. Twice, if you&rsquo;re us.</li>
<li><strong>Small disks, surprising layouts.</strong> Nested hosts with 64 GB disks ship
ESX-OSDATA at essentially the whole disk. A post-provision step relocates
scratch, or stage 2 fills the disk with logs.</li>
<li><strong>DNS is a prerequisite, not a step.</strong> The installer&rsquo;s pre-flight wants
every record resolvable before it starts. A one-shot script creates the
per-environment records ahead of the request.</li>
<li><strong>vRO&rsquo;s content-source lag.</strong> A new or changed workflow takes 15–20
minutes of data collection before the catalog sees it. Publish, wait,
<em>then</em> test, or you&rsquo;ll debug a ghost.</li>
<li><strong>Wrapper inputs are duplicated by necessity.</strong> vRO requires every
sub-workflow input to be passed explicitly, so adding an input to a
component means updating the wrapper&rsquo;s call too. Null-guards in each
component turn a forgotten field into a loud failure instead of a silent
default.</li>
</ul>
<h2 id="why-bother">Why bother?</h2>
<p>Because the payoff compounds. Once a full VCF instance is a catalog request,
everything downstream changes character:</p>
<ul>
<li>upgrade rehearsals happen on freshly built instances instead of precious
pets;</li>
<li>a broken environment is redeployed, not repaired;</li>
<li>the lab stops being a collection of snowflakes and becomes a <em>product</em>:
versioned, validated, reproducible.</li>
</ul>
<p><img alt="Three pods, identical IP plans, no route between them" loading="lazy" src="/images/product-01-hook.jpg">
<em>Where this is heading: the factory&rsquo;s output feeding per-student pods with identical addressing.</em></p>
<p>The factory&rsquo;s next customers, funnily enough, are the isolated VPC pods from
<a href="/series/the-vpc-pod-papers/">the other series on this blog</a>. Same
philosophy, one layer further down.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>A complete VCF instance from one form changes what an environment costs
to have. Environments used to be precious, because building one took a
week. Now they&rsquo;re disposable, and a lot follows from that:</p>
<ul>
<li><strong>Proofs of concept</strong> run on an environment built for the customer&rsquo;s
scenario, not on whatever happens to be free.</li>
<li><strong>Upgrade and migration rehearsals</strong> happen on a fresh instance of the
right version, then it&rsquo;s deleted.</li>
<li><strong>Training and enablement</strong> get a real VCF per person or per team.</li>
<li><strong>Reference builds</strong> exist for every supported release, on demand.</li>
</ul>
<p>This is how Comms-care provides a dedicated instance to every consultant.
The same factory, pointed at a customer&rsquo;s requirements, is a repeatable
way to deliver environments, rather than a one-off project each time.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Derive names, subnets and VLANs from a single environment number; generate
specs, never hand-edit them.</li>
<li>Respect the request-duration ceiling: start long work, return a task
handle, re-attach to watch. Never block a wrapper on an hours-long task.</li>
<li><code>validateOnly</code> on every item, cascaded by the wrapper: dry-run the whole
stack before touching anything.</li>
<li>Componentise failure: one broken step re-runs alone.</li>
<li>Pre-create DNS; expect HCL friction on virtual hardware; budget for vRO&rsquo;s
content-source lag.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/deployment/deploying-a-new-vmware-cloud-foundation-or-vmware-vsphere-foundation-private-cloud-/use-a-json-specification-to-deploy-vmware-cloud-foundation-or-vmware-vsphere-foundation.html">Use a JSON Specification File to Deploy VMware Cloud Foundation or vSphere Foundation</a>: deploying VCF 9.1 from a JSON spec, which the installer validates first</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/planning-and-preparation/vcf-components-fqdns-and-ip-addresses/first-vcf-instance-fqdns-and-ip-addresses.html">First VCF Instance FQDNs and IP addresses</a>: the FQDNs, static IPs and forward and reverse DNS every component needs</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/vcfa-overview/working-with-the-vcf-automation-catalog/service-broker-adding-content-to-the-catalog/service-broker-add-vrealize-orchestrator-workflows-to-the-catalog.html">Add VCF Operations Orchestrator Client workflows to the VCF Automation catalog</a>: vRO workflows as catalog items, through an Orchestrator content source</li>
<li><a href="https://knowledge.broadcom.com/external/article/408300/vsan-esa-deployment-override-hcl-validat.html">vSAN ESA Deployment: Override HCL Validation for Non-Certified Hardware</a>: the installer&rsquo;s vSAN ESA disk check against the HCL, and the documented override</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/release-notes/vmware-cloud-foundation-9-1-0-0-release-notes/vmware-cloud-foundation-bill-of-materials.html">Bill of Materials 9.1.0</a>: the components and builds of VCF 9.1.0</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/vcfa-overview/getting-started-with-organizations-for-vm-apps-in-vcf-automation/map-head-projects-adding-and-managing-projects/projects-how-do-i-add-a-project-for-my-development-team.html">Add a project for your VCF Automation for VM Apps development team</a>: a project&rsquo;s request Timeout on its Provisioning tab, two hours by default</li>
</ul>
<hr>
<p><em>Lab environment; opinions my own. The automation described builds nested
VCF 9.1 instances for lab and rehearsal use — patterns transfer, specifics
are ours.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>Driving the VCF Installer API from vRO: generate, validate, start, walk away</title>
      <link>https://thenestedlab.com/posts/driving-the-vcf-installer-api-from-vro/</link>
      <pubDate>Wed, 16 Sep 2026 06:10:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/driving-the-vcf-installer-api-from-vro/</guid>
      <description>Stage 2 of the lab factory: a vRO workflow turns an environment number into a VCF 9.1 deployment spec, validates it and starts bringup, because the request dies long before the eight-hour build does.</description>
      <content:encoded><![CDATA[<p>The <a href="/posts/porting-a-powershell-deploy-script/">nested hosts exist</a>. Now
they need to become a VCF instance: vCenter, NSX, SDDC Manager and the fleet
components. The VCF 9.1 Installer appliance does that through an API, from a
deployment spec of a few hundred lines of JSON.</p>
<p>This post is the vRO workflow that drives it. Three design constraints
shaped it:</p>
<ul>
<li>nobody edits the spec by hand;</li>
<li>the request gets two hours by default;</li>
<li>nested hosts fail hardware validation.</li>
</ul>
<p>Unlike stage 1, this isn&rsquo;t VM provisioning, so it&rsquo;s not a cloud template.
It&rsquo;s a <strong>vRO workflow published directly as a catalog item</strong> through an
<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/vcfa-overview/working-with-the-vcf-automation-catalog/service-broker-adding-content-to-the-catalog/service-broker-add-vrealize-orchestrator-workflows-to-the-catalog.html">Orchestrator content source</a>.</p>
<h2 id="the-spec-is-generated-never-edited">The spec is generated, never edited</h2>
<p>A vRO action, <code>buildVcfDeploymentSpec(environment, hostFqdns, labPassword, …)</code>, returns the whole spec as a string. Its structure was reconciled
against a <em>validated</em> export from a real bringup (the installer UI lets
you export the spec it accepted). Everything variable derives from the
environment number X:</p>
<table>
	<thead>
			<tr>
					<th>Element</th>
					<th>Pattern</th>
					<th>f03</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Names</td>
					<td><code>f0X-m01-*</code></td>
					<td><code>f03-m01-vc01.res.lab</code></td>
			</tr>
			<tr>
					<td>Subnets</td>
					<td><code>10.(20+X).&lt;sub&gt;.0/24</code></td>
					<td><code>10.23.1.0/24</code> (mgmt)</td>
			</tr>
			<tr>
					<td>VLANs</td>
					<td><code>2X0&lt;sub&gt;</code></td>
					<td>2301 mgmt … 2306 TEP</td>
			</tr>
			<tr>
					<td>Gateways</td>
					<td><code>.254</code></td>
					<td><code>10.23.1.254</code></td>
			</tr>
			<tr>
					<td>vMotion / vSAN ranges</td>
					<td><code>.1–.16</code></td>
					<td><code>10.23.3.1-16</code></td>
			</tr>
			<tr>
					<td>NSX TEP pool</td>
					<td><code>.6.1–.6.32</code></td>
					<td><code>10.23.6.1-32</code></td>
			</tr>
			<tr>
					<td>SDDC Manager</td>
					<td><code>f0X-vcf01.res.lab</code></td>
					<td><code>f03-vcf01.res.lab</code></td>
			</tr>
	</tbody>
</table>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-javascript" data-lang="javascript"><span class="line"><span class="cl"><span class="kd">var</span> <span class="nx">n</span>   <span class="o">=</span> <span class="nb">parseInt</span><span class="p">(</span><span class="nx">environment</span><span class="p">.</span><span class="nx">substring</span><span class="p">(</span><span class="mi">1</span><span class="p">),</span> <span class="mi">10</span><span class="p">);</span>   <span class="c1">// &#34;f03&#34; -&gt; 3
</span></span></span><span class="line"><span class="cl"><span class="kd">var</span> <span class="nx">pfx</span> <span class="o">=</span> <span class="nx">environment</span> <span class="o">+</span> <span class="s2">&#34;-m01&#34;</span><span class="p">;</span>
</span></span><span class="line"><span class="cl"><span class="kd">var</span> <span class="nx">net</span> <span class="o">=</span> <span class="s2">&#34;10.&#34;</span> <span class="o">+</span> <span class="p">(</span><span class="mi">20</span> <span class="o">+</span> <span class="nx">n</span><span class="p">);</span>
</span></span><span class="line"><span class="cl"><span class="kd">function</span> <span class="nx">vlan</span><span class="p">(</span><span class="nx">o</span><span class="p">)</span> <span class="p">{</span> <span class="k">return</span> <span class="mi">2000</span> <span class="o">+</span> <span class="p">(</span><span class="nx">n</span> <span class="o">*</span> <span class="mi">100</span><span class="p">)</span> <span class="o">+</span> <span class="nx">o</span><span class="p">;</span> <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="kd">function</span> <span class="nx">gw</span><span class="p">(</span><span class="nx">sub</span><span class="p">)</span> <span class="p">{</span> <span class="k">return</span> <span class="nx">net</span> <span class="o">+</span> <span class="s2">&#34;.&#34;</span> <span class="o">+</span> <span class="nx">sub</span> <span class="o">+</span> <span class="s2">&#34;.254&#34;</span><span class="p">;</span> <span class="p">}</span>
</span></span></code></pre></div><p>Some things are static across environments: DNS, NTP, subdomain, component
sizes, vSAN ESA FTT=1, and the component build versions pinned to the
installer binaries.</p>
<p>One lab password feeds every credential field. The UI export scrubs them,
and the action puts them back as the API schema expects.</p>
<p>Two spec-level decisions worth stealing:</p>
<ul>
<li><code>skipEsxThumbprintValidation: true</code> instead of carrying per-host
<code>sslThumbprint</code>. Supported, and the right trade-off for a lab.</li>
<li>Ops and Automation are <strong>checkboxes</strong> that add their blocks to the spec.
The installer only accepts a <code>licenseServerSpec</code> when Ops is present, so
the action adds them together or not at all.</li>
</ul>
<h2 id="the-workflow-authenticate--validate--start--return">The workflow: authenticate → validate → start → return</h2>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">1. POST /v1/tokens                 installer login
</span></span><span class="line"><span class="cl">2. POST /v1/sddcs/validations      the installer&#39;s OWN pre-flight on the spec
</span></span><span class="line"><span class="cl">   (poll until COMPLETED; fail on any FAILED check)
</span></span><span class="line"><span class="cl">3. if validateOnly -&gt; return the validation report; touch nothing
</span></span><span class="line"><span class="cl">4. POST /v1/sddcs                  start bringup -&gt; sddcTaskId
</span></span><span class="line"><span class="cl">5. return { sddcTaskId, installerUrl }
</span></span></code></pre></div><p>Step 2 is the <a href="/posts/validateonly-everywhere/">validateOnly</a> story made
concrete. In seconds, the installer will tell you that
<code>f03-m01-nsx01.res.lab</code> doesn&rsquo;t resolve, that an IP is in use, or that a
host isn&rsquo;t reachable. Two hours into a bringup is a bad time to learn that.
Seconds in, it&rsquo;s merely embarrassing.</p>
<p>Every one of the roughly 40 DNS records the pre-flight wants is created
ahead of time by a one-shot PowerShell script (<code>New-LabEnvDnsRecords.ps1</code>).
DNS is a prerequisite, not a step.</p>
<h2 id="the-two-hour-leash-and-how-to-slip-it">The two-hour leash, and how to slip it</h2>
<p>A request from the catalog carries a token with a roughly <strong>two-hour</strong>
lifetime, and a bringup takes around <strong>eight</strong>. The arithmetic is not
encouraging. When the token runs out, the workflow dies with
<code>Delegating token is not service token</code>.</p>
<p>The project&rsquo;s
<a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/vcfa-overview/getting-started-with-organizations-for-vm-apps-in-vcf-automation/map-head-projects-adding-and-managing-projects/projects-how-do-i-add-a-project-for-my-development-team.html">request timeout</a>
also defaults to two hours, and its Provisioning tab can raise it. We never
tried, because eight hours is too long to hold a request open.</p>
<p>So by default the workflow is fire-and-forget: <code>waitForCompletion=false</code>,
return the task id, and watch progress in the installer UI. A <code>watchTaskId</code>
input lets you re-attach later, and poll an already-running bringup from a
new request.</p>
<p>The wrapper that chains <em>everything</em> (the hosts, then bringup, then the day-N
components that need bringup to be finished) has a neater trick.</p>
<p>The catalog-bound parent deploys the hosts, submits bringup, and then
<strong>re-executes itself as a plain vRO run</strong> (Orchestrator &gt; Run: no catalog
token, no two-hour kill). It carries the hidden <code>bringupWatchTaskId</code> with
it. That continuation polls the installer task to completion (eight hours,
fine), and then runs certificates, fleet items, edge, supervisor and
identity.</p>
<p>Watch it under <em>Orchestrator &gt; Activity &gt; Runs</em>. The catalog request itself
completes in under an hour (47 minutes on the run pictured below), having
<em>started the work correctly and handed off</em>.</p>
<p><img alt="Orchestrator runs: the catalog-bound parent (13:12→14:00) and the continuation it spawned (14:00 → 01:56 next day)" loading="lazy" src="/images/ui/f5-f00-vro-runs-parent-continuation.jpg">
<em>Two rows, one build. The parent returns inside the catalog&rsquo;s window; the continuation waits out the bringup and does the day-N work.</em></p>
<p><img alt="The deployment&rsquo;s stackSummary output: hosts ready, bringup completed, continuation started — watch it in Orchestrator › Activity › Runs" loading="lazy" src="/images/ui/f4-f00-stack-outputs-continuation.jpg"></p>
<h2 id="nested-host-frictions">Nested-host frictions</h2>
<p>Three things a physical bringup never meets:</p>
<ul>
<li><strong>HCL validation vs virtual NVMe.</strong> The installer&rsquo;s hardware
compatibility list (HCL) check blocks the virtual NVMe controller. Fix it
at the vSphere Lifecycle Manager (vLCM) layer:
<code>enforce_hcl_validation=false</code> on the image policy. The vSAN health test
<code>nvmeonhcl</code> also complains. We silence it through the vSAN API,
best-effort, with a manual fallback.</li>
<li><strong>DVS compatibility appears late.</strong> After bringup, NSX takes 1–2 hours
to settle before the supervisor&rsquo;s zones endpoint stops returning 500.
If the supervisor stage fails with &ldquo;No compatible DVS&rdquo; on a fresh
instance, wait and re-run just that item. Patience, it turns out, is a
deployment step.</li>
<li><strong>TSM-SSH.</strong> Bringup wants SSH on the hosts, so the wrapper enables it
host-direct through SOAP before submitting.</li>
</ul>
<h2 id="stale-schema-the-failure-that-looks-like-a-bug-and-isnt">Stale schema: the failure that looks like a bug and isn&rsquo;t</h2>
<p>Add an input to the vRO workflow after the catalog item exists, and three
things happen. The form shows the new field. The request records its value.
And the workflow receives <strong>null</strong>, because Service Broker keeps the old
request schema until the content source re-imports. Two out of three, which
here counts as a fail.</p>
<p>The workflow null-guards every boolean, and aborts with &ldquo;inputs not mapped&rdquo;
rather than running with silently wrong options. The fix: re-import the
content source, confirm the schema, and submit a <em>new</em> request.
Resubmitting an old one reuses the old payload.</p>
<p>There are actually three asynchronous layers between &ldquo;publish&rdquo; and
&ldquo;mappable request&rdquo;:</p>
<ul>
<li>vRO processing the import;</li>
<li>the catalog schema after re-import;</li>
<li>the form service still enforcing the previous custom form for a minute
or two.</li>
</ul>
<p>Same symptom for all three. Check timing before assuming a bug.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>Repeatable, generated VCF deployments matter well beyond a lab: a second
site, a disaster-recovery instance, a new business unit, an environment per
supported release.</p>
<p>Generating the specification from a validated reference removes a whole
class of errors: the ones that come from editing hundreds of lines of JSON
by hand. Running the installer&rsquo;s own validation first turns &ldquo;find out in
hour two&rdquo; into &ldquo;find out in minute one&rdquo;. It&rsquo;s the difference between a VCF
deployment being a project and being a procedure.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li><strong>Generate the spec</strong> from a validated export plus one number. Nobody
hand-edits JSON at 2am.</li>
<li>Run the <strong>installer&rsquo;s own validation</strong> first, and make it a mode you
can request on its own.</li>
<li>Pre-create DNS. Enable SSH. Disable HCL enforcement on virtual NVMe.</li>
<li>Respect the request lifetime: <strong>start, return a task id, re-attach</strong>.
For a long chain, have the workflow re-run itself outside the catalog.</li>
<li>Null-guard every input and fail loudly. Stale schemas are a fact of life
after adding inputs.</li>
<li>On a fresh instance, give NSX an hour before you expect DVS
compatibility.</li>
</ul>
<h2 id="broadcom-documentation">Broadcom documentation</h2>
<ul>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/deployment/deploying-a-new-vmware-cloud-foundation-or-vmware-vsphere-foundation-private-cloud-/use-a-json-specification-to-deploy-vmware-cloud-foundation-or-vmware-vsphere-foundation.html">Use a JSON Specification File to Deploy VMware Cloud Foundation or vSphere Foundation</a>: the deployment spec, its validation and retrying failed tasks; points to the VCF Installer API reference</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/planning-and-preparation/vcf-components-fqdns-and-ip-addresses/first-vcf-instance-fqdns-and-ip-addresses.html">First VCF Instance FQDNs and IP addresses</a>: the FQDNs and forward and reverse DNS records to plan for every component</li>
<li><a href="https://knowledge.broadcom.com/external/article/408300/vsan-esa-deployment-override-hcl-validat.html">vSAN ESA Deployment: Override HCL Validation for Non-Certified Hardware</a>: the installer&rsquo;s vSAN ESA HCL check, and the documented override for non-certified disks</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/vcfa-overview/working-with-the-vcf-automation-catalog/service-broker-adding-content-to-the-catalog/service-broker-add-vrealize-orchestrator-workflows-to-the-catalog.html">Add VCF Operations Orchestrator Client workflows to the VCF Automation catalog</a>: a vRO workflow as a catalog item, through an Orchestrator content source</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/release-notes/vmware-cloud-foundation-9-1-0-0-release-notes/vmware-cloud-foundation-bill-of-materials.html">Bill of Materials 9.1.0</a>: the component versions and builds of VCF 9.1.0</li>
<li><a href="https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/organization-management/vcfa-overview/getting-started-with-organizations-for-vm-apps-in-vcf-automation/map-head-projects-adding-and-managing-projects/projects-how-do-i-add-a-project-for-my-development-team.html">Add a project for your VCF Automation for VM Apps development team</a>: a project&rsquo;s request Timeout on its Provisioning tab, two hours by default</li>
</ul>
<p><em>Part of <a href="/series/the-lab-factory/">The Lab Factory</a>. Previously:
<a href="/posts/porting-a-powershell-deploy-script/">porting the host script</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own. Bringup verified end-to-end on a
rebuilt environment: 305/305 tasks, <code>COMPLETED_WITH_SUCCESS</code>.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
