<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Automation on The Nested Lab</title>
    <link>https://thenestedlab.com/tags/automation/</link>
    <description>Recent content in Automation on The Nested Lab</description>
    <generator>Hugo</generator>
    <language>en-gb</language>
    <lastBuildDate>Wed, 16 Sep 2026 06:40:00 +0100</lastBuildDate>
    <atom:link href="https://thenestedlab.com/tags/automation/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>One catalog item, one VCF instance: building a lab factory</title>
      <link>https://thenestedlab.com/posts/one-catalog-item-one-vcf-instance/</link>
      <pubDate>Wed, 16 Sep 2026 06:40:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/one-catalog-item-one-vcf-instance/</guid>
      <description>How an interactive PowerShell script grew into a catalog-driven factory that stands up complete nested VCF 9.1 instances — hosts, bringup, supervisor, fleet components — from a single request form. The design rules that made it survivable, and the traps that shaped them.</description>
      <content:encoded><![CDATA[<p>Every nested VCF lab starts the same way: a heroic PowerShell script.
Ours was <code>esxihostdeploy.ps1</code> — ovftool plus PowerCLI, an interactive menu
asking which environment, which ESX version, which role, how many hosts.
It worked. It also lived on one person&rsquo;s machine, prompted for credentials,
and knew nothing about everything that comes <em>after</em> the hosts exist.</p>
<p>This post is about what it became: a set of VCF Automation catalog items
where requesting <strong>one form</strong> produces a complete nested VCF 9.1 instance —
ESXi hosts, bringup (vCenter, NSX, SDDC Manager), a vSphere Supervisor,
VCF Automation, Operations, identity — with the environment number as
practically the only real input.</p>
<h2 id="the-shape-of-the-factory">The shape of the factory</h2>
<p>Three stages, each a catalog item, plus a wrapper that chains them:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">Stage 1  Nested ESX Hosts        VM Apps template + vRO actions
</span></span><span class="line"><span class="cl">         (the old script, reborn declaratively)
</span></span><span class="line"><span class="cl">Stage 2  Deploy VCF 9.1 Instance vRO workflow driving the VCF Installer API
</span></span><span class="line"><span class="cl">         (spec generated, validated, bringup started)
</span></span><span class="line"><span class="cl">Day-N    Supervisor · NSX Edge · VCF Automation · Ops Logs/Networks/RTM ·
</span></span><span class="line"><span class="cl">         Identity (AD)           one catalog item each
</span></span><span class="line"><span class="cl">Wrapper  &#34;Deploy VCF Stack&#34;      one form, checkbox per component
</span></span></code></pre></div><p><img alt="The factory catalog: hosts, bringup, every day-N component, and the wrapper — ten tiles" loading="lazy" src="/images/ui/f1-f00-factory-catalog.jpg"></p>
<p>The wrapper&rsquo;s form has a checkbox per component; ticking one reveals that
component&rsquo;s tab with every field pre-populated. One lab password feeds every
credential. Tick everything, click request, and go get coffee.</p>
<h2 id="rule-1-derive-everything-from-one-number">Rule 1: derive everything from one number</h2>
<p>Each lab environment is <code>f0X</code>, and <em>everything</em> scales from X by formula:</p>
<table>
	<thead>
			<tr>
					<th>Element</th>
					<th>Pattern</th>
					<th>f03 example</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Names</td>
					<td><code>f0X-m01-*</code></td>
					<td><code>f03-m01-vc01.res.lab</code></td>
			</tr>
			<tr>
					<td>Subnets</td>
					<td><code>10.(20+X).&lt;sub&gt;.0/24</code></td>
					<td><code>10.23.1.0/24</code> (mgmt)</td>
			</tr>
			<tr>
					<td>VLANs</td>
					<td><code>2X0n</code></td>
					<td>2307 (edge TEP)</td>
			</tr>
	</tbody>
</table>
<p>The bringup spec — hundreds of lines of JSON the VCF Installer wants — is
generated by a vRO action from a known-good reference spec plus X. Nobody
edits a deployment spec by hand, which means nobody typo-breaks a bringup at
2am. When the old script did this, the formulas lived in string
concatenation; now they live in one action with the reference spec beside it.</p>
<p>The same philosophy carried into stage 1: the script&rsquo;s &ldquo;next free esxNN
index&rdquo; scan became a vRO action bound to the request form, its VLAN/IP
arithmetic became template expressions, its <code>--prop:guestinfo.*</code> flags
became <code>ovfProperties</code> in the template. Porting a script isn&rsquo;t rewriting
it — it&rsquo;s finding the declarative home for each behaviour.</p>
<h2 id="rule-2-never-wait-for-anything-you-can-watch-instead">Rule 2: never wait for anything you can watch instead</h2>
<p>The hard constraint that shaped the whole design: a VCF Automation request
gets about <strong>two hours</strong> before the platform gives up on it. A full VCF
bringup takes longer than that. So the wrapper <em>never waits</em>:</p>
<ul>
<li>Bringup is <strong>fire-and-forget</strong> — the workflow authenticates to the
installer, validates the spec, starts the task, and hands back a
<code>watchTaskId</code>. Re-attach any time to check on it.</li>
<li>Fleet deployments (VCF Automation, Ops for Logs/Networks, metrics) are
server-side tasks; the items submit with <code>waitForCompletion=false</code>.</li>
<li>The supervisor item submits enablement and returns; vCenter carries on.</li>
<li>Only fast, deterministic steps (identity configuration, minutes) run to
completion inside the request.</li>
</ul>
<p><img alt="F06-Mgmt-VCF: the wrapper deployment, Create Successful, 13:12 → 14:05" loading="lazy" src="/images/ui/f3-f00-stack-deployment-success.jpg">
<em>A whole VCF instance as one deployment record — the request finished in under an hour while the build ran on for twelve.</em></p>
<p>Result: a full-stack kick-off <em>completes</em> as a request in well under an hour (53 minutes on the run pictured),
while the actual multi-hour build continues as watchable server-side tasks.
The request&rsquo;s job isn&rsquo;t to do the work — it&rsquo;s to <strong>start the work
correctly</strong> and tell you where to watch it.</p>
<p>The corollary: a failed component is recorded and the remaining components
still run. You fix one thing and re-run one item, not the world.</p>
<h2 id="rule-3-plan-mode-for-infrastructure">Rule 3: plan mode for infrastructure</h2>
<p>Every item in the chain supports <code>validateOnly</code> — and the wrapper cascades
it. Tick everything, set validateOnly, and the entire stack is <em>planned</em>
against the live environment with zero changes: specs generated,
prerequisites checked, name/IP collisions caught. A smoke runner exercises
exactly this on every change to the automation itself.</p>
<p>If you build nothing else into your lab automation, build this. The number
of 2am bringups saved by a five-minute dry run is not small.</p>
<h2 id="the-traps-that-shaped-the-rules">The traps that shaped the rules</h2>
<p>Some of the design above exists because of scars:</p>
<ul>
<li><strong>Hardware validation hates virtual NVMe.</strong> Bringup&rsquo;s HCL check will
block nested hosts; the spec generator has to account for it, or you
discover it two hours in — twice, if you&rsquo;re us.</li>
<li><strong>Small disks, surprising layouts.</strong> Nested hosts with 64 GB disks ship
ESX-OSDATA at essentially the whole disk; a post-provision step relocates
scratch or stage-2 fills the disk with logs.</li>
<li><strong>DNS is a prerequisite, not a step.</strong> The installer&rsquo;s pre-flight wants
every record resolvable before it starts; a one-shot script creates the
per-environment records ahead of the request.</li>
<li><strong>vRO&rsquo;s content-source lag.</strong> A new or changed workflow takes 15–20
minutes of data-collection before the catalog sees it. Publish, wait,
<em>then</em> test — or you&rsquo;ll debug a ghost.</li>
<li><strong>Wrapper inputs are duplicated by necessity.</strong> vRO requires every
sub-workflow input to be passed explicitly, so adding an input to a
component means updating the wrapper&rsquo;s call too. Null-guards in each
component turn a forgotten field into a loud failure instead of a silent
default.</li>
</ul>
<h2 id="why-bother">Why bother?</h2>
<p>Because the payoff compounds. Once a full VCF instance is a catalog request,
everything downstream changes character: upgrade rehearsals happen on
freshly-built instances instead of precious pets; a broken environment is
redeployed, not repaired; and the lab stops being a snowflake collection and
becomes a <em>product</em> — versioned, validated, reproducible.</p>
<p><img alt="Three pods, identical IP plans, no route between them" loading="lazy" src="/images/product-01-hook.jpg">
<em>Where this is heading: the factory&rsquo;s output feeding per-student pods with identical addressing.</em></p>
<p>The factory&rsquo;s next customers, funnily enough, are the isolated VPC pods from
<a href="/series/the-vpc-pod-papers/">the other series on this blog</a> — same
philosophy, one layer further down.</p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>A complete VCF instance from one form changes what an environment costs
to have. Environments that used to be precious — because building one took
a week — become disposable, and a lot follows from that:</p>
<ul>
<li><strong>Proofs of concept</strong> run on an environment built for the customer&rsquo;s
scenario, not on whatever happens to be free.</li>
<li><strong>Upgrade and migration rehearsals</strong> happen on a fresh instance of the
right version, then it&rsquo;s deleted.</li>
<li><strong>Training and enablement</strong> get a real VCF per person or per team.</li>
<li><strong>Reference builds</strong> exist for every supported release, on demand.</li>
</ul>
<p>This is how Comms-care provides a dedicated instance to every consultant.
The same factory, pointed at a customer&rsquo;s requirements, is a repeatable
way to deliver environments rather than a one-off project each time.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Derive names, subnets and VLANs from a single environment number; generate
specs, never hand-edit them.</li>
<li>Respect the request-duration ceiling: start long work, return a task
handle, re-attach to watch. Never block a wrapper on an hours-long task.</li>
<li><code>validateOnly</code> on every item, cascaded by the wrapper — dry-run the whole
stack before touching anything.</li>
<li>Componentise failure: one broken step re-runs alone.</li>
<li>Pre-create DNS; expect HCL friction on virtual hardware; budget for vRO&rsquo;s
content-source lag.</li>
</ul>
<hr>
<p><em>Lab environment; opinions my own. The automation described builds nested
VCF 9.1 instances for lab and rehearsal use — patterns transfer, specifics
are ours.</em></p>
]]></content:encoded>
    </item>
    <item>
      <title>validateOnly everywhere: plan mode for infrastructure</title>
      <link>https://thenestedlab.com/posts/validateonly-everywhere/</link>
      <pubDate>Wed, 16 Sep 2026 06:30:00 +0100</pubDate>
      <guid>https://thenestedlab.com/posts/validateonly-everywhere/</guid>
      <description>Terraform has plan. Kubernetes has &amp;ndash;dry-run. Your vRO workflows have nothing — unless you give them a validateOnly input and make the wrapper cascade it. A short argument for the single most valuable checkbox in the lab factory, with the failures it caught.</description>
      <content:encoded><![CDATA[<p>Terraform has <code>plan</code>. Kubernetes has <code>--dry-run=server</code>. Ansible has
<code>--check</code>. Every mature infrastructure tool grew a way to say &ldquo;tell me what
you&rsquo;d do, then don&rsquo;t&rdquo; — because the alternative is finding out at 2am, two
hours into a bringup, that a hostname doesn&rsquo;t resolve.</p>
<p>vRO workflows don&rsquo;t come with one. This is the case for adding it to every
single one you write, and cascading it through every wrapper.</p>
<h2 id="the-shape">The shape</h2>
<p>Every catalog item in <a href="/posts/one-catalog-item-one-vcf-instance/">the lab factory</a>
has a boolean input, <code>validateOnly</code>, default false. When true the workflow
does <em>everything it can without changing anything</em>:</p>
<ul>
<li>authenticate to every endpoint it would touch</li>
<li>resolve every name it would use, and fail on the ones that don&rsquo;t</li>
<li>generate every spec it would submit, and run the target&rsquo;s own validation
API on it where one exists (the VCF Installer has one; use it)</li>
<li>check for collisions — names, IPs, existing objects</li>
<li>report what it <em>would</em> have created, then return <code>CREATE_SUCCESSFUL</code></li>
</ul>
<p>The wrapper — the one form that chains hosts, bringup, supervisor, fleet
components, identity — has the same checkbox, and <strong>cascades</strong> it to every
component. Tick everything, tick validateOnly, request. Thirty seconds to
a few minutes later you have a full-stack plan against the <em>live</em>
environment, and nothing has moved.</p>
<h2 id="what-it-caught">What it caught</h2>
<p>Not hypothetically. On a built environment, the cascaded dry run of the
whole stack reported:</p>
<ul>
<li>edge cluster: <strong>already exists</strong> — correctly recorded, wrapper carried on</li>
<li>Ops for Logs: <strong>IP_IN_USE</strong> on the planned address — right, it&rsquo;s deployed</li>
<li>Ops for Networks: same</li>
<li>supervisor: the existing one would be reused; the per-service plan
listed which services were already active</li>
<li>identity: bind succeeded, group resolved, no changes needed</li>
</ul>
<p>That&rsquo;s a plan output. On a <em>fresh</em> environment the same run has caught, at
various times: a DNS record missing for one of ~40 required names (the
installer&rsquo;s own pre-flight found it, in seconds, instead of bringup
finding it in hour two); a stale content-library image ID; a form field
arriving <code>null</code> because a custom form hadn&rsquo;t finished re-importing — which
is a <em>publishing</em> bug the dry run surfaced before anyone requested
anything real.</p>
<h2 id="the-argument-against-answered">The argument against, answered</h2>
<p>&ldquo;It doubles the code.&rdquo; It doesn&rsquo;t — it moves the <code>if (!validateOnly)</code> guard
around the mutating call, and the validation logic is code you should have
had anyway. What it <em>does</em> force is separating &ldquo;compute what to do&rdquo; from
&ldquo;do it&rdquo;, which is how the workflows should have been structured in the
first place.</p>
<p>&ldquo;Some things can&rsquo;t be validated without doing them.&rdquo; True. Say so in the
result summary — &ldquo;would deploy X; no pre-validation available&rdquo; — rather
than skipping the item. Partial plans are still plans.</p>
<p>&ldquo;We have a test environment.&rdquo; You have <em>a</em> test environment. A dry run
against the <em>target</em> is what catches the collision with the thing that&rsquo;s
already there.</p>
<h2 id="make-it-the-smoke-test">Make it the smoke test</h2>
<p>The best consequence: a validateOnly request against a known environment
is a <strong>regression test for the automation itself</strong>, runnable on every
change. The factory&rsquo;s smoke runner does exactly this — request every item
with <code>validateOnly: true</code>, assert <code>CREATE_SUCCESSFUL</code>, diff the plan
summary against the last run. It takes minutes and it has caught more
bugs in the workflows than any amount of code review.</p>
<p><img alt="Deploy VCF Stack request form: one checkbox per component, and validateOnly" loading="lazy" src="/images/ui/f2-f00-stack-form-validateonly.jpg">
<em>The same form, real or dry-run. One checkbox decides.</em></p>
<h2 id="why-this-matters-outside-the-lab">Why this matters outside the lab</h2>
<p>For anyone who has sat through a failed change window, the value is
obvious: a full dry run against the <em>real</em> estate before anything moves.
Fewer failed changes, shorter windows, and a plan output that answers the
change board&rsquo;s questions before they&rsquo;re asked. It also gives auditors
something they rarely get from infrastructure automation — evidence of what
was going to happen, produced by the same tooling that then did it.</p>
<h2 id="rules-learned">Rules learned</h2>
<ul>
<li>Add <code>validateOnly</code> to <strong>every</strong> workflow. Default false. Wrappers
cascade it.</li>
<li>Dry-run does everything but mutate: auth, resolve, generate, call the
target&rsquo;s validator, check collisions, report.</li>
<li>Where a step truly can&rsquo;t be pre-validated, <em>say so</em> in the summary.
Never skip it silently.</li>
<li>A dry run against the real target is a plan. A dry run on every change
is a smoke test. Same checkbox.</li>
<li>The refactor it forces — compute, <em>then</em> act — is the one you wanted.</li>
</ul>
<p><em>Part of <a href="/series/the-lab-factory/">The Lab Factory</a>.</em></p>
<hr>
<p><em>Lab environment; opinions my own.</em></p>
]]></content:encoded>
    </item>
  </channel>
</rss>
