Nutanix: Invisible Wires: Agentic Cloud
NC2 across AWS, Azure and Google Cloud: active/active/standby on the v4 API, with a native service beside each cluster
The point of this post is narrower than the diagram makes it look: a Nutanix UVM is just a VM on a standard VPC or VNet subnet, so a native cloud service next to it is one networking hop away, not a Nutanix integration project. The active/active/standby pattern below is the proof, because DR is the workload most likely to expose sloppy networking — if a data lake, a Function and a BI tool can keep working across a live failover without a single Nutanix-specific hook, the easy case (a UVM calling S3 on a Tuesday) was never in question.
The three build posts (AWS, Azure, Google Cloud) each stopped at one running VM. The question underneath all three was always the next one: put a workload on all three, keep two of them serving and one caught up and cold, and do the failover with the same v4 API that built the VM. Nutanix's cross-cluster building blocks are protection policies and recovery plans, defined in the datapolicies namespace and executed through dataprotection, spoken over Prism Central's v4 API exactly like the vmm and networking calls in the earlier posts. The API never names a cloud, only Prism Centrals and clusters. The documentation is narrower: the Nutanix Disaster Recovery guide lists NC2 to NC2 recovery plans as "(Same cloud)" in its recovery plan support table, with no cross-cloud row, so a Google Cloud standby for AWS and Azure (and the AWS to Azure leg between the two active sides) is my design, to validate with Nutanix; I found no documented cross-cloud pairing.
Worth being honest about where those earlier posts actually left their customers. The AWS post built the cluster and Prism Central through the NC2 console (a UI, not the v4 API) and only picked up v4 for what came after: VPCs, subnets, the VM. The Azure post went further on automation (Bicep, no console click) but the organization, cloud account and cluster still came from the NC2 v2 API, with v4 arriving at the same point AWS's did. Neither customer has fully migrated: the lifecycle layer (org, account, cluster) is still v1/v2, and only the operational layer (network, compute, storage) is v4. Only the Google Cloud day-one post ran the whole sequence, cluster included, on v4 alone. This post's DR half (Prism Central pairing, protection policies, recovery plans) is v4 on all three clouds regardless of how the cluster itself was born, because it runs on Prism Central, not on the NC2 lifecycle API.
Why the migration is worth finishing
- One API family, every layer. Today an AWS or Azure customer scripts the cluster in one API generation and the network/compute/storage in another — two auth flows, two response shapes, two things to version. Full v4 collapses that to one.
- One task-polling model. Every v4 mutation in this series — VM create, VPC create, protection policy, recovery plan — returns the same task envelope, polled the same way. v1/v2 console operations do not; they are UI state, not a resource you can poll from your own pipeline.
- The native-service story gets easier, not harder. Because v4
networkingis what defines the subnets a UVM lives in, the same IaC that makes those calls can also create the S3 gateway endpoint, the Event Grid topic's private link, or the BigQuery-facing service connection: one pipeline, not a handoff from a console-driven network to a hand-built peering. - Idempotent and diffable: the part that actually changes how a team operates. A v4 resource body is the full desired state, not a verb, and the API is built to be retried: every write carries an
NTNX-Request-Id, which the spec calls "an idempotence token for safely retrying requests," and every update carriesIf-Match, so a stale one fails with a412instead of overwriting a newer change. That is what lets a tool like OpenTofu compute a plan (additions, changes, destroys) before anything happens, against the live cluster, not against a changelog someone kept by hand. Console clicks and v2's imperative calls ("create this," "attach that") have no such plan step: there is no diff to review, only an audit log after the fact of what already happened. Three practical effects follow directly: drift detection (a manual console change shows up as a diff on the nextplan, instead of silently rotting the IaC's picture of reality); safe re-runs (a failed apply, a flaky network call, a retried CI job: the tool reads live state first and the API takes retries, so re-running is the fix, not a new problem to reason about); and code review for infrastructure changes (aplanoutput is a diff, which means it can go through the same pull-request review as application code, instead of a screen-share of someone's console session). The AWS and Azure posts both hit this wall exactly at the console/v2 boundary: everything after cluster creation is diffable; the creation step itself is not, and has to be treated as a one-time, undiffable fact the rest of the pipeline works around.
Why this matters most to a customer vacating an SDDC
The customer this pattern is actually for is not greenfield — it is the one mid-migration off a VMware SDDC, still running vSphere/NSX/vSAN on one side and standing NC2 up on the other, for as long as that takes. Three things about v4 matter disproportionately to exactly that customer:
- The migration itself becomes a diffable plan, not a runbook. A
planagainst v4 resources shows precisely what a wave of VMs moving off the SDDC will create, in advance, reviewable before a single workload actually cuts over — the same confidence aterraform plangives an AWS migration, applied to the Nutanix side of a hybrid estate that used to have no plan step on either end. - Idempotency forgives a migration's inevitable false starts. Migration waves get retried, paused, and re-run against partially-moved state more than almost any other IaC workload; a plan that finds an already-created VM unchanged, plus a retry the API recognizes by its
NTNX-Request-Id, is the difference between a safe re-run and a bespoke cleanup script every time a wave stalls. - One API for the destination end of the migration, matching the source end's own IaC. A team already scripting the SDDC side with Terraform/vSphere providers is not learning a new operational model for the NC2 side — v4's declarative, pollable, diffable shape is the same shape, so the migration tooling on both ends of the move looks like one discipline instead of two.
This is also, not incidentally, where the DR pattern above stops being a nice-to-have: a customer vacating an SDDC rarely gets to cut over instantly, so an SDDC-side workload and its NC2-side replacement often have to run as an active/active or active/standby pair for the migration window itself, before the SDDC is ever decommissioned. Where the on-prem leg already runs on Nutanix under Prism Central, that's the same protection policy and recovery plan shape, with one leg on-prem instead of a third cloud, and on-prem to NC2 is a pairing the DR guide documents. A vSphere/vSAN leg isn't a Prism Central location, so those VMs reach NC2 through Nutanix Move, whose 6.3 guide lists ESXi to NC2 on all three clouds.
The second half of this post is the other direction: what the clouds are for, once NC2 is not doing it for you, and how little stands between a UVM and each one. A data lake (AWS), an event-driven function (Azure), a BI tool over a warehouse (Google Cloud): three services with no Nutanix equivalent and no reason to fake one. They sit beside the cluster, reading and writing over the cloud's own VPC and VNet routing to the subnets the v4 networking calls already defined, never inside the failover path, so a region failing over does not also fail over your BigQuery-backed dashboard, and never touched by anything Nutanix-specific, because there is nothing Nutanix-specific to touch.
Everything below is a design pattern built from the published v4 OpenAPI specifications (datapolicies and dataprotection at v4.1 and v4.2, prism, networking, vmm and clustermgmt at v4.1), the Nutanix Disaster Recovery guide (pc.7.6) and the NC2 Deployment and User Guides for each cloud, read 2026-09-24. The recovery plan calls are v4.2 and list a minimum of Prism Central 7.5 and Prism Element 7.5. I have not run this exact three-cloud sequence against three live clusters, and I am not writing it as though I had; where a guide and a block here disagree on your version, the guide wins.
The map
┌─ AWS · us-west-2 ── ACTIVE ─────────────┐ ┌─ Azure · westus2 ── ACTIVE ──────────────┐
│ NC2 cluster + Prism Central A │ │ NC2 cluster + Prism Central B │
│ └─ App/DB VMs, serving traffic │ │ └─ App/DB VMs, serving traffic │
│ beside it, not in the failover path: │ │ beside it, not in the failover path: │
│ S3 + Glue + Athena (data lake) ◀────┼─VPC─┤ Event Grid ─▶ Azure Function (FaaS) │
└───────────────┬──────────────────────────┘ └───────────────┬────────────────────────────┘
│ datapolicies v4.1: protection policy │
│ (sync/async), recovery points │
▼ ▼
┌──────────────────────────────────────────────────────────────┐
│ PC-to-PC availability zone pairing (mutual) │
│ A ⇄ B replicate to each other; both replicate │
│ to the standby below │
└───────────────────────────┬────────────────────────────────┘
│ datapolicies v4.2: recovery plans
│ (ordered VM boot, network mapping)
▼
┌─ Google Cloud · us-west1 ── STANDBY ───────────────────┐
│ NC2 cluster + Prism Central C │
│ └─ Recovery points only; VMs restored at failover │
│ beside it, reachable once promoted: │
│ BigQuery / Cloud SQL ─▶ Looker (BI) │
└─────────────────────────────────────────────────────────┘
Failover: recovery plans run on C ─▶ power on in stage priority order ─▶
re-IP or re-attach floating IP ─▶ traffic follows ─▶ A/B unregister
Failback: resync, then the same plans run from A and B
Step 1: three clusters that already exist
Nothing here builds a cluster; that is the three earlier posts, one per cloud:
- NC2 on AWS, as code — NC2 console + OpenTofu on the v4 API.
- NC2 on Azure, as code — Bicep landing zone + NC2 v2 API.
- NC2 on Google Cloud, day one — pure v4 API, no console.
The precondition for everything below: three Prism Central instances, each answering on :9440, each with at least one registered cluster and a VM image already in place.
Step 2: pair the Prism Centrals (availability zones)
Cross-cluster DR in Prism Central is built on "availability zone" pairing between Prism Centrals, not between clusters directly. Each pair is a one-time trust setup, done from either side: PC A and PC B register each other as a remote availability zone, then PC B and PC C do the same, then A and C. The console does it under Administration > Availability Zones; the v4 prism namespace has the matching call, POST /api/prism/v4.1/management/domain-managers/{extId}/$actions/register, which the spec describes as registering "a domain manager (Prism Central) instance to other entities like PE and PC," with a DomainManagerRemoteClusterSpec body that takes the remote Prism Central's address and credentials. Each Prism Central's own extId comes from GET /api/prism/v4.1/config/domain-managers, and the policies and plans below name the Prism Centrals by those IDs. Three registrations for a three-way mesh:
PC A (AWS) ───pair─── PC B (Azure)
\ /
\__pair────pair___/
PC C (GCP)
Step 3: protection policies, v4 datapolicies
A protection policy is what makes a VM's disks eligible for replication and states where recovery points land and how often. For the two active legs replicating into the standby, the policy is asynchronous (RPO measured in minutes, not synchronous mirroring — cross-cloud WAN latency rules out zero-RPO here):
POST https://<PC-A>:9440/api/datapolicies/v4.1/config/protection-policies
{
"name": "aws-to-gcp-standby",
"replicationLocations": [
{ "label": "PC-A-AWS", "domainManagerExtId": "<PC A extId>", "isPrimary": true },
{ "label": "PC-C-GCP", "domainManagerExtId": "<PC C extId>" }
],
"replicationConfigurations": [
{ "sourceLocationLabel": "PC-A-AWS", "remoteLocationLabel": "PC-C-GCP",
"schedule": { "recoveryPointObjectiveTimeSeconds": 900 } },
{ "sourceLocationLabel": "PC-C-GCP", "remoteLocationLabel": "PC-A-AWS",
"schedule": { "recoveryPointObjectiveTimeSeconds": 900 } }
],
"categoryIds": [ "<extId of app-tier:active-active-core>" ]
}
The same body runs against PC B with PC-B-Azure as the primary location. Two policies, one target, both async, both scoped by categoryIds to the app-tier:active-active-core category, so any VM tagged with it is covered without a per-VM edit later. The spec asks for both directions in replicationConfigurations ("Connections from both source-to-target and target-to-source should be specified"), so the C-to-A leg that failback needs is in the policy from the start.
Step 4: recovery plans, v4.2 datapolicies
The recovery plan is the failover script: which VMs power on, in what order, on what network, at the standby side. It is a resource, not a one-shot action, so it can be validated long before it is ever executed. A plan has one primary and one recovery location, so this design needs two (AWS to Google Cloud, Azure to Google Cloud), and the DR guide's advice is to create them at the primary site, PC A and PC B, from where they synchronize to the recovery side:
POST https://<PC-A>:9440/api/datapolicies/v4.2/config/recovery-plans
{
"name": "promote-gcp-from-aws",
"primaryLocation": { "domainManagerExtId": "<PC A extId>" },
"recoveryLocation": { "domainManagerExtId": "<PC C extId>" }
}
Stages and network mappings are sub-resources of the plan. A stage selects VMs by the same category and sets their power-on priority (1 first; on a planned failover the spec shuts the highest number down first):
POST /api/datapolicies/v4.2/config/recovery-plans/{recoveryPlanExtId}/stages
{
"entityType": "VM",
"categoryExtIds": [ "<extId of app-tier:active-active-core>" ],
"priority": 1,
"postActions": [ { "config": { "$objectType": "datapolicies.v4.config.DelayAction", "delaySecs": 120 } } ]
}
Network mappings go to .../network-mappings the same way, one primaryNetwork and recoveryNetwork pair per subnet. Validate before you ever need it. The action lives in dataprotection, and like the failovers it takes a job name and the direction between the two Prism Centrals:
POST /api/dataprotection/v4.2/operations/recovery-plans/{recoveryPlanExtId}/$actions/validate
{
"name": "validate-aws-to-gcp",
"failoverDirections": [
{ "sourceDomainManagerExtId": "<PC A extId>", "targetDomainManagerExtId": "<PC C extId>" }
]
}
Rehearse with .../$actions/test-failover (same body), then .../$actions/clean-up-resources, which the spec describes as deleting the "entities recovered in the last test failover action."
Step 5: the failover call, and the poll
Executing a recovery plan is one call, made on the recovery side (PC C), and like every other v4 mutation in this series it returns a task to poll, not a result:
POST https://<PC-C>:9440/api/dataprotection/v4.2/operations/recovery-plans/{recoveryPlanExtId}/$actions/planned-failover
{
"name": "failover-aws-to-gcp",
"failoverDirections": [
{ "sourceDomainManagerExtId": "<PC A extId>", "targetDomainManagerExtId": "<PC C extId>" }
]
}
planned-failover assumes the primary is reachable; in the spec's words, "VMs are powered off on the source before migrating it to the target domain manager." unplanned-failover takes the same body plus an optional recoveryReferenceTime, restores "from the recovery points on the target domain manager," and is the one you use when AWS and Azure are both actually down, not drilled. Poll /api/prism/v4.1/config/tasks/{extId} the same way the single-cloud posts did until status: SUCCEEDED. The VMs come up on PC C in stage priority order, and floating IPs re-attach on the Google Cloud side per the earlier GCP post's networking calls.
Failback reuses the same pieces. The policy already carries the C-to-A leg, and the DR guide says "the same recovery plan applies to both the failover and the failback operations," with failback run at the primary site: the same planned failover, called on PC A with the direction reversed, once the original region is confirmed healthy and resynced.
The three native services — beside the cluster, never inside the API path
None of these are Nutanix resources. That is the point: NC2's v4 API stops at the VM, the disk and the network. Everything below reads from or writes to the VMs over a VPC/VNet endpoint, on the active side only, and is entirely absent from the failover call above.
AWS (active) ── VM writes events ──▶ S3 (raw) ──▶ Glue Catalog ──▶ Athena
"the data lake": land it, catalog it, query it in place
Azure (active) ── VM emits ──▶ Event Grid topic ──▶ Azure Function
"FaaS": react to an event, no server to patch or fail over
Google Cloud (standby, promoted) ── BigQuery / Cloud SQL ──▶ Looker
"the BI layer": stays cold with the cluster, wakes with it
- AWS data lake: the active-side VMs write append-only events to an S3 bucket in the same VPC (gateway endpoint, no NAT egress); a Glue crawler catalogs the schema; Athena queries it directly. No compute to fail over — S3 and Glue Catalog are already regional-durable, so this leg needs nothing from the DR plan above.
- Azure Function: the active-side VMs publish to an Event Grid topic; a Function subscribes and does the small stateless thing (notify, transform, write a row) that would be wasteful to run on a VM 24/7. If Azure is the side that fails over, the Function goes dark with it — deliberately not part of the recovery plan, because there is nothing to recover, only to redeploy from its own IaC once the region is back.
- Google Looker: sits over BigQuery/Cloud SQL on the Google Cloud side. While GCP is standby this is quiet; the moment a recovery plan promotes PC C, the underlying data source is live NC2 VMs again and Looker's dashboards start reflecting them; no Looker-specific step in the failover.
What this pattern does not claim
Three-way availability-zone pairing, two async protection policies and two recovery plans is the minimum shape for active/active/standby; it is not a substitute for reading your own RPO/RTO requirement against async replication's real floor (minutes, driven by cross-cloud WAN, not seconds). Nor is it a documented Nutanix configuration: the DR guide's recovery plan table lists NC2 to NC2 within the same cloud, so the cross-cloud legs are my design, to validate with Nutanix before anyone relies on them. A drill is test-failover and clean-up-resources; planned-failover moves production, and unplanned-failover on real data loss is a different operation with a different blast radius, a last resort, not a monthly exercise. And the native-service half is illustrative, not exhaustive: any AWS/Azure/Google Cloud managed service that talks to a VPC endpoint fits the same "beside the cluster, not inside the API path" rule.
Personal blog. Alex works at Nutanix; the opinions here are his own and nothing here is Nutanix confidential: every fact is public or his own field experience.
Follow along
Get the next post without leaving this page: copy the feed into any reader, or get it by email. Live builds and the lab's music stream on Twitch.
Comments