Sign in →
CloudLab Works emblem: Waku the orca ringed by CloudLab and WorksCLOUDLAB WORKScrossed whale bones, one end a wrenchCloud City, headwaters of the agentic cloud revolution!

Nutanix: Invisible Wires: Agentic Cloud

NC2 across AWS, Azure and Google Cloud: active/active/standby on the v4 API, with a native service beside each cluster

· By Alex Alvord

The point of this post is narrower than the diagram makes it look: a Nutanix UVM is just a VM on a standard VPC or VNet subnet, so a native cloud service next to it is one networking hop away, not a Nutanix integration project. The active/active/standby pattern below is the proof, because DR is the workload most likely to expose sloppy networking — if a data lake, a Function and a BI tool can keep working across a live failover without a single Nutanix-specific hook, the easy case (a UVM calling S3 on a Tuesday) was never in question.

The three build posts (AWS, Azure, Google Cloud) each stopped at one running VM. The question underneath all three was always the next one: put a workload on all three, keep two of them serving and one caught up and cold, and do the failover with the same v4 API that built the VM. Nutanix's cross-cluster building blocks are protection policies and recovery plans, defined in the datapolicies namespace and executed through dataprotection, spoken over Prism Central's v4 API exactly like the vmm and networking calls in the earlier posts. The API never names a cloud, only Prism Centrals and clusters. The documentation is narrower: the Nutanix Disaster Recovery guide lists NC2 to NC2 recovery plans as "(Same cloud)" in its recovery plan support table, with no cross-cloud row, so a Google Cloud standby for AWS and Azure (and the AWS to Azure leg between the two active sides) is my design, to validate with Nutanix; I found no documented cross-cloud pairing.

Worth being honest about where those earlier posts actually left their customers. The AWS post built the cluster and Prism Central through the NC2 console (a UI, not the v4 API) and only picked up v4 for what came after: VPCs, subnets, the VM. The Azure post went further on automation (Bicep, no console click) but the organization, cloud account and cluster still came from the NC2 v2 API, with v4 arriving at the same point AWS's did. Neither customer has fully migrated: the lifecycle layer (org, account, cluster) is still v1/v2, and only the operational layer (network, compute, storage) is v4. Only the Google Cloud day-one post ran the whole sequence, cluster included, on v4 alone. This post's DR half (Prism Central pairing, protection policies, recovery plans) is v4 on all three clouds regardless of how the cluster itself was born, because it runs on Prism Central, not on the NC2 lifecycle API.

Why the migration is worth finishing

Why this matters most to a customer vacating an SDDC

The customer this pattern is actually for is not greenfield — it is the one mid-migration off a VMware SDDC, still running vSphere/NSX/vSAN on one side and standing NC2 up on the other, for as long as that takes. Three things about v4 matter disproportionately to exactly that customer:

This is also, not incidentally, where the DR pattern above stops being a nice-to-have: a customer vacating an SDDC rarely gets to cut over instantly, so an SDDC-side workload and its NC2-side replacement often have to run as an active/active or active/standby pair for the migration window itself, before the SDDC is ever decommissioned. Where the on-prem leg already runs on Nutanix under Prism Central, that's the same protection policy and recovery plan shape, with one leg on-prem instead of a third cloud, and on-prem to NC2 is a pairing the DR guide documents. A vSphere/vSAN leg isn't a Prism Central location, so those VMs reach NC2 through Nutanix Move, whose 6.3 guide lists ESXi to NC2 on all three clouds.

The second half of this post is the other direction: what the clouds are for, once NC2 is not doing it for you, and how little stands between a UVM and each one. A data lake (AWS), an event-driven function (Azure), a BI tool over a warehouse (Google Cloud): three services with no Nutanix equivalent and no reason to fake one. They sit beside the cluster, reading and writing over the cloud's own VPC and VNet routing to the subnets the v4 networking calls already defined, never inside the failover path, so a region failing over does not also fail over your BigQuery-backed dashboard, and never touched by anything Nutanix-specific, because there is nothing Nutanix-specific to touch.

Everything below is a design pattern built from the published v4 OpenAPI specifications (datapolicies and dataprotection at v4.1 and v4.2, prism, networking, vmm and clustermgmt at v4.1), the Nutanix Disaster Recovery guide (pc.7.6) and the NC2 Deployment and User Guides for each cloud, read 2026-09-24. The recovery plan calls are v4.2 and list a minimum of Prism Central 7.5 and Prism Element 7.5. I have not run this exact three-cloud sequence against three live clusters, and I am not writing it as though I had; where a guide and a block here disagree on your version, the guide wins.

The map

 ┌─ AWS · us-west-2 ── ACTIVE ─────────────┐     ┌─ Azure · westus2 ── ACTIVE ──────────────┐
 │  NC2 cluster + Prism Central A          │     │  NC2 cluster + Prism Central B            │
 │   └─ App/DB VMs, serving traffic        │     │   └─ App/DB VMs, serving traffic          │
 │  beside it, not in the failover path:   │     │  beside it, not in the failover path:     │
 │   S3 + Glue + Athena  (data lake)  ◀────┼─VPC─┤   Event Grid ─▶ Azure Function  (FaaS)    │
 └───────────────┬──────────────────────────┘     └───────────────┬────────────────────────────┘
                 │ datapolicies v4.1: protection policy            │
                 │ (sync/async), recovery points                   │
                 ▼                                                  ▼
          ┌──────────────────────────────────────────────────────────────┐
          │           PC-to-PC availability zone pairing (mutual)        │
          │           A ⇄ B replicate to each other; both replicate      │
          │           to the standby below                              │
          └───────────────────────────┬────────────────────────────────┘
                                      │ datapolicies v4.2: recovery plans
                                      │ (ordered VM boot, network mapping)
                                      ▼
                 ┌─ Google Cloud · us-west1 ── STANDBY ───────────────────┐
                 │  NC2 cluster + Prism Central C                        │
                 │   └─ Recovery points only; VMs restored at failover    │
                 │  beside it, reachable once promoted:                  │
                 │   BigQuery / Cloud SQL ─▶ Looker  (BI)                │
                 └─────────────────────────────────────────────────────────┘

 Failover:  recovery plans run on C ─▶ power on in stage priority order ─▶
            re-IP or re-attach floating IP ─▶ traffic follows ─▶ A/B unregister
 Failback:  resync, then the same plans run from A and B

Step 1: three clusters that already exist

Nothing here builds a cluster; that is the three earlier posts, one per cloud:

The precondition for everything below: three Prism Central instances, each answering on :9440, each with at least one registered cluster and a VM image already in place.

Step 2: pair the Prism Centrals (availability zones)

Cross-cluster DR in Prism Central is built on "availability zone" pairing between Prism Centrals, not between clusters directly. Each pair is a one-time trust setup, done from either side: PC A and PC B register each other as a remote availability zone, then PC B and PC C do the same, then A and C. The console does it under Administration > Availability Zones; the v4 prism namespace has the matching call, POST /api/prism/v4.1/management/domain-managers/{extId}/$actions/register, which the spec describes as registering "a domain manager (Prism Central) instance to other entities like PE and PC," with a DomainManagerRemoteClusterSpec body that takes the remote Prism Central's address and credentials. Each Prism Central's own extId comes from GET /api/prism/v4.1/config/domain-managers, and the policies and plans below name the Prism Centrals by those IDs. Three registrations for a three-way mesh:

   PC A (AWS) ───pair─── PC B (Azure)
       \                    /
        \__pair────pair___/
              PC C (GCP)

Step 3: protection policies, v4 datapolicies

A protection policy is what makes a VM's disks eligible for replication and states where recovery points land and how often. For the two active legs replicating into the standby, the policy is asynchronous (RPO measured in minutes, not synchronous mirroring — cross-cloud WAN latency rules out zero-RPO here):

POST https://<PC-A>:9440/api/datapolicies/v4.1/config/protection-policies
{
  "name": "aws-to-gcp-standby",
  "replicationLocations": [
    { "label": "PC-A-AWS", "domainManagerExtId": "<PC A extId>", "isPrimary": true },
    { "label": "PC-C-GCP", "domainManagerExtId": "<PC C extId>" }
  ],
  "replicationConfigurations": [
    { "sourceLocationLabel": "PC-A-AWS", "remoteLocationLabel": "PC-C-GCP",
      "schedule": { "recoveryPointObjectiveTimeSeconds": 900 } },
    { "sourceLocationLabel": "PC-C-GCP", "remoteLocationLabel": "PC-A-AWS",
      "schedule": { "recoveryPointObjectiveTimeSeconds": 900 } }
  ],
  "categoryIds": [ "<extId of app-tier:active-active-core>" ]
}

The same body runs against PC B with PC-B-Azure as the primary location. Two policies, one target, both async, both scoped by categoryIds to the app-tier:active-active-core category, so any VM tagged with it is covered without a per-VM edit later. The spec asks for both directions in replicationConfigurations ("Connections from both source-to-target and target-to-source should be specified"), so the C-to-A leg that failback needs is in the policy from the start.

Step 4: recovery plans, v4.2 datapolicies

The recovery plan is the failover script: which VMs power on, in what order, on what network, at the standby side. It is a resource, not a one-shot action, so it can be validated long before it is ever executed. A plan has one primary and one recovery location, so this design needs two (AWS to Google Cloud, Azure to Google Cloud), and the DR guide's advice is to create them at the primary site, PC A and PC B, from where they synchronize to the recovery side:

POST https://<PC-A>:9440/api/datapolicies/v4.2/config/recovery-plans
{
  "name": "promote-gcp-from-aws",
  "primaryLocation":  { "domainManagerExtId": "<PC A extId>" },
  "recoveryLocation": { "domainManagerExtId": "<PC C extId>" }
}

Stages and network mappings are sub-resources of the plan. A stage selects VMs by the same category and sets their power-on priority (1 first; on a planned failover the spec shuts the highest number down first):

POST /api/datapolicies/v4.2/config/recovery-plans/{recoveryPlanExtId}/stages
{
  "entityType": "VM",
  "categoryExtIds": [ "<extId of app-tier:active-active-core>" ],
  "priority": 1,
  "postActions": [ { "config": { "$objectType": "datapolicies.v4.config.DelayAction", "delaySecs": 120 } } ]
}

Network mappings go to .../network-mappings the same way, one primaryNetwork and recoveryNetwork pair per subnet. Validate before you ever need it. The action lives in dataprotection, and like the failovers it takes a job name and the direction between the two Prism Centrals:

POST /api/dataprotection/v4.2/operations/recovery-plans/{recoveryPlanExtId}/$actions/validate
{
  "name": "validate-aws-to-gcp",
  "failoverDirections": [
    { "sourceDomainManagerExtId": "<PC A extId>", "targetDomainManagerExtId": "<PC C extId>" }
  ]
}

Rehearse with .../$actions/test-failover (same body), then .../$actions/clean-up-resources, which the spec describes as deleting the "entities recovered in the last test failover action."

Step 5: the failover call, and the poll

Executing a recovery plan is one call, made on the recovery side (PC C), and like every other v4 mutation in this series it returns a task to poll, not a result:

POST https://<PC-C>:9440/api/dataprotection/v4.2/operations/recovery-plans/{recoveryPlanExtId}/$actions/planned-failover
{
  "name": "failover-aws-to-gcp",
  "failoverDirections": [
    { "sourceDomainManagerExtId": "<PC A extId>", "targetDomainManagerExtId": "<PC C extId>" }
  ]
}

planned-failover assumes the primary is reachable; in the spec's words, "VMs are powered off on the source before migrating it to the target domain manager." unplanned-failover takes the same body plus an optional recoveryReferenceTime, restores "from the recovery points on the target domain manager," and is the one you use when AWS and Azure are both actually down, not drilled. Poll /api/prism/v4.1/config/tasks/{extId} the same way the single-cloud posts did until status: SUCCEEDED. The VMs come up on PC C in stage priority order, and floating IPs re-attach on the Google Cloud side per the earlier GCP post's networking calls.

Failback reuses the same pieces. The policy already carries the C-to-A leg, and the DR guide says "the same recovery plan applies to both the failover and the failback operations," with failback run at the primary site: the same planned failover, called on PC A with the direction reversed, once the original region is confirmed healthy and resynced.

The three native services — beside the cluster, never inside the API path

None of these are Nutanix resources. That is the point: NC2's v4 API stops at the VM, the disk and the network. Everything below reads from or writes to the VMs over a VPC/VNet endpoint, on the active side only, and is entirely absent from the failover call above.

 AWS (active) ── VM writes events ──▶ S3 (raw) ──▶ Glue Catalog ──▶ Athena
                                         "the data lake": land it, catalog it, query it in place

 Azure (active) ── VM emits ──▶ Event Grid topic ──▶ Azure Function
                                         "FaaS": react to an event, no server to patch or fail over

 Google Cloud (standby, promoted) ── BigQuery / Cloud SQL ──▶ Looker
                                         "the BI layer": stays cold with the cluster, wakes with it

What this pattern does not claim

Three-way availability-zone pairing, two async protection policies and two recovery plans is the minimum shape for active/active/standby; it is not a substitute for reading your own RPO/RTO requirement against async replication's real floor (minutes, driven by cross-cloud WAN, not seconds). Nor is it a documented Nutanix configuration: the DR guide's recovery plan table lists NC2 to NC2 within the same cloud, so the cross-cloud legs are my design, to validate with Nutanix before anyone relies on them. A drill is test-failover and clean-up-resources; planned-failover moves production, and unplanned-failover on real data loss is a different operation with a different blast radius, a last resort, not a monthly exercise. And the native-service half is illustrative, not exhaustive: any AWS/Azure/Google Cloud managed service that talks to a VPC endpoint fits the same "beside the cluster, not inside the API path" rule.

Personal blog. Alex works at Nutanix; the opinions here are his own and nothing here is Nutanix confidential: every fact is public or his own field experience.

Comments

  1. Loading comments…

Comments are read by Alex before they appear. No email address needed; your name shows as you type it. See privacy.

← All posts