Sign in →
CloudLab Works emblem: Waku the orca ringed by CloudLab and WorksCLOUDLAB WORKScrossed whale bones, one end a wrenchCloud City, headwaters of the agentic cloud revolution!

Nutanix: Invisible Wires: Agentic Cloud

Where should the model run? A placement decision for a regulated RAG assistant across NCI, NC2 and the hyperscalers

· By Alex Alvord

Every AI conversation I have with a regulated customer starts in the same place. Not "which model", and not "how many GPUs". It starts with "where is this allowed to run, and what happens when we have to leave?"

This is one worked answer to that question, written the way I would hand it to an architecture review board: as a decision record. Seven decisions, each with what it rejected, what it costs the architecture, and the fact that would change it. It builds on three posts I have already published: the active/active/standby pattern across three clouds, resilience as a regulated number, and IAM Roles Anywhere on NC2.

The customer and the brief

An insurer with entities in the EU and the UK, so DORA and the PRA's operational resilience rules both apply. It already runs:

The brief: a retrieval-augmented (RAG) assistant for about 2,000 internal users, peaking at 64 requests in flight, with answers grounded in those documents. Prompts will carry regulated data. Nothing leaves the EU.

The demand, as planning assumptions: 1,536 tokens of retrieved context in, 512 tokens out, a 2-second time-to-first-token target at p95, and at least 10 tokens per second per user. At peak that is 2,560 tokens per second (derived: 64 × 2,048 tokens ÷ 51.2 s per answer).

The map

Where the model runs: the decision record on one page
Where the model runs: the decision record on one page
The same map as a flow, decisions D1 to D7
The same map as a flow, decisions D1 to D7

The decision record

#DecisionRejectedWhat it costs the architecture
D1Primary inference on Nutanix Enterprise AI (NAI) on EKS, 2× g6e.4xlarge (L40S), FrankfurtNC2 g4dn.metal; on-prem GPU now; managed APIA second platform (EKS) beside NC2; the NC2 GPU is off NAI's supported list
D2Managed API (Bedrock) for non-regulated work onlyManaged API for everythingTwo inference paths to govern and audit
D3Open-weights model: Llama 3.1 8B InstructA proprietary managed model; 70B nowWe own upgrades and evaluation; 70B needs 4× the GPUs
D4Corpus stays on NCI; only the derived index travelsCopying the corpus to the cloudAn index pipeline to keep in sync; NAI's storage floor is 3.15 TiB
D5Cold standby on AKS in an EU Azure region, exit tested quarterlySingle cloud; hot standbyA quarterly drill to staff; GPU capacity not guaranteed on exit day
D6IAM Roles Anywhere for anything outside AWS calling AWSLong-term access keysA certificate authority to operate; zero long-lived secrets
D7One control plane: NAI on EKS, AKS and, later, NKP on NCIA native ML platform per cloudOne procedure to fail over, not two projects

Each one in turn.

D1 · Primary inference runs on NAI on EKS, beside the NC2 estate

Decision. Two g6e.4xlarge instances in Frankfurt, one NVIDIA L40S each, under NAI on EKS. One carries the load, the second is N+1. The NC2 application VMs keep native AWS IP addresses, so they reach the inference endpoint by plain VPC routing.

Why the GPU decides it. The only GPU instance NC2 on AWS supports is g4dn.metal, with eight NVIDIA T4s. NAI's supported-GPU list includes the L40S, H100, A100 and others. The T4 is not on it. Unlisted GPUs are not blocked, but they are "not officially supported". On EKS, NAI supports the G6e family, which is the L40S.

What Option A costs the architecture. Three g4dn.metal nodes put 24 T4 GPUs on a platform NAI does not officially support. Every NAI upgrade becomes a compatibility question nobody else has answered, and every support case starts with "unsupported GPU". That is tech debt taken on at day one, sized to the cluster rather than the workload: three bare-metal nodes for a load that fits on one L40S (D3). Option B adds a second platform to operate, EKS beside NC2, but it is the platform NAI is validated on, and it grows one GPU at a time rather than one bare-metal node at a time.

The storage fact that also rules out Option A. NAI's minimum storage is 150 GiB block, 2 TiB of shared file storage and 1 TiB of S3-compatible object storage: 3.15 TiB in total. Three g4dn.metal nodes give about 2.7 TB usable at RF2. It does not fit without adding storage elsewhere.

Rejected: on-prem GPUs now. The most residency-friendly answer, and the right long-term home if volume grows. It means buying and racking GPU nodes before demand is proven, with a lead time I could not confirm. It becomes the real option at the trigger below.

What would change it. GPU utilization above 60% sustained for a quarter: then size L40S nodes for the on-prem NCI estate. Or NAI validating a GPU on NC2, which would make Option A a one-platform answer.

D2 · The managed API is allowed, but not for regulated prompts

Decision. Amazon Bedrock stays available for work that carries no regulated data: drafting, public content, development and test.

The trade-off, stated honestly. At this volume the managed API is the cheaper thing to run, and I would tell the customer that before anything else. Its cost sits on the other side of the ledger. Prompts, evaluations and integrations get built against one provider's API and its model versions, and that is debt you pay down on exit day. The self-hosted path is justified by residency, control and the exit plan in D5, not by run cost.

What would change it. A clear legal position that a managed model in an EU region satisfies the insurer's data-processing obligations, plus confirmation that the model is offered in-region. I could not confirm Llama 3.1 8B availability or pricing in Frankfurt for this note.

D3 · An open-weights model, because the exit plan needs one

Decision. Llama 3.1 8B Instruct. The weights are 16.06 GB at FP16. The KV cache is 128 KiB per token, so 64 sequences of 2,048 tokens need 16 GiB. That fits one 48 GB L40S with headroom. FP8 on the L40S halves the weights again, after an accuracy regression test.

Why open weights. DORA Article 28(8) requires exit strategies for ICT services that support critical or important functions, and says firms must be able to exit "without disruption" to their business. An open-weights model can be copied to any cluster that runs NAI. A proprietary managed model has no copy to move.

Rejected: 70B now. NAI needs at least four L40S GPUs for Llama 3.1 70B. With N+1 that is two g6e.12xlarge in Frankfurt: eight L40S GPUs instead of two. That is four times the footprint to patch, schedule and fail over, for a gain nobody has measured yet.

What would change it. An evaluation on the insurer's own documents showing 8B misses the accuracy bar.

D4 · The corpus stays on NCI; only the derived index travels

Decision. The documents stay where they live today, on Nutanix Files and Objects in the insurer's own data centre. Embeddings are computed from them, and only that index is copied to the inference sites: encrypted, held in the EU, under the same retention as the source.

The trade-off. The index is derived from regulated documents, so it is regulated too. Copying it is a real decision, not a technicality. The alternative is retrieving over the WAN from on-prem at query time, which makes the on-prem estate a dependency of every answer and of every failover. That is the wrong way round for D5.

What would change it. A legal view that the index must not leave the data centre. Then retrieval stays on-prem and the inference tier moves on-prem with it, which pulls the D1 trigger forward.

D5 · A cold standby on a second cloud, and an exit you actually run

Decision. NAI on AKS in an EU Azure region, kept at zero GPU nodes. Same NAI release, same model files, same endpoint name. The index copy runs nightly, so retrieval can be up to 24 hours stale after a failover. Prompts are not stored, so there is nothing else to replicate.

Why a second cloud. An NC2 cluster cannot span availability zones, and the application tier lives in one. The resilience post makes the regulatory case: supervisors now expect firms to show they can lose a provider. So the AI tier gets its own exit path, built on the same idea as the active/active/standby pattern.

The exit, as a procedure:

The exit path, run on a schedule
The exit path, run on a schedule
The exit path as a flow, with the one real risk marked
The exit path as a flow, with the one real risk marked

The trade-off. A cold standby sits idle until it is used, so its cost is operational: a drill to run every quarter and a runbook to keep current. The real risk is step 2: GPU capacity in the standby region is not guaranteed on the day you need it. Reserving capacity removes that risk, in exchange for capacity that sits unused. I have left the standby GPU type open, because I did not verify which NAI-supported GPUs AKS offers in the target region.

What would change it. A first quarterly test that misses the impact tolerance. Then reserve the capacity, or move to warm standby.

D6 · No long-lived keys anywhere

Decision. Anything outside AWS that calls AWS uses IAM Roles Anywhere. That covers the NC2 VMs reaching Bedrock for non-regulated work, and on-prem jobs writing to the S3 model store. Workloads present an X.509 certificate from the insurer's own certificate authority and receive temporary credentials. No long-term access keys live on any cluster.

The trade-off. It needs a certificate authority and a certificate lifecycle that someone owns. The IAM Roles Anywhere post walks through the setup.

D7 · One control plane on every site

Decision. NAI runs the model everywhere: EKS now, AKS on standby, and NKP on the on-prem NCI estate when D1's trigger fires. That gives one model catalog, one endpoint shape and one runbook.

Rejected: a native ML platform per cloud. Two clouds run with two toolchains are not resilience. They are two single points of failure. With one control plane, the D5 exit is a procedure. With two, it is a migration project.

What would change it. Nutanix has announced it acquired Ryax, an AI workload orchestration company, and said the technology targets future NKP and NAI releases, for smarter scheduling and GPU placement. Nothing in this record depends on that. When it ships, placement across the three sites could become a policy rather than a manual decision.

What this record does not claim

The short version

At this volume the managed API is the cheaper thing to run, and I would say so first. The customer still self-hosts, for three reasons. The data has to stay in its control. The exit plan needs a model that can move. And one platform across both clouds and its own data centre turns a failover into a procedure that gets run every quarter. The run-cost argument says "rent". The regulation argument says "be able to leave". When they disagree, a regulated firm should follow the second, and write down the first so that nobody pretends it isn't true.

Sources: Nutanix, NC2 on AWS: Supported Regions and Bare-metal Instances and Installing NVIDIA GRID Host Driver (g4dn.metal is the only supported GPU instance; T4); Nutanix, Nutanix Enterprise AI 2.8 Requirements (supported Kubernetes platforms including EKS and AKS, supported GPUs, unlisted-GPU policy, storage minimums, 70B GPU minimum); Nutanix Bible, NC2 on AWS (single-AZ clusters, native networking); AWS public EC2 offer file, EU (Frankfurt), read 5 October 2026, for instance specifications only (GPU type and count for g4dn.metal, g6e.4xlarge, g6e.12xlarge); NVIDIA L40S and T4 product pages; Hugging Face, Llama 3.1 8B Instruct model card and config (parameter count, layers, KV heads); Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention (arXiv:2309.06180); AWS, What is IAM Roles Anywhere?; Regulation (EU) 2022/2554 (DORA), Article 28(8); Nutanix, Bringing Intelligent AI Workload Orchestration to NKP and Nutanix Enterprise AI (Ryax acquisition).

Personal blog. Alex works at Nutanix; the opinions here are his own and nothing here is Nutanix confidential: every fact is public or his own field experience.

Comments

  1. Loading comments…

Comments are read by Alex before they appear. No email address needed; your name shows as you type it. See privacy.

← All posts