Sign in →
CloudLab Works emblem: Waku the orca ringed by CloudLab and WorksCLOUDLAB WORKScrossed whale bones, one end a wrenchCloud City, headwaters of the agentic cloud revolution!

Nutanix: Invisible Wires: Agentic Cloud

The cheap seat: edge-class inference on NC2 on AWS with the NVIDIA T4

· By Alex Alvord

In Where should the model run? I rejected NC2 on AWS's GPU option for a regulated RAG assistant. That was the right call for that workload: one 8B chat model serving a whole company, and a platform, Nutanix Enterprise AI, that doesn't list the T4 as supported.

The same GPU is the right answer for a different workload, and it's a common one. This post is about that workload.

The workload: edge-class inference

By edge-class I mean inference that is small, fast and close to the thing that calls it. A vision model reading a camera feed. Speech-to-text on a support line. Embeddings and a reranker for search. OCR on scanned forms. A small 8B model doing classification or tool calls, not long conversations.

Each one is light on its own. There are usually a lot of them. And each wants to sit next to the application it serves, not across a network boundary on a separate platform.

The hardware you already have

The only GPU instance NC2 on AWS supports is the g4dn.metal: eight NVIDIA T4 GPUs in a bare-metal node, with 96 vCPUs, local NVMe and 100 Gbps networking. AWS positions G4dn as its lowest-cost GPU instances for machine-learning inference.

The T4 is built for exactly this class of work. NVIDIA's numbers: 16 GB of memory, 70 watts, tensor cores for FP16, INT8 and INT4, and hardware video decode for up to 38 full-HD streams. It's a 2018 design from the Turing generation, and that matters later.

If your NC2 cluster runs on g4dn.metal nodes for its VMs anyway, the GPUs are already there.

One NC2 node, two jobs: the application VMs and eight single-GPU model endpoints
One NC2 node, two jobs: the application VMs and eight single-GPU model endpoints

Five design rules

  1. One model per GPU, never split. The T4 sits on PCIe Gen3 with no NVLink between cards. Sharding one model across several T4s spends most of its time waiting on the bus. Size every model to fit one card.
  2. Quantize. The T4's tensor cores do INT8 and INT4. An 8B model is about 16 GB at FP16, which fills the card before any context. At INT8 it's about 8 GB, and at INT4 about 4 to 5 GB, leaving room for the KV cache.
  3. Keep contexts short. The T4 has no BF16 and no FP8, and FlashAttention-2 needs Ampere or newer; Turing is covered only by a separate project with a subset of the features. Long-context chat runs poorly here. Classification, extraction, tool calls and short answers run fine.
  4. Run the serving stack yourself. Because Nutanix Enterprise AI doesn't officially support the T4, serve the models with an open-source stack in VMs on the same cluster, or on Kubernetes if you run it there. vLLM supports compute capability 7.5 and up, and names the T4 as its floor. llama.cpp and NVIDIA's Triton are the other usual choices.
  5. Keep a spare card, not a spare node. One idle T4 per node is the N+1 for every endpoint on it. That's far cheaper for the architecture than holding a whole node in reserve.
What fits in 16 GB on one T4
What fits in 16 GB on one T4

Where the low cost comes from

None of it is a price list. It comes from four architecture choices.

What it costs the architecture

You take onWhy it's acceptable here
You own the serving stack: images, upgrades, monitoringThe models are small and few in kind. One serving image covers most of them
No official NAI support on the T4NAI isn't in the path. If the workload outgrows the T4, it moves to the platform in the decision record
A 2018 GPU: no BF16, no FP8, no FlashAttention-2Small models at short contexts don't need them. Plan the move to the next GPU instance type when NC2 on AWS qualifies one
GPUs scale with nodes: more GPUs means more bare metalEndpoints are sized to one card, so you add a node only when the VMs or the endpoints run out of room
A 16 GB ceiling per modelThat's the design: anything bigger belongs elsewhere

When to say no

The decision record's workload is the counter-example: one model serving 2,000 people at 2,560 tokens a second with retrieval context. That needs bigger cards and a supported inference platform, so the T4 is the wrong seat. Use the T4 when the models are small, many and close to the app. Use something else when one model is large, central and shared.

The architecture: one stack, a cloud end and an edge end

The serving images, quantized model files and endpoint layout built here don't care where the GPU lives. The same stack runs on an inference card in an on-prem edge cluster under AHV. NC2 on AWS becomes the cloud end of that estate: test a model there, ship the same files to the sites, and burst back to the cloud when a site is short.

The estate: NC2 on AWS as the cloud end, one set of versioned artifacts, and on-prem AHV sites as the edge end
The estate: NC2 on AWS as the cloud end, one set of versioned artifacts, and on-prem AHV sites as the edge end

Three layers, and only the middle one is new. The cloud end and the edge sites are AHV clusters you already know how to run. The artifacts (one serving image, the quantized model files, and a record of which model sits on which card) are versioned in one registry and published once. A site that runs short of GPU calls the cloud endpoint over a private link to the VPC (VPN or Direct Connect); the stack itself doesn't change to make that work.

What this does not claim

Sources: Nutanix, NC2 on AWS: Supported Regions and Bare-metal Instances and Installing NVIDIA GRID Host Driver (g4dn.metal is the only supported GPU instance; T4); Nutanix, Nutanix Enterprise AI 2.8 Requirements (supported GPUs; unlisted GPUs not officially supported); AWS, Amazon EC2 G4 Instances (up to 8 T4s, 96 vCPUs, 100 Gbps, local NVMe, bare metal; positioned for low-cost inference); NVIDIA, T4 Tensor Core GPU (16 GB GDDR6, 70 W, FP16 / INT8 / INT4, PCIe Gen3 x16, video decode); vLLM, GPU installation (compute capability 7.5 or higher, T4 named); Dao-AILab, flash-attention (FlashAttention-2 requires Ampere, Ada or Hopper; Turing in a separate subset project; BF16 needs Ampere or newer).

Personal blog. Alex works at Nutanix; the opinions here are his own and nothing here is Nutanix confidential: every fact is public or his own field experience.

Comments

  1. Loading comments…

Comments are read by Alex before they appear. No email address needed; your name shows as you type it. See privacy.

← All posts