AI Factory Sizer — Blackwell · SuperPOD

Capacity & economics for sovereign, compliant, on-prem AI factories. Size across four Blackwell architectures — CapEx, OpEx, TCO, buy-vs-rent & flagship-model workloads, live.
DugganUSA · sizing-grade

Target

Sizing to the 18 compute trays inside a GB300 NVL72 (4× B300 each), not the whole rack. HGX nodes are already node-granular.

CapEx & OpEx

Buy vs Rent

Flagship model workload

Compare — 4-yr TCO

metric: TCO · CapEx · Racks · Power · $/GPU-hr · $/EF

Why Penguin — same GPUs, lower total cost, faster to production

Relion (Intel) and Altus (AMD) run the identical NVIDIA B300/B200 HGX silicon as a DGX — so this isn't a performance trade. The win is density, efficiency, sourcing leverage, and an integrated build/manage wrap that lowers CapEx, OpEx, and time-to-production. Numbers below track your current sizing.

2.5× rack density → less floor, less infra CapEx

Relion/Altus pack 8×B300 into 4U vs a DGX B300's 10U. Same GPUs, ~40% of the rack space — which cuts racks, PDUs, structured cabling, and data-center floor, the CapEx that never shows on the GPU line item.

Direct-to-chip liquid → lower PUE, lower OpEx

DTC liquid cooling runs a ~1.1–1.15 PUE vs ~1.4–1.5 for air. At Blackwell's 1.4 kW/GPU that overhead compounds every hour for the asset's life — power is the dominant OpEx line, and cooling efficiency is where it's won.

Dual-vendor (AMD + Intel) → sourcing leverage

Altus (EPYC) and Relion (Xeon 6) let you right-size the CPU to the workload and play vendors against each other on price and lead time — leverage a fixed single-vendor appliance doesn't give you.

No single-vendor premium · configure to the workload

Design-build-deploy-manage → faster time-to-value

Penguin ships the outcome, not a BOM: reference-architecture-validated integration, ClusterWareAI / MemoryAI / ComputeAI, and managed services. Fewer months to first token, lower ops headcount, one throat to choke.

Weeks-to-production vs a DIY science project

Buy vs Rent — own the inference load, don't rent it

Reasoning & inference are a permanent, high-utilization load — the worst case to rent. At your utilization and cloud rate, here's owned-infra TCO vs paying a hyperscaler per GPU-hour over the horizon.

Owned TCO vs Cloud rental — over 4 years

bars = total spend over the horizon (shorter is cheaper). Owned = Penguin Relion build; Cloud = same GPU-hours rented.

The strategic case for on-prem Blackwell

Three forces are moving frontier compute back inside the walls — and B300 is built for the first of them.

Reasoning is an inference load

Blackwell Ultra (B300) exists for reasoning. Reasoning models burn compute at generation time — long chains of thought, test-time scaling, agentic loops — moving the cost center from a one-off training run to a permanent inference load. That load is cheapest on iron you own. The 288 GB HBM and FP4 throughput are built to serve it.

Sovereignty

Sovereign AI is the fastest-growing driver of on-prem GPU demand: nations and enterprises require models trained and served on infrastructure they own and control, in-country. An on-prem AI factory keeps weights, data, and inference inside your walls — no hyperscaler tenancy, no cross-border data flow, no capacity queue.

Compliance by architecture

Regulated data can't live in shared cloud. Healthcare (HIPAA), finance, and government (FedRAMP High / IL5+, EU AI Act, GDPR residency) need air-gap, residency, and auditability a multi-tenant hyperscaler can't guarantee. On-prem brings the model to the data — compliance as a property of the build.

The economic corollary

All three point the same way: sustained, sensitive, owned workloads. That's precisely the profile where buy beats rent and where an integrated partner — hardware + software + managed services — turns a capital project into a running AI factory. See the Buy vs Rent tab for the number.

Rack Power, Thermals & Cabling Planner

Build up to 5 racks. Per-phase current, breaker & connector; DTC liquid/air heat split, CFM & loop flow; and a full structured-cabling BOM (DAC/AOC/optic + fiber) from the standard 8-GPU port topology, with slack for slide-out service. Every field recalculates live. Save the BOM to JSON and reload it next session. Typical/reference values — override any field with real numbers.

Mode
Hand-configure individual racks: power, thermal, per-rack cabling and a SKU-rational BOM.
saved

Project totals

Rationalized SKU roll-up — what goes on the PO

QtySparesOrderSpeedPartConnectorLengthUsed by

Structured cabling — by function

QtyLink classSpeedMedia typeLengthDescription
⚠ Port topology is per-architecture, not assumed. HGX node (Relion / Altus / DGX B300): 8× compute (1 rail/GPU) · 2× storage · 2× mgmt · 1× BMC · 1× SMC. GB300 NVL72 per compute tray: 4× ConnectX-8 compute · 2× BlueField-3 B3240 storage/N-S · 1× mgmt · 1× BMC — ×18 trays, plus rack-fixed 9× switch-tray BMC and 2× SN2201 OOB uplinks. NVLink5 GPU↔GPU is the in-rack copper cartridge backplane and is deliberately absent from the BOM. SKU roll-up rounds each computed reach up to a stock length, splits optical links into 2× transceiver + 1× trunk, and merges by physical part rather than by function. DTC capture 75–80% to liquid, 20–25% residual to air. Per-phase A = kW·1000 ÷ (V·√3·PF); breaker at 80% continuous; CFM = air-BTU/hr ÷ (1.08·ΔT); loop L/min = liquid-kW·60 ÷ (4.186·ΔT_water). Sizing-grade, not an electrical/mechanical stamp or a quote.

🐧 MemoryAI — because a starving Blackwell is the most expensive idle object you own

Explain it like I'm five 🍳
Your GPU is a very fast chef. It can cook an unbelievable number of orders per second. You paid about the price of a house for this chef.
HBM is the counter space next to the chef. It's the fastest surface in the kitchen, and there is never enough of it.
The KV cache is the chef's notes on every order still in progress. Long conversations mean long notes. A hundred customers at once means a hundred sets of notes.
Run out of counter and the chef stops cooking. Either they walk to the basement filing cabinet (that's NVMe — a hundred thousand nanoseconds) or they bin the notes and redo the work. Either way: a very expensive person, standing still.
MemoryAI is a second counter bolted to the first. A slightly longer reach (200–500 ns — you'd never see it), far more room, and much cheaper per square foot than hiring a second chef.
The whole question is whether your chef is actually standing around. If they are cooking flat out, more counter will not help and another chef is the better spend. DCGM tells you which situation you are in — and this page will say so either way.
Everything below is that same idea with the arithmetic attached. Step 1 measures whether the chef is idle. Step 2 checks whether the notes actually fit on the counter. Step 3 compares the price of more counter against the price of another chef.

A B300 costs about the same as a house and can do 15.3 petaFLOPS. When the KV cache outgrows HBM, it does nothing at all — very quickly, while drawing 1.4 kW. This tab answers one question with numbers instead of vibes: are your GPUs actually starving, and if so is CXL memory the cheapest way to feed them? Measure first (DCGM below), then do the arithmetic. If compute is your real constraint, this tab will say so plainly and point you at GPUs instead.

🔒 generated in your browser · nothing uploaded · saves to your machine

Step 1 · Measure — the 60-second DCGM test

Don't guess whether you're memory-bound. NVIDIA Data Center GPU Manager already knows. Run this against a production serving node under real load — not a benchmark — for a few minutes at 1-second resolution:

dcgmi dmon -e 1001,1002,1004,1005,1009,1010,203,204,250,252,155 -d 1000

Or scrape the same field IDs from dcgm-exporter into Prometheus if you'd rather look at a week than a minute. The field IDs are what matter — the transport doesn't.

DCGM fieldIDReading that points hereWhat it actually means
DCGM_FI_DEV_GPU_UTIL203> 85%The one to distrust on its own — it only means a kernel was resident, not that math happened. Read it alongside 1004 rather than instead of it.
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE1004< 35%The truth. Fraction of cycles the tensor pipes were actually busy. High 203 + low 1004 = the GPU is holding a seat, not working.
DCGM_FI_PROF_DRAM_ACTIVE1005< 30%Not even HBM bandwidth is saturated. You're not compute-bound or bandwidth-bound — you're waiting on something off-chip. That's the memory wall.
DCGM_FI_PROF_DRAM_ACTIVE1005> 70% w/ low 1004A different wall. Bandwidth-bound rather than capacity-bound — added capacity won't move this one. Worth resolving first, then reassessing capacity.
DCGM_FI_DEV_FB_USED / FB_TOTAL252 / 250> 95%HBM is full. You are evicting and recomputing KV, or rejecting/truncating requests. This is the capacity wall MemoryAI is built for.
DCGM_FI_PROF_PCIE_RX_BYTES1010sustained + burstyYou are already paging KV cache across PCIe. You've built a worse version of CXL by hand. Textbook case.
DCGM_FI_DEV_POWER_USAGE155< 60% of TDPUnder 840 W on a 1,400 W B300 while "busy" suggests the silicon is waiting rather than working — corroborating evidence, not a verdict on its own.
DCGM_FI_PROF_SM_OCCUPANCY1003low + low 1004Batches too small to fill the machine — usually because KV headroom caps concurrency. More memory raises the batch ceiling.
203 high · 1004 low · 1005 low · 252 pinned→ Memory-capacity bound. This is the regime MemoryAI is built for — run the numbers below.
1004 already > 60%→ Compute-bound today. GPUs are the better dollar — revisit as concurrency and context grow.
1005 > 70% with 1004 low→ HBM bandwidth-bound. Address that constraint first, then re-run this.
252 comfortably under FB_TOTAL→ Headroom to spare today. Worth rechecking as context, concurrency, or tenant count grows.

Step 2 · The arithmetic — does the KV cache even fit?

KV cache per token is 2 × layers × KV-heads × head-dim × precision-bytes. Multiply by context and concurrency and compare against the HBM you have left after weights. Everything here is editable — put your real model in.

⚠ A note if this reads like memory-overcommit guidance

Most CXL literature is written in the fleet-efficiency register, and for good reason: Microsoft's Pond paper justified pooling with ~25% of Azure DRAM stranded and 50% of VMs never touching 50% of their rented memory. That is a textbook overcommit statistic, and for that use case the analogy is exact.

It transfers poorly to KV cache, and carrying the assumption across tends to under-provision the tier. Overcommit rests on two conditions that do not hold here:

Demand is anti-correlated. Tenants peak at different times, so peak-of-sum beats sum-of-peaks. But LLM serving follows the same diurnal curve on every replica and a viral hour hits all of them at once — the pool runs dry exactly when you need it.

There are cold pages to reclaim. Ballooning works because guests hoard idle memory. A KV cache is the hot working set of in-flight requests. Nothing about it is idle. There is nothing to steal.

Size this tier to the working set, not to average utilization — which is what Step 2 does, per replica, deliberately. And note the far tier is a design point, not a failure mode: at 200–500 ns it is NUMA-far, not swap-far. You plan to sit on it permanently.

Step 2b · Multi-tenancy — who can actually separate Pepsi from Coke

If two tenants on your iron are commercial enemies, the memory math changes — and the feature that saves you the most memory becomes the one you are obligated to switch off. Set the isolation model in the controls above and Step 2 recomputes.

Separates…MechanismEnforcementReality check
CapacityCXL MLD / DCDHardware — memory controllers, up to 16 logical devices by LD-ID, each to a different hostCXL's uncontested lane. Nothing in NVIDIA's stack pools memory across mutually-untrusting hosts. But it is still mostly pilots at hyperscale as of 2026.
ComputeNVIDIA MIGHardware — crossbar ports, L2 cache banks, memory controllers and DRAM address buses assigned uniquely per instanceCommonly mislabelled "software/logical" partitioning. It is not. It has shipped in production multi-tenant clouds since A100.
TrustNVIDIA Confidential ComputingHardware — signing key fused at manufacture, never exposed to software/firmware/host; remote attestation; encrypted VRAM; NVLink encryption across up to 8 Blackwell GPUsExplicitly a don't-trust-the-operator design. ~2–5% overhead published for H100 inference.
Nothing, togetherMIG + NVLinkMutually exclusiveEnabling MIG disables NVLink P2P. You choose hard intra-GPU partitioning or the fabric — never both on the same GPU. Any comparison treating "MIG + NVLink" as one multi-tenant story is comparing something that cannot be built.
🚨 Cross-tenant prefix caching — a documented side channel, not a free optimisation

Sharing KV for common prefixes across tenants is the single biggest memory saving in multi-tenant serving — and it is a documented cross-tenant data-leak vector, not an optimisation you get for free. Cache hits return measurably faster, so an attacker probes time-to-first-token, learns which prefixes are resident, and incrementally reconstructs another tenant's prompt — system instructions, PII, medical and financial detail — with no memory access and no special privilege.

This is not theoretical: it carries a vLLM security advisory (GHSA-4qjh-9fv9-r85r) and a research literature (The Early Bird Catches the Leak, InputSnatch), with PrefixWall and SafeKV as proposed mitigations.

Set isolation to hostile in the controls and the model forfeits that saving, because in a real Pepsi-and-Coke deployment you must. The forfeited capacity is shown as its own line — that number is the price of isolation, and it is usually the honest reason to buy more memory rather than a reason not to.

Step 3 · The money — CXL vs NVMe vs just buying more GPUs

The comparison everyone makes is CXL vs NVMe. That's the wrong axis — NVMe at ~100 µs was never in the inference hot path, it's where cold data goes to retire. The comparison that decides the PO is CXL vs buying more GPUs to reach the same served throughput.

Step 3b · Time to first token — the number in the SLA

Capacity and idle-GPU capital are the internal arguments. TTFT is the one your customer actually contracts on, and it's the metric the memory tier moves most directly. It's also fully derivable rather than asserted, so here it is derived.

TTFT is set by prefill. On a cache miss you recompute the prompt's KV — roughly 2·P·tokens of FLOPs across the replica. On a cache hit you instead move that KV from wherever it lives, so the tier's bandwidth decides. Which gives the test that actually settles the tier debate: a tier slower than recomputing isn't a cache at all — it's a slower way to arrive at the same answer.

Day one or later — adding the memory tier at install time

A separate question from whether, and one that lands earlier: specify the memory tier with the cluster, or retrofit it once the workload proves it needs it? The honest answer is that install time is cheaper and later is better-informed, and which matters more depends on how well you already know your workload.

For — specify it at install

One integration, one validation, one change window. Adding a memory tier to a live cluster means a fresh qualification cycle, a maintenance window against a running SLA, and re-tuning a serving stack that people now depend on. Doing it inside the original bring-up is materially less disruptive.

Supply is the real argument. DDR5 is tight across three suppliers and CXL capacity is allocated. A tier specified in the original order is a tier you have; a tier you decide you want in nine months is a lead-time conversation.

You size power, cooling and rack space once. Retrofitting means finding kW and U in a rack that was planned without them — often the binding constraint, and the one nobody costs in advance.

Against — wait and measure

You are buying against a forecast. Every number on this page is workload-specific, and pre-production traffic estimates are routinely wrong in both directions. Ninety days of real DCGM data is worth more than any sizing exercise, including this one.

Capital sits idle if you are wrong. Unused CXL is expensive shelf-space in a market where the same spend could have been GPUs — and if you turn out to be compute-bound, that is exactly the wrong asset.

The technology is moving. CXL 3.x features and controller generations are still landing. A tier bought a year later is a better tier, at a price that may well have moved.

The practical middle — what most buyers should actually do

Buy the headroom, defer the capacity. The costs that are painful to retrofit are structural: PCIe lanes, rack U, power envelope, cooling, and the validation cycle. The cost that is easy to defer is DRAM itself. Specify a platform and a rack design that can accept the memory tier — then populate it when your own telemetry says to.

Concretely: reserve the U and the kW at design time, confirm the platform's CXL AIC slot count with your Penguin rep, and treat the modules as a follow-on order gated on 90 days of production DCGM. That converts an expensive forecasting bet into a cheap option.

Two cases override this. If your workload is already in production elsewhere and you have the telemetry, you are not forecasting — specify it day one and take the integration saving. If you are on a hard SLA from launch with long context or high concurrency, the retrofit window may never be politically available, and day one is the only realistic answer.

When the smart spend is somewhere else first

You're compute-bound

Tensor-active (1004) already above ~60% means the silicon is working and there is little idle capacity for memory to unlock. GPUs are the better dollar today. Worth rechecking as you scale concurrency — capacity is commonly the next constraint after compute. This is the most common misdiagnosis, because GPU_UTIL 203 reads ~99% in both cases.

Batch, short-context, or latency-tolerant work

Offline batch inference, classification, embeddings, short prompts, low concurrency — cold data can comfortably afford a 100 µs fetch. NVMe is the right tier here, at roughly 150× less per GB. Take the cheaper win now; the calculus changes if the workload turns interactive or agentic.

Memflation & supply concentration

CXL rides DDR5, and 2026 DRAM is in a hard shortage — DDR5 roughly tripled-to-quadrupled from mid-2025, with three suppliers and no new entrants. NAND has a broader base and trended down. Dedicated CXL supply contracts carry real allocation risk, so the case is strongest where GPU spend is your dominant cost line. A current quote is worth having before you decide either way.

…and when it clearly is

High concurrency, long context (32K+), agentic multi-step loops, prefix-cache reuse across sessions — anywhere KV working set exceeds HBM and GPUs visibly stall. CXL costs materially less per GB than HBM while sitting ~200–500× closer than SSD. Modest memory spend unlocks GPU capacity that would otherwise cost far more to add in GPUs.

Sourcing. The efficiency multipliers used as defaults in Step 3 (throughput uplift, batch-size headroom, GPU-count reduction for KV storage) are vendor- and industry-reported figures, not DugganUSA measurements — they are exposed as editable inputs precisely so you can replace them with results from your own fleet. CXL ~200–500 ns vs NVMe ~100 µs is an order-of-magnitude architectural fact; the business multipliers are workload-specific and will not reproduce identically on your models. DCGM field IDs and semantics are NVIDIA's. Confirm capacity, pricing, and CXL supply commitments with Penguin Solutions before any of this becomes a purchase order. Sizing-grade, not a quote — and never above 95% certainty.

DCGM Upload — run the numbers on your fleet, not our defaults

Every figure on the MemoryAI tab is a guess until it's your telemetry. NVIDIA Data Center GPU Manager already collects everything the decision needs. Export it, drop the file here, and the model re-runs on measured data. The file never leaves this browser — parsing is local JavaScript, there is no upload, no server, no telemetry. Close the tab and it's gone.

📄
Drop a DCGM export here
or · .txt · .prom · .json · .csv · up to ~50 MB

How to produce the file — four ways, pick whichever you already have

All four are read-only and safe to run on a production node. Capture under real serving load — a quiet node or a synthetic benchmark will tell you a comfortable lie. Aim for at least 5 minutes; an hour is better; a week through Prometheus is best.

1

Prometheus exposition from dcgm-exporter easiest · recommended

If you run the DCGM exporter (it ships with NVIDIA GPU Operator on every Kubernetes GPU cluster), it's already serving exactly what we need on port 9400. One curl, one file:

curl -s http://localhost:9400/metrics > dcgm-snapshot.prom

On Kubernetes, port-forward first: kubectl -n gpu-operator port-forward svc/nvidia-dcgm-exporter 9400:9400

Caveat: this is an instantaneous snapshot — one moment in time. Good for a quick read, weak for a purchase decision. Take several minutes apart and upload the largest, or use method 2. If your exporter doesn't publish the profiling fields, set DCGM_EXPORTER_COLLECTORS to a counters file that includes DCGM_FI_PROF_PIPE_TENSOR_ACTIVE and DCGM_FI_PROF_DRAM_ACTIVE — they are commonly commented out of the default config, which is the single most frequent reason this tab reports "tensor field missing."

2

Prometheus range query best evidence

If dcgm-exporter feeds Prometheus, ask for a real time window. This gives averages and percentiles over days rather than one instant — which is what actually justifies capital:

curl -sG 'http://prometheus:9090/api/v1/query_range' \ --data-urlencode 'query={__name__=~"DCGM_FI_PROF_PIPE_TENSOR_ACTIVE|DCGM_FI_PROF_DRAM_ACTIVE|DCGM_FI_PROF_SM_OCCUPANCY|DCGM_FI_DEV_GPU_UTIL|DCGM_FI_DEV_FB_USED|DCGM_FI_DEV_FB_TOTAL|DCGM_FI_DEV_POWER_USAGE|DCGM_FI_PROF_PCIE_RX_BYTES"}' \ --data-urlencode "start=$(date -u -d '7 days ago' +%Y-%m-%dT%H:%M:%SZ)" \ --data-urlencode "end=$(date -u +%Y-%m-%dT%H:%M:%SZ)" \ --data-urlencode 'step=5m' > dcgm-week.json

On macOS use date -u -v-7d instead of date -u -d '7 days ago'. A week at 5-minute steps is a few MB — well within what this page parses. If the response is truncated, narrow the window or widen the step.

3

dcgmi dmon — no Prometheus required bare metal

Straight from the DCGM CLI on the node. This is the one to use on a standalone box or an air-gapped site. Field IDs are explicit, so nothing depends on exporter config:

dcgmi dmon -e 1001,1002,1003,1004,1005,1009,1010,203,204,250,252,155 -d 1000 -c 300 > dcgm-dmon.txt

-d 1000 = sample every 1000 ms, -c 300 = 300 samples ≈ 5 minutes. Raise -c for a longer window (-c 3600 ≈ 1 hour). Drop -c entirely and it runs until you Ctrl-C.

Requires nv-hostengine running (sudo systemctl start nvidia-dcgm). Profiling fields 1001–1010 need DCGM 2.0+ and are unavailable inside some containers unless the DCGM socket is mounted — if those columns come back N/A, run it on the host rather than in the pod.

4

CSV from Grafana or your own pipeline fallback

Any CSV with a header row works, as long as column names contain either the DCGM field name (DCGM_FI_PROF_PIPE_TENSOR_ACTIVE) or the dmon short name (TENSO). Grafana's panel → Inspect → Data → Download CSV produces this directly. Extra columns are ignored.

Minimum useful set: tensor-active and DRAM-active. Everything else sharpens the picture but the verdict can be reached without it.

What gets read, and what it decides

DCGM fieldIDdmonUnitWhat this page does with it
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE1004TENSOratio 0–1Required. Becomes "tensor-active now" — the idle-capital calculation rests entirely on this.
DCGM_FI_PROF_DRAM_ACTIVE1005DRAMAratio 0–1Required. Separates a capacity wall (CXL helps) from a bandwidth wall (CXL does not).
DCGM_FI_DEV_FB_TOTAL250FBTTLMiBSets HBM per GPU automatically — no need to tell us what you own.
DCGM_FI_DEV_FB_USED252FBUSDMiBPeak framebuffer occupancy. >95% is the capacity wall showing itself.
DCGM_FI_DEV_GPU_UTIL203GPUTL%Only used to expose the gap against 1004 — the "busy but not working" delta.
DCGM_FI_PROF_SM_OCCUPANCY1003SMOCCratio 0–1Low occupancy + low tensor = batch ceiling, usually KV-headroom imposed.
DCGM_FI_PROF_PCIE_RX_BYTES1010PCIRXbytes/sSustained inbound = you're already paging KV over PCIe by hand.
DCGM_FI_DEV_POWER_USAGE155POWERWDraw far under TDP while "busy" corroborates idle silicon.

GPU count is inferred from distinct entity/GPU identifiers in the file. Percentiles are computed across all samples and all GPUs; the per-GPU table below the summary is where you'll spot a single starving rank dragging an otherwise healthy fleet.

🔒 Privacy. The file is read with the browser's local FileReader and parsed in page JavaScript. It is never transmitted — no fetch, no XHR, no beacon, no analytics on this tab. The page works fully offline; disconnect your network and it behaves identically. Saved reports are written to your own downloads folder by the browser. What DCGM exports contains: GPU UUIDs, model names, hostnames, and utilization — treat the file as infrastructure metadata and handle it under your own classification rules. ⚠ Parsed telemetry is real; the conclusions still ride on the editable assumptions on the MemoryAI tab. Measurement removes the guesswork about your fleet's state, not about how a specific product will perform on it. Certainty capped at 95%.

🎯 Choose X when Y — the whole decision on one page

Six gates, in order. Most of them resolve somewhere other than a memory purchase — which is the point, because the fastest way to trust a recommendation is to see what would have overturned it. Each outcome carries the condition that would bring you back. Your current numbers from the MemoryAI tab light the path you're actually on; change them there or upload real telemetry on the DCGM tab and this redraws.

Drive it by
Workload shape

Choose the isolation layer — three different walls, three different products

These are routinely compared as if they compete. They don't — each separates a different thing, and the third one is a trap that gets drawn on architecture diagrams constantly.

Separate CAPACITY
CXL — MLD / DCD
Choose when: mutually-untrusting hosts must share one memory pool without stranding it. Up to 16 logical devices by LD-ID, each mapped to a different host, enforced in the memory controller.
Only option in this lane. Nothing in NVIDIA's stack pools memory across hostile hosts.
Separate COMPUTE
NVIDIA MIG
Choose when: several tenants share one GPU and need enforced performance isolation. Crossbar ports, L2 banks, memory controllers and DRAM buses are assigned uniquely per instance — hardware, not software.
Shipping in production multi-tenant clouds since A100. Costs you NVLink — see below.
Separate TRUST
NVIDIA Confidential Computing
Choose when: you must run on infrastructure whose operator you don't trust. Signing key fused at manufacture, remote attestation, encrypted VRAM, and NVLink encryption across up to 8 Blackwell GPUs.
~2–5% published overhead for H100 inference. The layer most often assumed to be missing.
⚠ Separates NEITHER
MIG + NVLink together
Worth knowing before it reaches a design review: enabling MIG disables NVLink peer-to-peer, in both NVLink and PCIe form, and MIG-backed vGPUs are not supported for NVSwitch.
Pick hard intra-GPU partitioning or the fabric. A diagram showing both is describing a configuration the hardware will not permit.

The short version

Choose…WhenBecause
GPUs firsttensor-active (1004) > 60%Silicon is working, so little idle capacity is available for memory to unlock. Recheck as concurrency grows. Most common misdiagnosis — GPU_UTIL reads ~99% either way.
Bandwidth firstDRAM-active (1005) > 70% with tensor lowBandwidth-bound rather than capacity-bound. Resolve that, then reassess — capacity is often the next wall.
Headroom todayKV working set fits HBM headroomNothing to solve right now. Step 2 gives the concurrency at which you would cross — check against it when traffic changes.
NVMe tierbatch, short context, latency-tolerantCold data affords 100 µs and NAND is ~150× less per GB. Reopens if the work turns interactive or agentic.
MemoryAI (CXL)KV spills, GPUs stall, work is interactive~200–500 ns keeps the GPU fed where NVMe cannot, and costs far less than the GPUs it frees.
MemoryAI + isolation…and tenants are hostilePrefix dedup must be off — it's a TTFT side channel. Budget the forfeited capacity explicitly.
Pilot firstreturn under ~2×Thin returns do not yet justify allocation risk in a short DRAM market. A current quote often moves this.
⚠ The gates are ordered by how cheaply they can be settled, not by how likely they are. Two DCGM reads resolve most cases before any capacity arithmetic runs, which is deliberate — a recommendation is worth more when you can see what would have overturned it. Every outcome carries the condition that brings you back. Thresholds (60% tensor-active, 70% DRAM-active, 2× return) are editable judgement calls, not physical constants; the DCGM field semantics behind them are NVIDIA's. Outcome colour is decorative only — every node carries an icon and a written verdict, so the tree reads correctly in greyscale. Certainty capped at 95%.

📚 References — the published numbers, and who published them

Every figure this tool leans on, with its source. Vendor numbers are labelled as vendor numbers and separated from what the tool derives itself — so when you quote something in front of a customer, you know exactly which kind of claim you're making and who stands behind it.

Penguin Solutions — MemoryAI KV cache server

MemoryAI KV cache server

Penguin Solutions · announced GTC, March 2026 · vendor-published
Faster than NVMe-based KV caching10×
Faster than RDMA-based approaches3.8×
Total memory per appliance11 TB
Composition3 TB DDR5 + 8× 1 TB AIC
Form factor4U
Product page →
These are the numbers to quote. This tool's TTFT model derives a comparable CXL-vs-NVMe advantage independently, but the derived figure is bandwidth-only — cite Penguin's published 10×, not ours.

The problem it addresses

Penguin Solutions · memory-wall-scaling · vendor-published
"Performance slows when memory-starved GPUs struggle to produce tokens and become idle"
Claimed outcomesTTFT ↓ · SLA ↑
Stated fitlarge models, long context
Memory wall → Why AI needs CXL →

CXL — latency, and what the spec actually guarantees

Access latency

CXL Consortium figures · independent measurement
Local DRAM (direct path)80–140 ns
CXL-attached memory170–250 ns
Measured, incl. remote-NUMA comparison200–400 ns
Load instruction vs NUMA access~35% longer
NVMe, for contrast~100 µs
CXL latency at HC34 → Dissecting CXL Memory Performance at Scale →
The ~200–500 ns range used elsewhere in this tool is the conservative envelope across these sources. CXL sits roughly 2–3× local DRAM and roughly 400× closer than NAND — that ratio, not the absolute number, is the architectural argument.

Multi-tenancy primitives

CXL 2.0 / 3.x specification
Logical devices per MLD (by LD-ID)up to 16
Hosts served simultaneouslyup to 16
Link encryption (CXL 2.0 IDE)AES-GCM 256
Confidential VMs (CXL 3.1)TSP
Dynamic Capacity (CXL 3.0)DCD add/release
CXL 2.0 memory pooling → CXL IDE →

The stranding problem — why pooling exists at all

Pond: CXL-based memory pooling for cloud

Microsoft Research · ASPLOS '23 · peer-reviewed
Azure DRAM strandedup to 25%
VMs never touching half their rented memory50%
Achieved performancesame-NUMA-node
Estimated server cost reduction4–5%
Pond (ASPLOS '23) →
This is the fleet-efficiency case, not the AI-inference case. Useful for credibility on CXL generally; do not use it to size a KV tier — see the overcommit note on the MemoryAI tab.

Deployment status, honestly

industry reporting · 2026
Hyperscale production poolingstill largely pilots
Microsoft CXL VM preview (SAP HANA)Nov 2025
First production CXL KV cache serverMar 2026
Samsung: CXL vs DRAM, 8-GPU config92% of DRAM perf
Worth knowing before a customer raises it: a well-circulated contrarian piece argues CXL underdelivered for AI. The counter is that KV cache offload is a different use case from fleet pooling — and it is the one now shipping.

NVIDIA — the measurement and isolation layer

DCGM field semantics

NVIDIA Data Center GPU Manager
PIPE_TENSOR_ACTIVE — the truth1004
DRAM_ACTIVE — capacity vs bandwidth1005
GPU_UTIL — residency, not work203
FB_USED / FB_TOTAL252 / 250
DCGM field IDs →

GB300 NVL72 reference architecture

NVIDIA enterprise reference architectures
Compute trays (4× B300 + 2× Grace each)18
NVLink switch trays9
Power shelves6–8 × 33 kW
Rack power — max / nominal142 / 132 kW
NVL72 components →

Isolation — and the exclusivity

NVIDIA MIG · Confidential Computing
MIG instances per GPUup to 7
MIG isolationhardware memory path
Blackwell CC — NVLink encryptionup to 8 GPUs
Published CC overhead, H100 inference~2–5%
MIG enabled → NVLink P2Pdisabled
MIG user guide → Secure AI whitepaper →

Prefix-cache sharing — the security literature

Cross-tenant KV cache timing side channel

vendor advisory + peer-reviewed research
vLLM security advisoryGHSA-4qjh-9fv9-r85r
Attack vectorTTFT timing probe
Proposed mitigationsPrefixWall · SafeKV
vLLM advisory → The Early Bird Catches the Leak → InputSnatch →
Raise this before a security team does. It converts a memory-saving feature into a capacity requirement in any multi-tenant deployment — which is an argument for the tier, quantified on the MemoryAI tab.

What this tool derives itself

DugganUSA · computed, not cited
KV per token2·L·H·D·bytes
Prefill / recompute cost≈2·P·tokens FLOPs
Tier fetch timeKV bytes ÷ bandwidth
Partial NVL72 powerfixed + per-tray
Cable BOMper-arch port topology
These are arithmetic from first principles on inputs you control — reproducible, and wrong in exactly the ways your inputs are wrong. They are not measurements of any product.
Three kinds of claim, kept separate on purpose. Vendor-published figures (Penguin's 10× and 3.8×, NVIDIA's specifications) belong to those vendors — quote them with attribution and they carry that vendor's weight. Peer-reviewed and independent measurement (Pond, the CXL latency work, the side-channel literature) is the strongest material available and is what to lead with in a technical audience. Derived figures are this tool's own arithmetic on your inputs — useful for showing a customer their own situation, never for claiming a product benchmark. Presenting a derived number as a measured one is the failure mode this page is structured to prevent. Certainty capped at 95%.

📋 Build Sheet — rows, racks and the tag taxonomy

The row-by-row breakdown an installer works from, and the asset tagging schema that has to exist before the first rack lands. Derived from whatever the Rack Planner's Advanced mode is currently set to — change the target there and this follows.

Row and rack breakdown

RowRacksIdentifiersContentsGPUsLoadNotes

Per-rack manifest

RackRowSUTypeDevicesPortsLoad

Asset tagging — sovereign adaptation

Microsoft's Cloud Adoption Framework and AWS's tagging guidance both assume a hyperscaler owns the substrate. The provider asserts jurisdiction, vets the staff, controls physical custody and attests to it. A sovereign on-prem AI factory inverts that: you are the substrate, so the guarantees a cloud provider would have made on your behalf have to become properties you carry yourself. That is the whole adaptation — the four commercial categories below are lifted straight from CAF and AWS, and the fifth is the one their frameworks have no reason to include.

🔒 generated in your browser · nothing uploaded
⚠ Row assignment is a naive fill at your stated racks-per-row and takes no account of your actual floor plan, hot/cold aisle orientation, structural loading or CRAC placement. A GB300 NVL72 rack is roughly 1,500 kg populated, which is a structural question in many buildings before it is a networking one. Tag keys are recommendations; tag values and any enforcement policy are yours. Nothing here is a compliance attestation — a tag records an intent, it does not prove one. 95% cap.
AI Factory Assistant
Blackwell · Penguin Solutions · powered by Mistral
Ask about Blackwell / SuperPOD sizing, Penguin Solutions Relion & Altus, TCO & buy-vs-rent, or reasoning / sovereignty / compliance.