Private cloud hardware decisions get made in the wrong order. A vendor configuration lands, the team argues about CPU models for three weeks, and nobody measures the workload the cluster is supposed to run.
The order that works is the reverse. Profile the workload, derive the constraint, then select hardware against the constraint. Almost every private cloud that ends up either starved or half empty got there by skipping the first step. Here is how to run the sequence, and where the money leaks.
Profile the workload before you open a configurator
You need four numbers per workload class, measured rather than assumed.
Steady-state CPU, at p95 over at least 30 days. Use the measurement, not the vCPU count allocated to the guest. Allocated vCPU is a budget somebody set once, usually generously, and sizing to it is how a cluster ends up twice the size it needs to be.
Working set, not provisioned memory. Memory is usually the first resource to run out in a virtualisation cluster, which is why it deserves a measurement rather than a rule of thumb. Overcommit is also platform-dependent: VMware supports memory overcommit, and Nutanix AHV does not allow it at all. That single difference can change your node count.
IOPS and the read/write split, at peak. Two workloads with the same average throughput and different write ratios need different drives.
East-west versus north-south traffic. A three-tier application with a chatty database layer generates far more traffic between nodes than out of the cluster. That determines the fabric, not your internet link.
Then apply headroom deliberately, and write down the reason for each allocation. N+1 for a node failure. A defined percentage for growth over the refresh window. Undocumented headroom compounds silently at every layer.
Compute: cores, clock, and an honest overcommit ratio
CPU overcommit is where hardware budgets are won or lost, and there is no universal ratio. VMware’s own architecture guidance is explicit that the achievable vCPU to physical core ratio depends on workload and hardware rather than a fixed number, and Nutanix takes the same position for AHV. The ratio has to come from your data. Practically, that means three sizing postures in one private cloud cluster:
- General-purpose virtual machines, web tiers and test environments tolerate meaningful overcommit and should be sized to measured p95.
- Latency-sensitive workloads such as transactional databases and real-time inference should be sized at or near 1:1 and pinned. Queuing here is a design failure rather than an efficiency gain.
- Licensed-by-core workloads should be sized to minimise core count, because the licence is usually more expensive than the silicon.
On core count versus clock speed, the deciding factor is whether the workload scales horizontally. Consolidation density favours high core counts. Single-threaded application performance favours fewer, faster cores. Buying a 64-core part to run software that cannot use more than eight threads per instance is a common and expensive error, particularly where that software is licensed per core.
For GPU nodes, the GPU is not the whole specification. Memory bandwidth, NVLink interconnect and storage throughput all shape job completion time. A GPU node with an undersized storage path is a GPU node that idles.
Storage: endurance and the replication tax
Two decisions dominate storage cost, and both have exact numbers attached.
The first is the data protection scheme. In Ceph, the default 3x replication carries a 3.0x overhead factor, which is 33.3% space efficiency. Erasure coding does better: a k=4, m=2 profile runs 1.5x overhead at 66.7% efficiency, and k=8, m=3 runs 1.375x at 72.7%. On the same raw disks that is two to three times the usable capacity.
The catch is real. Erasure coding uses more CPU and RAM on access and recovery, and each write has to land across k+m devices, which raises write latency. Red Hat’s guidance is to prefer replication for workloads with frequent partial writes, such as databases and block images, and where low write latency is critical. There is also a floor on cluster size: a k=4, m=2 profile needs at least six OSD hosts if your failure domain is the host. The practical split is to erasure code the object and archive tiers and replicate the block tier.
The second decision is endurance. Read-intensive enterprise SSDs, suitable for boot volumes and read-heavy analytics, are rated around 1 to 3 DWPD. Write-intensive parts, aimed at write-ahead logs and transactional databases, run 3 to 25 DWPD or higher. Validate the rating against your actual write pattern rather than the workload label, because small random writes, frequent snapshots and high metadata churn wear NAND faster than large sequential writes at the same total volume.
Network: sized for the fabric, not the uplink
Underspecifying the network is the most common structural mistake in private cloud builds, because the symptom presents as a storage problem. Distributed storage is network storage. Every write in a replicated or erasure-coded pool crosses the fabric, and every rebuild after a failed drive saturates it. If storage replication and virtual machine traffic share the same links, a rebuild during business hours becomes a production incident.
Three things to get right. Separate the storage replication network from the workload network, physically or with enforced bandwidth guarantees. Set port speed from east-west traffic and rebuild time, not from your internet uplink. And build redundancy at the fabric, since a top-of-rack switch is a single point of failure for every node in that rack.
Power and density: the constraint that is not on the spec sheet
Capacity is increasingly limited by kilowatts per rack rather than by rack units.
Uptime Institute’s 2025 global survey put the mean rack density at 7.6 kW, up from 6.8 kW in 2024, with 82% of respondents reporting their highest-density rack below 30 kW and only 9% at 50 kW or above. Current-generation AI hardware sits well outside that range: an AI rack draws between 40 kW and 142 kW depending on platform, with 72-GPU configurations rated at 130 kW and half-sized versions at 60 kW to 70 kW.
Confirm per rack power and cooling limits with your facility or private cloud provider before you select nodes, not after. A facility that cannot deliver 40 kW to a cabinet will force you to spread eight dense nodes across four racks, which changes your cabling, your fabric and your colocation bill.
Where teams overspend
Sizing to allocated resources rather than measured ones. Covered above, and worth repeating because it is the largest single cause of overbuild.
Compounding headroom. Twenty per cent for growth, plus N+1, plus a vendor’s own recommended buffer, plus a safety margin at the storage layer, is not 20% headroom. It is closer to double the cluster.
Licensing bought after the hardware. Broadcom’s April 2025 change set a 72-core minimum purchase per order, per VMware product, replacing the previous 16-core floor. The Register reported, from a distributor notice to partners, that a customer running a single eight-core processor would pay for 72 cores, meaning 64 cores of software they cannot use. Price the licences against the exact node configuration before you sign for the hardware, because a three-node cluster below the minimum pays for cores it does not have.
Uniform tiers. One drive class, one node type and one protection scheme across the whole estate means you are paying write-intensive prices for archive data and replication overhead for objects that would erasure code cleanly.
The pattern behind all four is the same. Each is a decision made once, applied everywhere, and never checked against a measurement. Measure first, tier deliberately, and price the whole stack including licences before the purchase order.
Related reading: Choosing an AI infrastructure partner: what to look for and Migrating AI workloads to cloud in Australia.
Frequently asked questions
What vCPU to physical core ratio should I use for a private cloud? There is no universal figure. VMware’s architecture guidance states the achievable ratio depends on workload and hardware, and Nutanix gives the same answer for AHV. Derive it from measured p95 CPU utilisation over at least 30 days, and split your estate: general-purpose virtual machines tolerate meaningful overcommit, latency-sensitive databases and inference should sit at or near 1:1, and core-licensed workloads should be sized to minimise core count.
Should I use erasure coding or replication for private cloud storage? Use both, on different tiers. Ceph 3x replication gives 33.3% space efficiency but low write latency, which suits block storage, databases and virtual machine images. Erasure coding gives 66.7% at k=4, m=2 and 72.7% at k=8, m=3, at the cost of extra CPU, RAM and write latency, which suits object and archive tiers. Erasure coding also needs a minimum cluster size: k=4, m=2 requires at least six OSD hosts with a host-level failure domain.
How much SSD endurance do I actually need? Match DWPD to your measured write pattern rather than to a workload label. Read-intensive enterprise drives are rated at roughly 1 to 3 DWPD and suit boot volumes and read-heavy analytics. Write-intensive drives run 3 to 25 DWPD or higher for write-ahead logs and transactional databases. Small random writes, frequent snapshots and metadata churn consume endurance faster than the same volume of sequential writes.
What is the most commonly missed cost in a private cloud build? Licensing priced after the hardware is chosen. Broadcom’s April 2025 change imposed a 72-core minimum purchase per order per VMware product, up from a 16-core floor, so a node running a single eight-core processor is licensed for 72 cores and pays for 64 it cannot use. Price licences against the exact node configuration before committing to hardware, and check per rack power limits at the same time, since density rather than floor space is now the usual physical constraint.