Private AI Cluster
Multiple GPU nodes, shared storage and a high-speed fabric. Priced per node - plan for three or more.
Roughly ~160B params in fp16, ~320B in 8-bit, or ~590B in 4-bit quantization.
- GPU8× NVIDIA RTX 6000 Adamemory: 48 GB GDDR6 ECC · bandwidth: 960 GB/s · interface: PCIe 4.0 x16$54,400Check price ↗
- CPUAMD EPYC 9354cores: 32 · threads: 64 · clock: 3.8 GHz boost$3,200Check price ↗
- MotherboardSupermicro H13SSL-N (SP5, EPYC 9004)chipset: SoC · memory: 12 x DDR5 RDIMM ECC, up to 3 TB · pcie: 8 x PCIe 5.0 x16 (slots + MCIO fan-out)$850Check price ↗
- RAMSamsung 512 GB (8x64) DDR5-4800 ECC RDIMMcapacity: 512 GB (8 x 64) · type: DDR5-4800 RDIMM · ecc: ECC$2,550Check price ↗
- StorageKioxia CD8-R 3.84 TB U.2 NVMecapacity: 3.84 TB · bus: PCIe 4.0 U.2 · read: 7000 MB/s$620Check price ↗Sabrent Rocket 4 Plus 8 TB NVMecapacity: 8 TB · bus: PCIe 4.0 NVMe · read: 7100 MB/s$700Check price ↗
- PSUFSP CRPS 2+2 redundant, 4000 W usablecapacity: 4000 W usable (2+2 hot-swap) · rating: 80+ Platinum · standard: CRPS redundant$1,400Check price ↗
- CoolingGeneric 4U high-static-pressure fan wall + shroudstype: rack forced-air · dissipation: chassis-wide · noise: loud (data-center)$220Check price ↗
- NetworkingNVIDIA ConnectX-6 Dx 100GbEspeed: 100 GbE · ports: 2 (QSFP56) · bus: PCIe 4.0 x16$650Check price ↗MikroTik CRS309 8x 10GbE SFP+ switchports: 8 x SFP+ 10 GbE + 1 GbE · throughput: non-blocking · form factor: desktop / 1U$280Check price ↗
- UPSEaton 9PX 3000 RT (online double-conversion, 2U)capacity: 3000 VA / 2700 W · topology: online double-conversion · runtime: ~7 min at full load$1,600Check price ↗
- ChassisSupermicro 4U 8-GPU server chassis (CRPS, fan wall)form factor: 4U rack · gpu slots: 8 dual-slot passive · power: 2+2 CRPS redundant$3,200Check price ↗
Component prices are hand-sourced and dated (checked 2026-09-16), not scraped. Confirm on the vendor's site before buying.
- · One node = 8 RTX 6000 Ada, 384 GB of VRAM. The numbers shown here are for a single node; a cluster is three or more of them.
- · The 100 GbE fabric (RoCEv2) carries tensor- and pipeline-parallel traffic between nodes and NVMe-over-fabric to shared storage - it is the part that makes it a cluster rather than a room of servers.
- · Shared storage lives on a separate node (not priced here): every GPU node mounts the same dataset and model store.
- · At this scale, colocation usually beats on-prem: metered power, cooling, and remote hands cost less than building it yourself.
- · Cluster management (Slurm or Kubernetes + a scheduler) is assumed - the hardware is the easy half.
- + Scales horizontally - add nodes, not rebuilds
- + 384 GB VRAM per node, shared model store
- + Fabric ready for multi-node training
- − Per-node price only - shared storage, fabric switch and colocation are extra
- − Needs an ops team: scheduling, monitoring, failover
- − PCIe GPUs, not SXM - frontier-scale pretraining still wants an H100/B200 cluster
A6000 nodes: same 384 GB per node at roughly 60 % of the GPU cost.
L40S nodes trade some throughput for datacenter-standard parts and support.
For frontier training, rent an H100/B200 pod - see GPU cloud pricing.
Over 36 months, at this build's power draw and the live cloud rate for its GPU.
- Hardware
- $69,670
- Electricity
- $6,606
- Total cost of ownership
- $76,276
- Equivalent cloud, monthly
- $2,698/mo
- Equivalent cloud, total
- $97,131
- Break-even
- 28 months
- · High-availability needs usually keep some cloud capacity in reserve even when owning makes sense - hybrid covers both.
- · Over 36 months, owning saves roughly $20855 versus the cloud equivalent.
"Buy on Amazon" links are affiliate links - as an Amazon Associate, Obolith earns from qualifying purchases.