<- Reference builds
Cluster node

Private AI Cluster

Multiple GPU nodes, shared storage and a high-speed fabric. Priced per node - plan for three or more.

aggregate throughputshared storagehorizontal growth
updated 16 Sept 2026
Total price
$69,670
VRAM
384 GB
RAM
512 GB
Storage
12 TB
Est. power draw
3047 W
AI capability

Roughly ~160B params in fp16, ~320B in 8-bit, or ~590B in 4-bit quantization.

small-model inferencemid-size inferencelocal RAGlarge-model inferenceLoRA fine-tuningconcurrent servingtraining
Components
  • GPU
    8× NVIDIA RTX 6000 Adamemory: 48 GB GDDR6 ECC · bandwidth: 960 GB/s · interface: PCIe 4.0 x16
  • CPU
    AMD EPYC 9354cores: 32 · threads: 64 · clock: 3.8 GHz boost
  • Motherboard
    Supermicro H13SSL-N (SP5, EPYC 9004)chipset: SoC · memory: 12 x DDR5 RDIMM ECC, up to 3 TB · pcie: 8 x PCIe 5.0 x16 (slots + MCIO fan-out)
  • RAM
    Samsung 512 GB (8x64) DDR5-4800 ECC RDIMMcapacity: 512 GB (8 x 64) · type: DDR5-4800 RDIMM · ecc: ECC
  • Storage
    Kioxia CD8-R 3.84 TB U.2 NVMecapacity: 3.84 TB · bus: PCIe 4.0 U.2 · read: 7000 MB/s
    Sabrent Rocket 4 Plus 8 TB NVMecapacity: 8 TB · bus: PCIe 4.0 NVMe · read: 7100 MB/s
  • PSU
    FSP CRPS 2+2 redundant, 4000 W usablecapacity: 4000 W usable (2+2 hot-swap) · rating: 80+ Platinum · standard: CRPS redundant
  • Cooling
    Generic 4U high-static-pressure fan wall + shroudstype: rack forced-air · dissipation: chassis-wide · noise: loud (data-center)
  • Networking
    NVIDIA ConnectX-6 Dx 100GbEspeed: 100 GbE · ports: 2 (QSFP56) · bus: PCIe 4.0 x16
    MikroTik CRS309 8x 10GbE SFP+ switchports: 8 x SFP+ 10 GbE + 1 GbE · throughput: non-blocking · form factor: desktop / 1U
  • UPS
    Eaton 9PX 3000 RT (online double-conversion, 2U)capacity: 3000 VA / 2700 W · topology: online double-conversion · runtime: ~7 min at full load
  • Chassis
    Supermicro 4U 8-GPU server chassis (CRPS, fan wall)form factor: 4U rack · gpu slots: 8 dual-slot passive · power: 2+2 CRPS redundant

Component prices are hand-sourced and dated (checked 2026-09-16), not scraped. Confirm on the vendor's site before buying.

Why this build
  • · One node = 8 RTX 6000 Ada, 384 GB of VRAM. The numbers shown here are for a single node; a cluster is three or more of them.
  • · The 100 GbE fabric (RoCEv2) carries tensor- and pipeline-parallel traffic between nodes and NVMe-over-fabric to shared storage - it is the part that makes it a cluster rather than a room of servers.
  • · Shared storage lives on a separate node (not priced here): every GPU node mounts the same dataset and model store.
  • · At this scale, colocation usually beats on-prem: metered power, cooling, and remote hands cost less than building it yourself.
  • · Cluster management (Slurm or Kubernetes + a scheduler) is assumed - the hardware is the easy half.
Pros
  • + Scales horizontally - add nodes, not rebuilds
  • + 384 GB VRAM per node, shared model store
  • + Fabric ready for multi-node training
Limits
  • − Per-node price only - shared storage, fabric switch and colocation are extra
  • − Needs an ops team: scheduling, monitoring, failover
  • − PCIe GPUs, not SXM - frontier-scale pretraining still wants an H100/B200 cluster
Alternatives
Cheaper

A6000 nodes: same 384 GB per node at roughly 60 % of the GPU cost.

Balanced

L40S nodes trade some throughput for datacenter-standard parts and support.

Performance

For frontier training, rent an H100/B200 pod - see GPU cloud pricing.

Buy vs rent vs hybrid

Over 36 months, at this build's power draw and the live cloud rate for its GPU.

Hybrid
Own it
Hardware
$69,670
Electricity
$6,606
Total cost of ownership
$76,276
Rent it (cloud)
Equivalent cloud, monthly
$2,698/mo
Equivalent cloud, total
$97,131
Break-even
28 months
  • · High-availability needs usually keep some cloud capacity in reserve even when owning makes sense - hybrid covers both.
  • · Over 36 months, owning saves roughly $20855 versus the cloud equivalent.

"Buy on Amazon" links are affiliate links - as an Amazon Associate, Obolith earns from qualifying purchases.

Concepts on this page