<- Reference builds
Company server

Company AI Server

A proper GPU server: 4-8 cards, redundant power, BMC management and fast networking, for a company that has outgrown a workstation.

throughputmonitoringserviceability
updated 16 Sept 2026
Total price
$51,270
VRAM
192 GB
RAM
512 GB
Storage
12 TB
Est. power draw
1947 W
AI capability

Roughly ~80B params in fp16, ~160B in 8-bit, or ~295B in 4-bit quantization.

small-model inferencemid-size inferencelocal RAGlarge-model inferenceLoRA fine-tuningconcurrent servingtraining
Components
  • GPU
    4× NVIDIA L40Smemory: 48 GB GDDR6 ECC · bandwidth: 864 GB/s · interface: PCIe 4.0 x16
  • CPU
    AMD EPYC 9354cores: 32 · threads: 64 · clock: 3.8 GHz boost
  • Motherboard
    Supermicro H13SSL-N (SP5, EPYC 9004)chipset: SoC · memory: 12 x DDR5 RDIMM ECC, up to 3 TB · pcie: 8 x PCIe 5.0 x16 (slots + MCIO fan-out)
  • RAM
    Samsung 512 GB (8x64) DDR5-4800 ECC RDIMMcapacity: 512 GB (8 x 64) · type: DDR5-4800 RDIMM · ecc: ECC
  • Storage
    Kioxia CD8-R 3.84 TB U.2 NVMecapacity: 3.84 TB · bus: PCIe 4.0 U.2 · read: 7000 MB/s
    Sabrent Rocket 4 Plus 8 TB NVMecapacity: 8 TB · bus: PCIe 4.0 NVMe · read: 7100 MB/s
  • PSU
    FSP CRPS 2+2 redundant, 4000 W usablecapacity: 4000 W usable (2+2 hot-swap) · rating: 80+ Platinum · standard: CRPS redundant
  • Cooling
    Generic 4U high-static-pressure fan wall + shroudstype: rack forced-air · dissipation: chassis-wide · noise: loud (data-center)
  • Networking
    NVIDIA ConnectX-6 Dx 100GbEspeed: 100 GbE · ports: 2 (QSFP56) · bus: PCIe 4.0 x16
    MikroTik CRS309 8x 10GbE SFP+ switchports: 8 x SFP+ 10 GbE + 1 GbE · throughput: non-blocking · form factor: desktop / 1U
  • UPS
    Eaton 9PX 3000 RT (online double-conversion, 2U)capacity: 3000 VA / 2700 W · topology: online double-conversion · runtime: ~7 min at full load
  • Chassis
    Supermicro 4U 8-GPU server chassis (CRPS, fan wall)form factor: 4U rack · gpu slots: 8 dual-slot passive · power: 2+2 CRPS redundant

Component prices are hand-sourced and dated (checked 2026-09-16), not scraped. Confirm on the vendor's site before buying.

Why this build
  • · Four L40S: 192 GB of VRAM, passively cooled by the chassis fan wall - the density and airflow model a real server is built around.
  • · EPYC 9354 gives 128 PCIe 5.0 lanes and 12 memory channels: the whole server has bandwidth, not just the GPUs.
  • · 512 GB of ECC memory keeps large models and their KV cache addressable outside VRAM when needed.
  • · Redundant CRPS power and a BMC (IPMI / Redfish) mean remote power-cycling, health telemetry, and a hot-swap PSU - the difference between a server and a big PC.
  • · A 100 GbE NIC is there for the day this becomes node one of a cluster.
Pros
  • + 192 GB VRAM, serviceable and monitored
  • + Redundant power, remote management
  • + Headroom to 8 GPUs in the same chassis
Limits
  • − Loud - a dedicated room with cooling is not optional
  • − Still one node: cluster-grade availability needs at least two
  • − L40S has no NVLink - large-model training wants SXM parts
Alternatives
Cheaper

Four A6000s: same 192 GB, blower-cooled, roughly half the GPU cost.

Balanced

Four RTX 6000 Ada: 192 GB with the best single-card throughput short of datacenter SXM.

Performance

Multiple 8-GPU nodes with shared storage and a 100 GbE fabric.

Open ↗
Buy vs rent vs hybrid

Over 36 months, at this build's power draw and the live cloud rate for its GPU.

Hybrid
Own it
Hardware
$51,270
Electricity
$4,221
Total cost of ownership
$55,491
Rent it (cloud)
Equivalent cloud, monthly
$2,289/mo
Equivalent cloud, total
$82,388
Break-even
24 months
  • · High-availability needs usually keep some cloud capacity in reserve even when owning makes sense - hybrid covers both.
  • · Over 36 months, owning saves roughly $26897 versus the cloud equivalent.

"Buy on Amazon" links are affiliate links - as an Amazon Associate, Obolith earns from qualifying purchases.

Concepts on this page