Compute
What model are you serving, to how many concurrent users, and at what latency? This drives GPU count and everything downstream.
Deployment
Workload
Service level
50
Milliseconds per output token. Caps batch size, so it trades per-user speed against GPU count.
1000
Networking
The fabric follows from the node count. East-west carries GPU-to-GPU traffic and is non-blocking; north-south carries client and storage traffic at 4:1.
Topology
Tensor parallelism spanning nodes needs 1:1. Anything less throttles collectives.
Storage
The POSIX/S3/AI tier feeds the GPUs. Size it by capacity, by the throughput the fleet needs, or by matching the reference architecture.
WEKA tier
4
GB/s per GPU. 2–4 suits inference and RAG; training checkpoints want more.
TB net
WEKA sizes the WEKApod at 5+2 with one virtual hot spare.
Results
The whole solution: compute, fabric, storage, facility and rack elevations.