InferX Guided Sizing

Compute → networking → storage → results
Compute Storage

Compute

What model are you serving, to how many concurrent users, and at what latency? This drives GPU count and everything downstream.

Deployment

Workload

Service level

50
Milliseconds per output token. Caps batch size, so it trades per-user speed against GPU count.
1000