Network design and software support shape AI-cloud performance
Network design is a major source of performance differences between providers, and networks built for 100,000 GPUs differ substantially from those for 10,000 GPUs or smaller clusters. Some providers build custom networks, with Amazon EFA cited as an example. The source names InfiniBand and “Rocky” as two NVIDIA networking options.
New inference features in open-source tools such as vLLM and SGLang may take months to work well on custom networks. The delay can frustrate customers waiting for features such as KVCache offloading and lead them to prefer NVIDIA’s standard networking. Configuration, software and firmware support, reliability, and storage can also affect performance.
