Independent network infrastructure sourcing
Contact Us|
English
Solutions

Separate Storage and Compute Fabrics for GPU Training

Independent editorial analysis of third-party public reports. No SwitchInfra implementation or performance claim.

NVIDIA networking
Overview

Separate Storage and Compute Fabrics for GPU Training

Separate Storage and Compute Fabrics for GPU Training

Overview

A GPU cluster moves several kinds of data at once. Training workers exchange model information, storage systems deliver input data, and checkpoints create periodic write bursts. A useful network design begins by measuring those flows separately.

A documented architecture

CoreWeave describes H200 nodes with BlueField-3 storage Ethernet and a separate ConnectX-7/Quantum-2 GPU fabric. Its 64-node, CPU-driven fio read test reported roughly 500 GiB/s aggregate. This was a storage test, not a training-speed comparison or a GPUDirect Storage measurement. CoreWeave’s test and methodology CoreWeave’s B200 production announcement independently identifies one BlueField-3 DPU and eight ConnectX-7 HCAs per eight-GPU instance. B200 architecture

Turn workload traces into a network brief

For a proposed deployment, record input-read volume, checkpoint size, checkpoint interval and the number of ranks that write simultaneously. Include data-loader activity during ordinary training, not just the first minute of a job. Then measure GPU collective traffic while the storage test runs. Keep three acceptance questions separate: A design that passes an isolated bandwidth test may still fail one of these questions. Testing the combined workload makes the result more useful to the application team.

  • Can each node obtain the storage bandwidth its application needs?
  • Does checkpoint activity change collective completion time or training-step duration?
  • Can one tenant create a measurable slowdown for another under the intended isolation policy?

Allocate functions deliberately

In the procurement brief, assign each interface a role: GPU communication, storage, external access or management. State whether the DPU will run infrastructure services and which software image supplies them. NVIDIA’s CoreWeave case describes DOCA-based networking and tenant-isolation functions; the presence of a BlueField card alone does not establish that those services are configured. NVIDIA customer case

Validate before expanding

Run the same test on a single node, one rack and the intended production slice. Preserve the block sizes, read/write mix, queue depth, process count, dataset size and cache state. Report both per-node and aggregate results. Add a controlled link-failure test with an agreed recovery objective. Use this evidence to select adapter count, fabric capacity and storage connectivity. Ask for the complete order code, server qualification, protocol, firmware and cable pairing before issuing a purchase order. This solution is an engineering framework informed by third-party deployments. It does not reproduce CoreWeave’s proprietary implementation or establish that any particular catalog variant was deployed there. Sources checked: 10 October 2026. This article discusses third-party evidence or an explicitly labeled engineering scenario; it does not establish a reseller’s delivery history, current stock or exact-SKU deployment.