Contact Us|
English
Home/HPC Networking/100/200G AI Storage Network Solution

100/200G AI Storage Network Solution

Fully supports RoCEv2, ensuring lossless transmission of GPU-Direct RDMA, Ceph, and related traffic

Contact Us

100/200G AI Storage Network Solution

SwitchInfra 100/200G AI storage network solution delivers lossless, ultra-low-latency Ethernet networking purpose-built for AI workloads. Powered by PicOS® switches with deep buffers and intelligent congestion control, the solution eliminates I/O bottlenecks that starve GPUs of data. It fully supports RoCEv2 with PFC, ECN, and DCQCN to ensure lossless transmission of GPU-Direct RDMA, Ceph, and related storage traffic. Standard Ethernet architecture enables unified operations, seamless scaling, and lower O&M costs compared to proprietary interconnects.

100/200G Low-Latency Switching

PicOS® switches with PFC, ECN, and deep buffers sustain lossless throughput for AI storage workloads. Hardware-based flow control prevents packet loss during microbursts, keeping GPUs fed with data.

GPU-Direct RDMA Support

RoCEv2 with DCQCN enables GPU-Direct RDMA, bypassing CPU and system memory to dramatically reduce latency. Accelerates Ceph distributed storage access and NCCL collective operations.

Unified Ethernet Fabric

Standard Ethernet architecture eliminates vendor lock-in, simplifies operations, and reduces total cost of ownership. Unified management via AmpCon-DC provides end-to-end visibility across compute and storage networks.

Why SwitchInfra AI Storage Network Solution

Lossless RoCEv2

DCQCN, PFC, and ECN work together to eliminate packet loss. Shared buffer architecture absorbs microbursts that would otherwise cause drops and retransmissions in AI storage traffic.

Deep Buffers

Substantial on-chip buffer capacity handles the bursty, synchronized I/O patterns typical of distributed AI training, preventing incast congestion collapse at the storage aggregation layer.

Intelligent Load Balancing

DLB and flowlet-based hashing distribute storage traffic evenly across all available links, avoiding hot spots and maximizing bisectional bandwidth for parallel file system access.

Telemetry-Driven Ops

AmpCon-DC provides real-time visibility into PFC counters, ECN marks, and buffer utilization. Proactive congestion alerts let operators intervene before job performance degrades.

Solution Architecture

The AI storage network follows a two-tier spine-leaf architecture optimized for east-west storage traffic. Storage leaf switches connect directly to GPU servers via 100G/200G links, while spine switches provide non-blocking interconnect between storage leaves and the parallel file system. RDMA over Converged Ethernet (RoCEv2) runs end-to-end, with PFC providing lossless Ethernet transport and ECN marking congestion before queues overflow.

Storage Leaf Layer

100G/200G PicOS® switches with deep buffers connect directly to GPU servers. PFC ensures lossless RoCEv2 transport. Each leaf supports 32-64 server ports for high-density AI clusters.

Storage Spine Layer

Non-blocking spine switches interconnect all storage leaves with 100G/200G uplinks. DLB distributes traffic across all available paths, maximizing throughput for parallel file system workloads.

Management Plane

AmpCon-DC automates switch provisioning and continuously monitors PFC counters, ECN marks, and buffer utilization. REST APIs integrate with existing DCIM/ITSM tools for unified operations.

Contact Us

Ready to eliminate AI storage bottlenecks?

SwitchInfra 100/200G AI storage network solution fully supports RoCEv2, ensuring lossless transmission of GPU-Direct RDMA, Ceph, and related traffic. Connect with our HPC experts for a tailored design.

+86 132 6592 2480

Expert customer service, 5x24 Phone Support

[email protected]
0/5000