| Oracle Cloud | Bare metal GPU shapes, including 8 x H100 80 GB with 8 x 2 x 200 Gbps RDMA, and 8 x H200 141 GB with 8 x 400 Gbps RDMA. Oracle states that bare metal GPU shapes support cluster networking. | Is cluster networking available in your region, at the node count you need? Confirm reservation terms and delivery date. | Compute shapes ↗ |
| CoreWeave | Managed Kubernetes on bare metal nodes, with the hypervisor layer removed. Its GPU compute page lists 3,200 Gbps of NVIDIA Quantum-2 InfiniBand networking. | Does your team run on Kubernetes? Confirm the GPU model, node count, storage, and which operations are included. | Kubernetes service ↗ GPU compute ↗ |
| Lambda | 1-Click Clusters of NVIDIA HGX B200 or H100, from 16 to 2,000+ GPUs, on Quantum-2 InfiniBand with SHARP. Published terms run from 2 weeks to 1 year, with custom pricing for 1 year and longer. | The page does not state whether nodes are bare metal. Ask, and confirm network bandwidth per GPU for your cluster size. | 1-Click Clusters ↗ |
| Microsoft Azure | ND H100 v5 virtual machines with eight H100 GPUs and a 400 Gbps InfiniBand connection per GPU, scaling to thousands of GPUs. Not bare metal. | A fit when you want InfiniBand without managing hardware. Confirm quota, region, and whether capacity is reserved. | ND H100 v5 ↗ |
| AWS | P5 instances with eight H100 GPUs and 3,200 Gbps of EFA networking, in UltraClusters of up to 20,000 H100 or H200 GPUs. The P5 page lists no bare metal size. | EFA is AWS’s own network, not InfiniBand. Check your software stack supports it, and ask how capacity is reserved. | P5 instances ↗ |