DEDICATED SERVERS. A FAST FABRIC.

Bare metal GPU clusters.
Built for multi-node training.

Rent whole GPU servers with no hypervisor, joined by an InfiniBand or RDMA network. See what bare metal changes, what the network has to deliver, and what to confirm before you sign.

01 / WHAT BARE METAL CHANGES

No hypervisor.
The whole server.

A bare metal GPU server is a physical machine allocated to one customer, with no virtualization layer between your software and the hardware. Ask the provider to confirm that in the proposal. The label alone is not a specification.

Bare metal server

You get the entire physical server, usually eight GPUs, with direct access to the GPUs and network cards. That often means your team manages the operating system, drivers, and libraries. Confirm who is responsible for firmware, driver updates, and hardware failures before you commit.

GPU virtual machine

A virtual machine adds a layer between your software and the server. It can still give you dedicated GPUs and a fast network. Azure’s ND H100 v5 is a virtual machine with eight H100 GPUs and a dedicated 400 Gbps InfiniBand connection for each GPU. Bare metal and InfiniBand are separate questions.

02 / THE NETWORK IS THE CLUSTER

Inside the server.
Between the servers.

Multi-node training spends much of its time exchanging data between GPUs. Two links matter, and providers specify them separately.

Inside each server

GPUs in the same server talk over NVLink. NVIDIA lists 900 GB/s of NVLink bandwidth for H100 SXM and 600 GB/s for H100 NVL. Confirm which variant each node carries.

Between servers

Servers talk over InfiniBand or an RDMA network. Published examples differ in form. Oracle lists 8 x 2 x 200 Gbps of RDMA for its bare metal H100 shape. AWS lists 3,200 Gbps of its own EFA networking for P5 instances. Compare bandwidth per GPU, not just the total.

How the fabric is built

Ask whether the network is non-blocking or oversubscribed, and by how much. Ask whether all your servers sit on the same fabric. AWS describes its UltraClusters as a petabit-scale nonblocking network. Get the same kind of statement in writing for your cluster.

Ask for proof, not a brochure.

Before acceptance, ask for collective communication test results across the exact number of servers you will rent. NVIDIA’s open source NCCL tests check both the performance and the correctness of these operations. Agree on the target result in the contract, then run the test again yourself on day one.

03 / BARE METAL OR VIRTUAL MACHINES

When bare metal
is worth it.

Neither option is faster in every case. The right choice depends on how long the work runs, how much control your team needs, and who will operate it.

Bare metal tends to fit

  • Multi-node training that runs for weeks or months
  • Custom drivers, libraries, schedulers, or kernel settings
  • Strict isolation or security review requirements
  • Steady use that justifies a reserved term

Virtual machines tend to fit

  • Single-server jobs, fine-tuning, and development
  • Inference that scales up and down with demand
  • Teams that want the provider to manage drivers and the operating system
  • Short or uncertain timelines

Performance differences depend on the workload. Benchmark your own training step on the proposed configuration. For single servers and shorter terms, see our GPU server rental guide.

04 / PUBLISHED CLUSTER OPTIONS

Five providers.
Different cluster models.

Use these examples to see how providers package multi-node GPU capacity. They are a starting point for questions, not a ranking.

Scroll the table horizontally to see the full comparison and official sources.

Five GPU cloud providers, their published multi-node cluster options, questions to confirm, and official sources.
ProviderPublished optionWhat to confirmOfficial sources
Oracle CloudBare metal GPU shapes, including 8 x H100 80 GB with 8 x 2 x 200 Gbps RDMA, and 8 x H200 141 GB with 8 x 400 Gbps RDMA. Oracle states that bare metal GPU shapes support cluster networking.Is cluster networking available in your region, at the node count you need? Confirm reservation terms and delivery date.Compute shapes
CoreWeaveManaged Kubernetes on bare metal nodes, with the hypervisor layer removed. Its GPU compute page lists 3,200 Gbps of NVIDIA Quantum-2 InfiniBand networking.Does your team run on Kubernetes? Confirm the GPU model, node count, storage, and which operations are included.Kubernetes service GPU compute
Lambda1-Click Clusters of NVIDIA HGX B200 or H100, from 16 to 2,000+ GPUs, on Quantum-2 InfiniBand with SHARP. Published terms run from 2 weeks to 1 year, with custom pricing for 1 year and longer.The page does not state whether nodes are bare metal. Ask, and confirm network bandwidth per GPU for your cluster size.1-Click Clusters
Microsoft AzureND H100 v5 virtual machines with eight H100 GPUs and a 400 Gbps InfiniBand connection per GPU, scaling to thousands of GPUs. Not bare metal.A fit when you want InfiniBand without managing hardware. Confirm quota, region, and whether capacity is reserved.ND H100 v5
AWSP5 instances with eight H100 GPUs and 3,200 Gbps of EFA networking, in UltraClusters of up to 20,000 H100 or H200 GPUs. The P5 page lists no bare metal size.EFA is AWS’s own network, not InfiniBand. Check your software stack supports it, and ask how capacity is reserved.P5 instances

Examples for evaluation, not an exhaustive directory. Listing a provider does not imply a reseller relationship or guaranteed access to capacity. The sources describe published options, not confirmed inventory for your project. Ask each provider to confirm the configuration, location, support, and terms in its proposal. For more providers, see GPU cloud providers. For billing examples, see GPU cloud pricing.

Planning to buy your own GPUs instead? High-density colocation providers such as TierPoint host customer-owned clusters. Our sister site Colocation Scout compares colocation options.

05 / MOVING FROM VIRTUAL MACHINES

From GPU VMs
to a cluster.

Most teams move to bare metal after outgrowing virtual machines. Plan the move so training keeps running while the new cluster proves itself.

1. Benchmark first

Run your real training step and a communication test on two nodes of the proposed configuration. Scale the test up before you sign for the full cluster.

2. Package the environment

Build container images with pinned CUDA, NCCL, and driver versions. On bare metal, confirm who installs and updates the drivers on the servers themselves.

3. Move data and checkpoints

Stage datasets and the latest checkpoints on the new cluster’s storage before cutover. Confirm read throughput per server, and any charges to move data out of the old environment.

Cut over when the new cluster passes.

Resume from a checkpoint on the new cluster and compare results with the old environment. Keep the old capacity until the new cluster passes the acceptance tests you agreed. Ask each provider what migration help, if any, is included in the proposal.

06 / YOUR CLUSTER CHECKLIST

Bring a clear brief.
Get comparable proposals.

Share these details with an advisor, even if some are still estimates.

Compute and network

  • GPU model, variant, and memory per GPU
  • Number of servers and GPUs per server
  • Network type and bandwidth per GPU
  • Fabric design, oversubscription, and server placement
  • Connection between GPUs inside each server

Operations and contract

  • Scheduler, such as Slurm or Kubernetes, and who runs it
  • Storage type and read throughput per server
  • Responsibility for drivers, firmware, and the operating system
  • Acceptance tests, failed-server replacement time, and support hours
  • Start date, term, reserved capacity, and exit terms

Provider availability and configurations change. Request confirmation for your dates and region. Technical sources checked .

LET’S FIND YOUR WAY FORWARD

Your next move
starts with a conversation.

Meet with an advisor to discuss your cluster requirements, explore realistic options, and agree on the next step for your project.

Our advisory services are free to you. We’re compensated by whichever provider you choose through us.

30-minute consultationNo obligation
YOUR GPU CLUSTER CONSULTATION

Bring your questions.
We’ll bring perspective.

A few details help us make the most of your time:

  • The model or job you plan to run
  • GPU count and cluster size, if known
  • Target location and start date
  • Budget range or an existing proposal
Meet with an advisor

Book securely on Calendly. Times appear in your time zone.

Prefer to call? 844-506-2299

BARE METAL CLUSTER QUESTIONS

Before you
reserve a cluster.

Bring your requirements or an existing quote.
Meet with an advisor

Is bare metal faster than a GPU virtual machine?

Not in every case. The GPUs do the same work either way. Differences tend to show up in network setup, driver control, and what else shares the server. Benchmark your own job on each proposed configuration.

Do I need bare metal to get InfiniBand?

No. Azure’s ND H100 v5 virtual machines include a 400 Gbps InfiniBand connection for each GPU. Confirm the network separately from whether the servers are bare metal.

What is the smallest cluster I can rent?

It varies by provider. Lambda’s published 1-Click Clusters start at 16 GPUs. Other providers rent a single eight-GPU server. Ask for the minimum configuration and whether you can add servers later on the same network.

How long are GPU cluster contracts?

Terms vary. Lambda publishes terms from 2 weeks to 1 year, with custom pricing beyond 1 year. Compare the total obligation, and ask whether the contract actually reserves capacity or only commits spend.

What does your advisory service cost?

Our advisory services cost you nothing. We’re compensated by whichever provider you choose through us. You pay the provider for your infrastructure and services, with no obligation to choose a provider.