01Capacity and growth
The capacity misses your launch date
A proposal may depend on hardware delivery, power availability or commissioning that has not finished. A promised start date then becomes a project dependency.
Evidence to request
Ask whether the proposed GPUs are installed, powered, commissioned and allocated to your project. Request the site, GPU count, acceptance milestones and dependencies. Hardware elsewhere in the fleet does not establish your allocation.
What to test or model
Validate access to the proposed environment before the production deadline. Keep a fallback plan for a smaller allocation, another region or another provider.
Terms to discuss
Discuss a firm acceptance date, staged payments and remedies for missed delivery, including termination or refund rights where agreed. Credits cannot restore a missed launch window.
Ask the provider: “What is ready today, what still needs to happen, and what happens to our payment and commitment if delivery slips?”
02Operational continuity
The provider cannot sustain the service
Financial distress, a change of operator or disruption to a supplier can affect continued access. A funding announcement alone does not answer how your service would continue.
Evidence to request
Identify the contracting entity and infrastructure operator. Request available financial information, ownership and financing context, relevant customer references and a business-continuity plan. Public information may not resolve the exposure.
What to test or model
Keep recoverable copies of critical data and deployment materials outside the service where appropriate. Test an export and estimate how long a transition would take.
Terms to discuss
Evaluate prepayment exposure, milestone billing, termination triggers and transition assistance with your procurement and legal teams. Contract rights do not guarantee recovery of prepaid funds.
Ask the provider: “If this service stops or changes hands, how do we recover the workload and how much prepaid money remains exposed?”
03Workload performance
The cluster does not perform as expected
The same GPU can deliver a different project result depending on networking, storage, software and failures. A machine being available does not establish that a training job will finish on time.
Evidence to request
Request GPU topology, inter-node networking, storage throughput assumptions, failure handling and node-replacement targets for the actual configuration being quoted.
What to test or model
Agree a representative proof of concept on the proposed cluster. Measure successful output, elapsed time, throughput, recovery and the complete bill using your software and data.
Terms to discuss
Define acceptance criteria and discuss measurable performance and replacement commitments. Read exclusions, claim procedures and remedies: an uptime SLA may provide service credits without promising your job outcome.
Ask the provider: “What must our workload demonstrate before we accept the allocation, and who resolves a failed test?”
Example: AWS Compute SLA defines availability and service-credit remedies ↗
04Cost and flexibility
The commitment outlasts the need
A project may stall, demand may change, or a more efficient model may need fewer GPUs. The original commitment can remain payable even when capacity sits idle.
Evidence to request
Model a baseline, a likely growth range and a downside scenario. Show total committed spend in each, including minimum allocations and paid idle time.
What to test or model
Use the pilot to refine utilization and capacity needs before choosing a longer term. Compare flexibility and capacity assurance separately from a discount.
Terms to discuss
Discuss shorter initial terms, phased ramps, resize or step-down rights, transfer rights and price-review mechanisms. Record which are accepted; do not assume they are standard or available.
Ask the provider: “What can we change if demand is half the forecast, twice the forecast, or delayed by six months?”
Example: CoreWeave describes different capacity purchase models ↗
05Cost and flexibility
The full bill exceeds the headline quote
Compute may be only one line item. Data loading, checkpoint storage, transfers, support, licenses and paid idle capacity can change the project budget.
Evidence to request
Request an itemized estimate using real storage volumes, transfer directions, checkpoint frequency, runtime and support needs. Separate GPU, whole-node, serverless and managed-service billing.
What to test or model
Reconcile the pilot invoice with the estimate. Include setup, idle periods, failed work and shutdown behavior. Check which storage charges continue after compute stops.
Terms to discuss
Record the unit prices, included allowances, billing increments, overages, taxes, price-validity period and any agreed pricing protections.
Ask the provider: “Show the full bill for our expected workload and our demand-surge scenario, including the cost to leave.”
Example: Runpod documents storage types, persistence and billing ↗
06Data and responsibilities
Security assumptions leave a gap
A provider-level statement may not cover the service, facility, tenant model or configuration you plan to use. Your team also retains operating responsibilities.
Evidence to request
Have your security team review the relevant reports and their scope, physical data location, isolation model, access controls, encryption, subprocessors and incident process. A logo or broad compliance claim is not the complete review.
What to test or model
Map who configures and monitors each control before uploading sensitive data. For HIPAA workloads, determine whether the selected service and required business associate agreement support your obligations.
Terms to discuss
Document the service scope, permitted locations, shared responsibilities, incident notification and required data agreements with your security and legal teams.
Ask the provider: “Which controls apply to our exact deployment, and which must our team implement?”
HHS guidance on HIPAA and cloud computing ↗
07Capacity and growth
Your allocation or expansion has lower priority
A discount, reservation label or ability to request more GPUs may not establish an enforceable allocation. Expansion can remain subject to future inventory.
Evidence to request
Separate the GPUs committed for launch from an option to request additional capacity. Ask for location, maximum concurrent allocation, notice periods, allocation priority and any reallocation or interruption rights.
What to test or model
Walk through a demand spike with the provider: the request process, approval steps, expected lead time, migration needs and fallback if the region is full.
Terms to discuss
Discuss dedicated capacity or a defined equivalent allocation, substitution rules, expansion commitments and remedies. Naming hardware should not prevent timely replacement of a failed node.
Ask the provider: “What additional capacity can you commit to, by when, in which region, and on what terms?”
08Data and responsibilities
Renewal becomes difficult to walk away from
Large datasets, provider-specific services and deployment dependencies can make migration expensive or slow. A renewal decision then depends on more than the next hourly rate.
Evidence to request
Inventory data volumes, export formats, container images, orchestration, identity, storage APIs and other dependencies. Kubernetes or Slurm can help standardize parts of a deployment, but they do not prove portability.
What to test or model
Restore a representative workload elsewhere. Measure transfer time and cost, validate checkpoints and record which components must change.
Terms to discuss
Discuss renewal notice, agreed pricing limits, export access, transition assistance, retention and deletion timing. Maintain an appropriate independent recovery copy.
Ask the provider: “How long would leaving take, what would it cost, and can we demonstrate that the workload still runs?”
09Operational continuity
A facility interruption stops the workload
A power, cooling or connectivity event can interrupt a long-running job. A second region on a provider map does not establish spare capacity or a working recovery path.
Evidence to request
Request the proposed site's power, cooling and connectivity design, maintenance approach, recovery dependencies and incident escalation process. Confirm whether alternate capacity is actually reserved.
What to test or model
Test checkpoint creation and restoration, including recovery outside the failed allocation where required. Measure the amount of work that can be lost and the time to resume.
Terms to discuss
Agree recovery responsibilities and discuss restoration targets, alternate capacity and remedies. Have counsel review outage exclusions and force-majeure language against the proposed recovery plan.
Ask the provider: “If this site is unavailable, where can we resume, how soon, and from which recoverable checkpoint?”