Skip to main content
Start your own AI-powered blog — freeGet started →

H100 Cloud Rental Cost 2026: Provider Price Guide

H100 Cloud Rental Cost 2026: Provider Price Guide
Photo by Jakub Zerdzicki on pexels

H100 Cloud Rental Cost 2026: Provider Price Guide

Stack of Polish zloty banknotes on financial documents with a pen, indicating monetary transactions in an office setting. Photo by Jakub Zerdzicki on Pexels

Quick Answer: In 2026, H100 cloud rental ranges from $1.20/hr (io.net spot) to $4.73/hr (AWS on-demand). The best value for most AI workloads is Lambda Labs at $2.49/hr (reliable, good support, consistent performance) or RunPod at $2.29/hr (good availability, community features). For the absolute cheapest H100 access, io.net spot at $1.20-1.80/hr is 75% cheaper than AWS but with 10-20% failure rates. For production serving with SLA requirements, Lambda Labs or AWS are the only viable choices. New H200 and B200 GPUs are available at 10-20% premium over H100 pricing.

H100 Pricing Comparison: All Providers

Single H100 Pricing ($/hr)

ProviderOn-DemandSpot/BidReserved (1mo)Reserved (1yr)
AWS (p5.xlarge)$4.73~$3.20 (1yr commit)
Google Cloud (a3-high)$4.52$2.85$3.10 (1yr commit)
Azure (NCads)$4.65$3.10$3.25 (1yr commit)
Lambda Labs$2.49$1,793/mo ($2.45/hr)$1,525/mo (6mo commit)
RunPod$2.29$1.85
Vast.ai$1.89$1.50
Akash$2.20-$3.00
io.net$1.89-$2.59$1.20-$1.80
Together AI$2.85Included in API
Fireworks$2.95Included in API

8x H100 Instance Pricing ($/hr)

ProviderInstance TypeOn-DemandSpotBest For
AWS p5.48xlarge8x H100$151.00Production, enterprise
Lambda Labs 8x8x H100$19.92Training at scale
RunPod 8x8x H100$17.84$14.50Budget training
Vast.ai 8x8x H100$13.50$10.50Experimentation
Azure ND H100 v58x H100$148.00$95.00Enterprise

8x H100 is where the price gap widens dramatically: AWS charges $151/hr for 8x H100 while Lambda Labs charges $19.92/hr (87% less). AWS includes NVLink, better networking, and enterprise support — but the 7.5x price premium is hard to justify for most workloads.

On-Demand vs Spot vs Reserved Pricing

Savings by Commitment Type

ProviderOn-Demand BaselineSpot Savings1-Month Reserved1-Year Reserved
AWS$4.73/hrN/AN/A~32% ($3.20/hr)
Lambda Labs$2.49/hrN/A~2% ($2.45/hr)~15% (6mo)
RunPod$2.29/hr~19% ($1.85/hr)N/AN/A
Vast.ai$1.89/hr~21% ($1.50/hr)N/AN/A
io.net$2.24/hr (avg)~33% ($1.50/hr avg)N/AN/A

When to Use Each

Pricing ModelBest ForWorst For
On-DemandProduction serving, consistent workloadsExperimentation (too expensive)
Spot/CheapBatch processing, hyperparameter sweepsProduction (preemption risk)
ReservedKnown weekly/monthly workload patternsVariable/spiky usage (commitment waste)

Spot Preemption Rates

ProviderPreemption Rate (spot)Average Runtime Before Preemption
RunPod spot5-15%12-24 hours
Vast.ai bid20-40%4-12 hours
io.net spot30-50%2-8 hours
GCP spot5-10%24+ hours

H100 vs H200 vs B100 Pricing

New GPU Pricing (2026)

GPUVRAMBest ProviderPrice/hrTokens/Sec (70B FP8)Cost/1M Tokens
H10080 GB HBM3Lambda Labs$2.4955$0.0126
H200141 GB HBM3eLambda Labs$2.9968$0.0122
B100120 GB HBM3eCoreWeave$3.5095$0.0102
B200180 GBCoreWeave$4.25120$0.0098

Is the H200/B100 Worth the Premium?

WorkloadH100 ($2.49/hr)H200 ($2.99/hr)B100 ($3.50/hr)Recommendation
70B inference (batch=1)55 tok/s68 tok/s95 tok/sH100 is cheapest per token
70B inference (batch=32)450 tok/s580 tok/s750 tok/sB100 if throughput matters
70B full fine-tune24 hrs18 hrs12 hrsB100 for faster experiments
7B inference high volume8,100 tok/s10,200 tok/s14,500 tok/sB100 at scale — savings add up

For most teams, H100 is the sweet spot. H200 gives 24% more throughput for 20% more cost — marginally better value. B100 is faster but at 40% premium over H100. Upgrade only if you're bottlenecked on throughput.

Three NVIDIA GeForce RTX graphics cards stacked on a surface, showcasing their sleek design and branding details. Photo by Andrey Matveev on Pexels

Cost per Million Tokens: The Real Metric

70B Model Serving Cost

ProviderGPU$/hrTokens/Sec (batch=1)Tokens/Sec (optimum batch)$/1M tokens
io.net (spot)H100$1.2055450 (batch=32)$0.0007
Vast.ai (bid)H100$1.5055450$0.0009
RunPod (spot)H100$1.8555450$0.0011
AkashH100$2.6055450$0.0016
RunPod (on-demand)H100$2.2955450$0.0014
Lambda LabsH100$2.4955450$0.0015
AWSH100$4.7355450$0.0029

8B Model Serving Cost

Provider$/hrTokens/Sec$/1M tokens
io.net (spot) — RTX 4090$0.2285$0.0007
Vast.ai (bid) — RTX 4090$0.2585$0.0008
RunPod — RTX 4090$0.5185$0.0017
Lambda Labs — H100$2.49305$0.0023
AWS — L4$1.5980$0.0055

Cheapest per-token inference: io.net spot H100 ($0.0007/1M tokens) for 70B. For small models, consumer GPUs on DePIN networks are dramatically cheaper per token than H100s.

Training Cost: H100 vs Consumer GPU Rental

Fine-Tuning Cost Comparison (Llama 3.1 70B, QLoRA, 1 epoch, 1000 samples)

GPUProvider$/hrTraining TimeTotal Cost
H100Lambda Labs$2.490.5 hrs$1.25
H100RunPod (spot)$1.850.5 hrs$0.93
H100AWS$4.730.5 hrs$2.37
RTX 5090RunPod$0.691.5 hrs$1.04
RTX 4090Vast.ai$0.352.5 hrs$0.88
RTX 4090RunPod$0.512.5 hrs$1.28
2x RTX 3090Vast.ai$0.502.0 hrs$1.00

Full Training Cost (70B, LoRA, 1 epoch, 10K samples)

GPUProvider$/hrTraining TimeTotal Cost
8x H100AWS$151.004 hrs$604
8x H100Lambda Labs$19.924 hrs$80
8x H100RunPod$17.844 hrs$71
8x H100Vast.ai$13.504 hrs$54

Hidden Costs: Storage, Transfer, and Engineering Time

Storage Costs

ProviderStorage IncludedAdditional StorageBlock Storage
AWS8GB (EBS boot)$0.08/GB-month (gp3)$0.08/GB-month
Lambda Labs200GB NVMe$0.10/GB-monthIncluded
RunPod5GB (template)$0.07/GB-month$7/TB-month
Vast.ai50GB (instance)$0.05/GB-monthVariable
io.net10GB$0.10/GB-monthN/A

Data Transfer (Egress)

ProviderEgress CostFree Tier
AWS$0.05-0.09/GB100GB/month
Lambda Labs$0.05/GB500GB/month
RunPod$0.01/GB1TB/month
Vast.ai$0.01/GB200GB/month
io.net$0.02/GB100GB/month

Real Monthly Cost Example

Workload: Batch inference, 70B, 50M tokens/day, H100

ProviderCompute (30 days)Storage (1TB)Egress (500GB)TotalEffective $/hr
AWS$3,405$80$35$3,520$4.88
Lambda Labs$1,793$100$25$1,918$2.66
RunPod$1,649$70$5$1,724$2.39
Vast.ai$1,361$50$5$1,416$1.97
io.net (on-demand)$1,361$100$10$1,471$2.04
io.net (spot)$864$100$10$974$1.35

Provider Rankings by Use Case

Best for Production Inference (SLA Required)

code
1. Lambda Labs  ★★★★★  $2.49/hr, 99.5% uptime, good support
2. RunPod       ★★★★☆  $2.29/hr, 98% uptime, community support
3. AWS          ★★★☆☆  $4.73/hr, 99.9% uptime, enterprise pricing

Best for Training

code
1. RunPod       ★★★★★  $2.29/hr (on-demand), $1.85/hr (spot)
2. Lambda Labs  ★★★★★  $2.49/hr, consistent performance
3. Vast.ai      ★★★★☆  $1.89/hr, but variable quality

Best for Batch Inference (Cheapest)

code
1. io.net spot  ★★★★★  $1.20/hr — can't beat the price
2. Vast.ai bid  ★★★★☆  $1.50/hr — more reliable than io.net
3. Akash        ★★★★☆  $2.20-3.00/hr — most decentralized

Best for Experimentation

code
1. Vast.ai      ★★★★★  $1.50-1.89/hr, widest selection
2. RunPod       ★★★★★  $1.85-2.29/hr, best UX
3. io.net       ★★★★☆  $1.20-2.59/hr, cheapest spot

Related Reads

Optimizing H100 Workloads for Cost Efficiency

Beyond choosing the right provider, small adjustments to workload configuration can yield outsized cost savings. For inference, batch size is the most critical lever: increasing batch size from 1 to 32 on a 70B model boosts throughput by 8x (from 55 to 450 tokens/sec) while only doubling GPU memory usage. This reduces cost per million tokens by 75%—transforming a $0.0029/1M token AWS deployment into a $0.0007/1M token operation on io.net. However, larger batches introduce latency; for real-time applications, benchmark your model’s latency vs. throughput tradeoff at batch sizes of 1, 2, 4, 8, 16, and 32 to find the optimal balance.

For training, mixed precision (FP8) and gradient checkpointing can cut costs by 30-50% without sacrificing model quality. FP8 reduces memory usage by 50% compared to FP16, enabling larger batch sizes or fitting larger models on a single H100. Gradient checkpointing trades compute for memory, reducing VRAM usage by 30-40% at the cost of a 20-30% increase in training time. When combined, these techniques can reduce the cost of a 70B fine-tuning run from $80 to $40 on RunPod. Providers like Lambda Labs and RunPod offer pre-configured templates with these optimizations enabled, while AWS requires manual setup via custom AMIs or Docker containers.

Network Topology and Multi-GPU Scaling

The performance gap between single H100s and 8x H100 instances isn’t linear due to networking bottlenecks. AWS’s p5.48xlarge instances include NVLink and 3.2Tbps networking, enabling near-linear scaling for distributed training (e.g., 7.8x speedup for 8x GPUs). In contrast, providers like Lambda Labs and RunPod use PCIe 4.0 or 5.0 for multi-GPU communication, which caps scaling efficiency at 6-7x for 8x GPUs. For workloads like full fine-tuning or large-scale inference, this difference can add 20-30% to runtime costs. Benchmark your workload’s scaling efficiency by comparing single-GPU performance to 2x, 4x, and 8x configurations—if scaling efficiency drops below 80%, consider splitting workloads across multiple single-GPU instances instead of using a single 8x instance.

For inference, multi-GPU setups are rarely cost-effective unless you’re serving at scale. A single H100 can handle ~500 tokens/sec for a 70B model at batch=32, which translates to ~130M tokens/day—enough for most production workloads. If you need higher throughput, deploy multiple single-GPU instances behind a load balancer (e.g., Nginx or Traefik) rather than using an 8x H100 instance. This approach improves fault tolerance and reduces costs by 30-50% compared to a single 8x instance, as you can scale horizontally with spot instances.

Provider-Specific Quirks and Workarounds

Each H100 provider has unique limitations that can impact cost and performance if not accounted for:

  • AWS: EBS boot volumes are slow (100-200 MB/s); use instance storage for datasets or cache. Egress costs ($0.05-0.09/GB) can exceed compute costs for large datasets—compress data or use AWS’s free tier (100GB/month) for transfers.
  • Lambda Labs: No spot instances, but their 1-month reserved pricing ($1,793/mo) is effectively a 2% discount. Their NVMe storage is included up to 200GB, making them ideal for storage-heavy workloads like video processing.
  • RunPod: Spot instances have a 5-minute preemption warning, but their API doesn’t expose this—use their webhook feature to trigger checkpointing. Their $0.01/GB egress is the cheapest among providers, but transfers are throttled to 1Gbps.
  • Vast.ai: Provider quality varies wildly; filter for "H100 PCIe 4.0" or "NVLink" in the search bar to avoid underperforming instances. Their bid system can save 20-30% over on-demand, but lowball bids may get preempted quickly.
  • io.net: Spot instances are the cheapest ($1.20/hr) but have no preemption warning—implement a heartbeat system to detect failures. Their storage is ephemeral, so use external storage (e.g., S3 or Backblaze B2) for datasets.

For production workloads, test each provider’s networking performance with tools like ib_write_bw (for InfiniBand) or iperf3 (for TCP). AWS and Azure offer 100Gbps+ networking, while alternative providers typically cap at 25-50Gbps. If your workload involves frequent data transfers (e.g., distributed training), this can add 10-20% to runtime costs. For storage-bound workloads, benchmark disk I/O with fio—Lambda Labs’ NVMe storage delivers 3-5GB/s, while AWS’s gp3 EBS tops out at 1GB/s.

Key Takeaways

  • For most AI workloads, Lambda Labs ($2.49/hr) or RunPod ($2.29/hr) offer the best balance of cost, reliability, and support—avoid AWS ($4.73/hr) unless you need enterprise SLAs or tight cloud integration.
  • Spot instances (e.g., io.net at $1.20/hr) can cut costs by 75% but carry 30-50% preemption risk; use them only for fault-tolerant workloads like batch inference or hyperparameter sweeps with checkpointing every 10-15 minutes.
  • H100 remains the cost-performance sweet spot in 2026: H200 offers 24% more throughput for a 20% price premium, while B100’s 40% premium is only justified for throughput-bottlenecked workloads like high-volume 70B inference or full fine-tuning.
  • For 70B inference, io.net spot H100s deliver the lowest cost per million tokens ($0.0007), but consumer GPUs (e.g., RTX 4090 on Vast.ai at $0.0008/1M tokens) are dramatically cheaper for smaller models like 8B.
  • Hidden costs add up: AWS charges $0.08/GB-month for storage and $0.05-0.09/GB egress, while providers like RunPod ($0.01/GB egress) or Lambda Labs (500GB free egress) can reduce total monthly costs by 20-40%.
  • For full fine-tuning (LoRA, 10K samples), expect $50-80 on RunPod/Vast.ai vs $600 on AWS; spot instances can further reduce training costs by 15-30% if checkpointing is implemented.

Frequently Asked Questions

Where is the cheapest place to rent an H100?

io.net spot at $1.20/hr is the cheapest on paper. However, with 30-50% preemption rates and 10-20% job failure rates, the effective cost is higher. For practical cheapest H100: Vast.ai bid at $1.50/hr with good provider selection. For reliable cheap: RunPod spot at $1.85/hr.

Is AWS H100 worth the premium?

For production serving with SLA requirements: yes. AWS offers 99.9%+ uptime, 24/7 enterprise support, and seamless integration with other AWS services. For batch processing and training: no — Lambda Labs or RunPod provide 95% of the reliability at 50% of the cost.

What's the difference between H100 and H200 for cloud rental?

H200 has 141 GB VRAM (vs 80 GB on H100) and slightly faster HBM3e memory. For inference, H200 yields ~24% more throughput. For training, H200 fits larger models without offloading. H200 costs ~20% more per hour, making it marginally better value.

How much does it cost to train Llama 3.1 70B on cloud H100s?

A full continued pre-training run ($500K-2M) is only for large AI labs. For fine-tuning (LoRA, 10K samples, 1 epoch), expect $50-80 on Vast.ai/RunPod or $600 on AWS. For full fine-tuning (all parameters, small dataset), expect $500-2,000 on alternative providers vs $5,000+ on AWS.

Can I save money by using spot instances for training?

Yes — if you implement checkpointing (every 10-15 minutes). Spot savings of 20-40% are real, and even with occasional preemption, the effective cost is still 15-30% less than on-demand. Without checkpointing, spot is risky for training.

S
Synor

1 followers

Deep dives on GPUs, decentralized AI, crypto, and open-source ML — buying guides, benchmarks, and tax/compliance explainers.

Comments

Sign in to join the conversation

No comments yet. Be the first to share your thoughts!

More from Synor

Recommended for you