H100 Cloud Rental Cost 2026: Provider Price Guide

H100 Cloud Rental Cost 2026: Provider Price Guide
Photo by Jakub Zerdzicki on Pexels
Quick Answer: In 2026, H100 cloud rental ranges from $1.20/hr (io.net spot) to $4.73/hr (AWS on-demand). The best value for most AI workloads is Lambda Labs at $2.49/hr (reliable, good support, consistent performance) or RunPod at $2.29/hr (good availability, community features). For the absolute cheapest H100 access, io.net spot at $1.20-1.80/hr is 75% cheaper than AWS but with 10-20% failure rates. For production serving with SLA requirements, Lambda Labs or AWS are the only viable choices. New H200 and B200 GPUs are available at 10-20% premium over H100 pricing.
H100 Pricing Comparison: All Providers
Single H100 Pricing ($/hr)
| Provider | On-Demand | Spot/Bid | Reserved (1mo) | Reserved (1yr) |
|---|---|---|---|---|
| AWS (p5.xlarge) | $4.73 | — | — | ~$3.20 (1yr commit) |
| Google Cloud (a3-high) | $4.52 | $2.85 | — | $3.10 (1yr commit) |
| Azure (NCads) | $4.65 | $3.10 | — | $3.25 (1yr commit) |
| Lambda Labs | $2.49 | — | $1,793/mo ($2.45/hr) | $1,525/mo (6mo commit) |
| RunPod | $2.29 | $1.85 | — | — |
| Vast.ai | $1.89 | $1.50 | — | — |
| Akash | $2.20-$3.00 | — | — | — |
| io.net | $1.89-$2.59 | $1.20-$1.80 | — | — |
| Together AI | $2.85 | — | Included in API | — |
| Fireworks | $2.95 | — | Included in API | — |
8x H100 Instance Pricing ($/hr)
| Provider | Instance Type | On-Demand | Spot | Best For |
|---|---|---|---|---|
| AWS p5.48xlarge | 8x H100 | $151.00 | — | Production, enterprise |
| Lambda Labs 8x | 8x H100 | $19.92 | — | Training at scale |
| RunPod 8x | 8x H100 | $17.84 | $14.50 | Budget training |
| Vast.ai 8x | 8x H100 | $13.50 | $10.50 | Experimentation |
| Azure ND H100 v5 | 8x H100 | $148.00 | $95.00 | Enterprise |
8x H100 is where the price gap widens dramatically: AWS charges $151/hr for 8x H100 while Lambda Labs charges $19.92/hr (87% less). AWS includes NVLink, better networking, and enterprise support — but the 7.5x price premium is hard to justify for most workloads.
On-Demand vs Spot vs Reserved Pricing
Savings by Commitment Type
| Provider | On-Demand Baseline | Spot Savings | 1-Month Reserved | 1-Year Reserved |
|---|---|---|---|---|
| AWS | $4.73/hr | N/A | N/A | ~32% ($3.20/hr) |
| Lambda Labs | $2.49/hr | N/A | ~2% ($2.45/hr) | ~15% (6mo) |
| RunPod | $2.29/hr | ~19% ($1.85/hr) | N/A | N/A |
| Vast.ai | $1.89/hr | ~21% ($1.50/hr) | N/A | N/A |
| io.net | $2.24/hr (avg) | ~33% ($1.50/hr avg) | N/A | N/A |
When to Use Each
| Pricing Model | Best For | Worst For |
|---|---|---|
| On-Demand | Production serving, consistent workloads | Experimentation (too expensive) |
| Spot/Cheap | Batch processing, hyperparameter sweeps | Production (preemption risk) |
| Reserved | Known weekly/monthly workload patterns | Variable/spiky usage (commitment waste) |
Spot Preemption Rates
| Provider | Preemption Rate (spot) | Average Runtime Before Preemption |
|---|---|---|
| RunPod spot | 5-15% | 12-24 hours |
| Vast.ai bid | 20-40% | 4-12 hours |
| io.net spot | 30-50% | 2-8 hours |
| GCP spot | 5-10% | 24+ hours |
H100 vs H200 vs B100 Pricing
New GPU Pricing (2026)
| GPU | VRAM | Best Provider | Price/hr | Tokens/Sec (70B FP8) | Cost/1M Tokens |
|---|---|---|---|---|---|
| H100 | 80 GB HBM3 | Lambda Labs | $2.49 | 55 | $0.0126 |
| H200 | 141 GB HBM3e | Lambda Labs | $2.99 | 68 | $0.0122 |
| B100 | 120 GB HBM3e | CoreWeave | $3.50 | 95 | $0.0102 |
| B200 | 180 GB | CoreWeave | $4.25 | 120 | $0.0098 |
Is the H200/B100 Worth the Premium?
| Workload | H100 ($2.49/hr) | H200 ($2.99/hr) | B100 ($3.50/hr) | Recommendation |
|---|---|---|---|---|
| 70B inference (batch=1) | 55 tok/s | 68 tok/s | 95 tok/s | H100 is cheapest per token |
| 70B inference (batch=32) | 450 tok/s | 580 tok/s | 750 tok/s | B100 if throughput matters |
| 70B full fine-tune | 24 hrs | 18 hrs | 12 hrs | B100 for faster experiments |
| 7B inference high volume | 8,100 tok/s | 10,200 tok/s | 14,500 tok/s | B100 at scale — savings add up |
For most teams, H100 is the sweet spot. H200 gives 24% more throughput for 20% more cost — marginally better value. B100 is faster but at 40% premium over H100. Upgrade only if you're bottlenecked on throughput.
Photo by Andrey Matveev on Pexels
Cost per Million Tokens: The Real Metric
70B Model Serving Cost
| Provider | GPU | $/hr | Tokens/Sec (batch=1) | Tokens/Sec (optimum batch) | $/1M tokens |
|---|---|---|---|---|---|
| io.net (spot) | H100 | $1.20 | 55 | 450 (batch=32) | $0.0007 |
| Vast.ai (bid) | H100 | $1.50 | 55 | 450 | $0.0009 |
| RunPod (spot) | H100 | $1.85 | 55 | 450 | $0.0011 |
| Akash | H100 | $2.60 | 55 | 450 | $0.0016 |
| RunPod (on-demand) | H100 | $2.29 | 55 | 450 | $0.0014 |
| Lambda Labs | H100 | $2.49 | 55 | 450 | $0.0015 |
| AWS | H100 | $4.73 | 55 | 450 | $0.0029 |
8B Model Serving Cost
| Provider | $/hr | Tokens/Sec | $/1M tokens |
|---|---|---|---|
| io.net (spot) — RTX 4090 | $0.22 | 85 | $0.0007 |
| Vast.ai (bid) — RTX 4090 | $0.25 | 85 | $0.0008 |
| RunPod — RTX 4090 | $0.51 | 85 | $0.0017 |
| Lambda Labs — H100 | $2.49 | 305 | $0.0023 |
| AWS — L4 | $1.59 | 80 | $0.0055 |
Cheapest per-token inference: io.net spot H100 ($0.0007/1M tokens) for 70B. For small models, consumer GPUs on DePIN networks are dramatically cheaper per token than H100s.
Training Cost: H100 vs Consumer GPU Rental
Fine-Tuning Cost Comparison (Llama 3.1 70B, QLoRA, 1 epoch, 1000 samples)
| GPU | Provider | $/hr | Training Time | Total Cost |
|---|---|---|---|---|
| H100 | Lambda Labs | $2.49 | 0.5 hrs | $1.25 |
| H100 | RunPod (spot) | $1.85 | 0.5 hrs | $0.93 |
| H100 | AWS | $4.73 | 0.5 hrs | $2.37 |
| RTX 5090 | RunPod | $0.69 | 1.5 hrs | $1.04 |
| RTX 4090 | Vast.ai | $0.35 | 2.5 hrs | $0.88 |
| RTX 4090 | RunPod | $0.51 | 2.5 hrs | $1.28 |
| 2x RTX 3090 | Vast.ai | $0.50 | 2.0 hrs | $1.00 |
Full Training Cost (70B, LoRA, 1 epoch, 10K samples)
| GPU | Provider | $/hr | Training Time | Total Cost |
|---|---|---|---|---|
| 8x H100 | AWS | $151.00 | 4 hrs | $604 |
| 8x H100 | Lambda Labs | $19.92 | 4 hrs | $80 |
| 8x H100 | RunPod | $17.84 | 4 hrs | $71 |
| 8x H100 | Vast.ai | $13.50 | 4 hrs | $54 |
Hidden Costs: Storage, Transfer, and Engineering Time
Storage Costs
| Provider | Storage Included | Additional Storage | Block Storage |
|---|---|---|---|
| AWS | 8GB (EBS boot) | $0.08/GB-month (gp3) | $0.08/GB-month |
| Lambda Labs | 200GB NVMe | $0.10/GB-month | Included |
| RunPod | 5GB (template) | $0.07/GB-month | $7/TB-month |
| Vast.ai | 50GB (instance) | $0.05/GB-month | Variable |
| io.net | 10GB | $0.10/GB-month | N/A |
Data Transfer (Egress)
| Provider | Egress Cost | Free Tier |
|---|---|---|
| AWS | $0.05-0.09/GB | 100GB/month |
| Lambda Labs | $0.05/GB | 500GB/month |
| RunPod | $0.01/GB | 1TB/month |
| Vast.ai | $0.01/GB | 200GB/month |
| io.net | $0.02/GB | 100GB/month |
Real Monthly Cost Example
Workload: Batch inference, 70B, 50M tokens/day, H100
| Provider | Compute (30 days) | Storage (1TB) | Egress (500GB) | Total | Effective $/hr |
|---|---|---|---|---|---|
| AWS | $3,405 | $80 | $35 | $3,520 | $4.88 |
| Lambda Labs | $1,793 | $100 | $25 | $1,918 | $2.66 |
| RunPod | $1,649 | $70 | $5 | $1,724 | $2.39 |
| Vast.ai | $1,361 | $50 | $5 | $1,416 | $1.97 |
| io.net (on-demand) | $1,361 | $100 | $10 | $1,471 | $2.04 |
| io.net (spot) | $864 | $100 | $10 | $974 | $1.35 |
Provider Rankings by Use Case
Best for Production Inference (SLA Required)
1. Lambda Labs ★★★★★ $2.49/hr, 99.5% uptime, good support
2. RunPod ★★★★☆ $2.29/hr, 98% uptime, community support
3. AWS ★★★☆☆ $4.73/hr, 99.9% uptime, enterprise pricing
Best for Training
1. RunPod ★★★★★ $2.29/hr (on-demand), $1.85/hr (spot)
2. Lambda Labs ★★★★★ $2.49/hr, consistent performance
3. Vast.ai ★★★★☆ $1.89/hr, but variable quality
Best for Batch Inference (Cheapest)
1. io.net spot ★★★★★ $1.20/hr — can't beat the price
2. Vast.ai bid ★★★★☆ $1.50/hr — more reliable than io.net
3. Akash ★★★★☆ $2.20-3.00/hr — most decentralized
Best for Experimentation
1. Vast.ai ★★★★★ $1.50-1.89/hr, widest selection
2. RunPod ★★★★★ $1.85-2.29/hr, best UX
3. io.net ★★★★☆ $1.20-2.59/hr, cheapest spot
Related Reads
- RunPod vs Vast.ai vs Lambda Labs: Best GPU Cloud for LLMs
- Akash Network H100 Pricing: How Cheap vs AWS?
- Self-Host Cost Per Million Tokens: Calculator & Provider Data
Optimizing H100 Workloads for Cost Efficiency
Beyond choosing the right provider, small adjustments to workload configuration can yield outsized cost savings. For inference, batch size is the most critical lever: increasing batch size from 1 to 32 on a 70B model boosts throughput by 8x (from 55 to 450 tokens/sec) while only doubling GPU memory usage. This reduces cost per million tokens by 75%—transforming a $0.0029/1M token AWS deployment into a $0.0007/1M token operation on io.net. However, larger batches introduce latency; for real-time applications, benchmark your model’s latency vs. throughput tradeoff at batch sizes of 1, 2, 4, 8, 16, and 32 to find the optimal balance.
For training, mixed precision (FP8) and gradient checkpointing can cut costs by 30-50% without sacrificing model quality. FP8 reduces memory usage by 50% compared to FP16, enabling larger batch sizes or fitting larger models on a single H100. Gradient checkpointing trades compute for memory, reducing VRAM usage by 30-40% at the cost of a 20-30% increase in training time. When combined, these techniques can reduce the cost of a 70B fine-tuning run from $80 to $40 on RunPod. Providers like Lambda Labs and RunPod offer pre-configured templates with these optimizations enabled, while AWS requires manual setup via custom AMIs or Docker containers.
Network Topology and Multi-GPU Scaling
The performance gap between single H100s and 8x H100 instances isn’t linear due to networking bottlenecks. AWS’s p5.48xlarge instances include NVLink and 3.2Tbps networking, enabling near-linear scaling for distributed training (e.g., 7.8x speedup for 8x GPUs). In contrast, providers like Lambda Labs and RunPod use PCIe 4.0 or 5.0 for multi-GPU communication, which caps scaling efficiency at 6-7x for 8x GPUs. For workloads like full fine-tuning or large-scale inference, this difference can add 20-30% to runtime costs. Benchmark your workload’s scaling efficiency by comparing single-GPU performance to 2x, 4x, and 8x configurations—if scaling efficiency drops below 80%, consider splitting workloads across multiple single-GPU instances instead of using a single 8x instance.
For inference, multi-GPU setups are rarely cost-effective unless you’re serving at scale. A single H100 can handle ~500 tokens/sec for a 70B model at batch=32, which translates to ~130M tokens/day—enough for most production workloads. If you need higher throughput, deploy multiple single-GPU instances behind a load balancer (e.g., Nginx or Traefik) rather than using an 8x H100 instance. This approach improves fault tolerance and reduces costs by 30-50% compared to a single 8x instance, as you can scale horizontally with spot instances.
Provider-Specific Quirks and Workarounds
Each H100 provider has unique limitations that can impact cost and performance if not accounted for:
- AWS: EBS boot volumes are slow (100-200 MB/s); use instance storage for datasets or cache. Egress costs ($0.05-0.09/GB) can exceed compute costs for large datasets—compress data or use AWS’s free tier (100GB/month) for transfers.
- Lambda Labs: No spot instances, but their 1-month reserved pricing ($1,793/mo) is effectively a 2% discount. Their NVMe storage is included up to 200GB, making them ideal for storage-heavy workloads like video processing.
- RunPod: Spot instances have a 5-minute preemption warning, but their API doesn’t expose this—use their webhook feature to trigger checkpointing. Their $0.01/GB egress is the cheapest among providers, but transfers are throttled to 1Gbps.
- Vast.ai: Provider quality varies wildly; filter for "H100 PCIe 4.0" or "NVLink" in the search bar to avoid underperforming instances. Their bid system can save 20-30% over on-demand, but lowball bids may get preempted quickly.
- io.net: Spot instances are the cheapest ($1.20/hr) but have no preemption warning—implement a heartbeat system to detect failures. Their storage is ephemeral, so use external storage (e.g., S3 or Backblaze B2) for datasets.
For production workloads, test each provider’s networking performance with tools like ib_write_bw (for InfiniBand) or iperf3 (for TCP). AWS and Azure offer 100Gbps+ networking, while alternative providers typically cap at 25-50Gbps. If your workload involves frequent data transfers (e.g., distributed training), this can add 10-20% to runtime costs. For storage-bound workloads, benchmark disk I/O with fio—Lambda Labs’ NVMe storage delivers 3-5GB/s, while AWS’s gp3 EBS tops out at 1GB/s.
Key Takeaways
- For most AI workloads, Lambda Labs ($2.49/hr) or RunPod ($2.29/hr) offer the best balance of cost, reliability, and support—avoid AWS ($4.73/hr) unless you need enterprise SLAs or tight cloud integration.
- Spot instances (e.g., io.net at $1.20/hr) can cut costs by 75% but carry 30-50% preemption risk; use them only for fault-tolerant workloads like batch inference or hyperparameter sweeps with checkpointing every 10-15 minutes.
- H100 remains the cost-performance sweet spot in 2026: H200 offers 24% more throughput for a 20% price premium, while B100’s 40% premium is only justified for throughput-bottlenecked workloads like high-volume 70B inference or full fine-tuning.
- For 70B inference, io.net spot H100s deliver the lowest cost per million tokens ($0.0007), but consumer GPUs (e.g., RTX 4090 on Vast.ai at $0.0008/1M tokens) are dramatically cheaper for smaller models like 8B.
- Hidden costs add up: AWS charges $0.08/GB-month for storage and $0.05-0.09/GB egress, while providers like RunPod ($0.01/GB egress) or Lambda Labs (500GB free egress) can reduce total monthly costs by 20-40%.
- For full fine-tuning (LoRA, 10K samples), expect $50-80 on RunPod/Vast.ai vs $600 on AWS; spot instances can further reduce training costs by 15-30% if checkpointing is implemented.
Frequently Asked Questions
Where is the cheapest place to rent an H100?
io.net spot at $1.20/hr is the cheapest on paper. However, with 30-50% preemption rates and 10-20% job failure rates, the effective cost is higher. For practical cheapest H100: Vast.ai bid at $1.50/hr with good provider selection. For reliable cheap: RunPod spot at $1.85/hr.
Is AWS H100 worth the premium?
For production serving with SLA requirements: yes. AWS offers 99.9%+ uptime, 24/7 enterprise support, and seamless integration with other AWS services. For batch processing and training: no — Lambda Labs or RunPod provide 95% of the reliability at 50% of the cost.
What's the difference between H100 and H200 for cloud rental?
H200 has 141 GB VRAM (vs 80 GB on H100) and slightly faster HBM3e memory. For inference, H200 yields ~24% more throughput. For training, H200 fits larger models without offloading. H200 costs ~20% more per hour, making it marginally better value.
How much does it cost to train Llama 3.1 70B on cloud H100s?
A full continued pre-training run ($500K-2M) is only for large AI labs. For fine-tuning (LoRA, 10K samples, 1 epoch), expect $50-80 on Vast.ai/RunPod or $600 on AWS. For full fine-tuning (all parameters, small dataset), expect $500-2,000 on alternative providers vs $5,000+ on AWS.
Can I save money by using spot instances for training?
Yes — if you implement checkpointing (every 10-15 minutes). Spot savings of 20-40% are real, and even with occasional preemption, the effective cost is still 15-30% less than on-demand. Without checkpointing, spot is risky for training.

Comments
Sign in to join the conversation
No comments yet. Be the first to share your thoughts!