SEED STAGE ($0-5M RAISED)
Seed-stage teams of 2-10 engineers should use cloud GPU credits ($100-350K via startup programs), spot instances (60-80% discount), and GPU marketplaces. Monthly GPU budget: $2,000-15,000. Total seed infra spend over 18 months: $36-270K.
Recommended: 1-4 A100/H100 via spot/marketplace. Zero reserved capacity - committing locks up 20-40% of seed round before PMF validation.
SERIES A: SCALING ($5-20M RAISED)
50% reserved (4-16 H100 at $2.60-3.10/hr on 1yr), 50% spot/marketplace. Monthly GPU: $15-80K. Annual infra: $300K-1.5M. Platform build-out: $100-200K one-time plus $15-30K/month.
GPU budget cap: 20% of monthly OpEx. At $400K/month burn typical for $10M raise over 24 months, GPU cap is $80K/month.
SERIES B: PRODUCTION AT SCALE ($20-100M RAISED)
Inference becomes 40-60% of GPU spend. Typical: 32-64 H100 for inference, 32-64 for training. Monthly GPU: $80-300K. Strategy: 60% reserved, 10% spot, 30% flexible.
Deploy speculative decoding (30-50% per-token cost reduction), GPU-aware autoscaling. ML infra team of 3-5 people. Total annual infra: $1.5-5M.
SERIES C+: ENTERPRISE OPTIMIZATION
Monthly GPU spend exceeds $300K up to $2M+. Negotiate 2-3 year reservations at 30-45% discount. Direct contracts with providers achieve $1.80-2.20/hr for H100.
Combined optimization (auto-scheduling, right-sizing, quantization, bin-packing) achieves 30-50% reduction from baseline on-demand costs. At $1M/month, 5-10% optimization saves $50-100K/month.