Stack Cost AI

Foundation Modelprofessional stack.

For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.

GPU Supercluster / Training Compute (A100/H100/B200)

Pricing & free-tier limitsFree: Colab T4 12GB 5h/day limited, Kaggle 30h P100/wk, Lightning 22h T4/mo. Cheaper: Vast.ai 4090 $0.69/hr, 2xH100 $3.98/hr; RunPod Secure H100 $2.89/hr, Community $1.99/hr. Paid: CoreWeave H100 $3.96-4.76/hr per GPU ($31.68-38.08/8x), AWS p5.48xlarge $98.32/hr ($12.29/GPU), Lambda 8xH100 $32.08/hr ($4.01/GPU), Azure ND96 $99.76/hr. Spot 40-60% off. B200 ~$7-9/hr/GPU in 2026. Overage: inter-node egress $0.01/GB.

Distributed Training Framework

Pricing & free-tier limitsFree: PyTorch DDP, DeepSpeed OSS, Accelerate, Colossal-AI free self-host unlimited. Cheaper: Self-host on RunPod/Vast at $0.69/hr infra, no license. Paid: Anyscale free 20 core-hrs/mo then $0.55/vCPU/hr + $2000/mo platform. Databricks Mosaic training overhead $0.00065/token (~$9750 for 15T tokens). Deterministic AI $2k/mo.

Data Lake / Object Storage (Petabyte Scale)

Pricing & free-tier limitsFree: MinIO OSS unlimited self-host, R2 Free 10GB, 10M Class A ops, 10M Class B, zero egress. S3 Free Tier 5GB 12mo. Cheaper: B2 $6/TB/mo storage, free downloads 3x stored/day, $0.01/GB over. Wasabi $6.99/TB zero egress, min 90d retention. Paid: S3 $23/TB/mo + $0.005/1k PUT + $0.09/GB egress. GCS $20/TB + egress $0.12/GB. 100PB dataset = $2.3M/mo S3. R2 $15/TB + $4.50/M operations.

Data Pipeline / ETL for Training Data (Spark, Ray Data)

Pricing & free-tier limitsFree: Spark OSS self-host on K8s free, Ray Data OSS free, DuckDB free single node. Cheaper: Modal free $30 credit/mo then CPU $0.000046/sec, Coiled free 5k OCPU-hrs. Dagster OSS free. Paid: Databricks $0.40/DBU + EC2 = $0.75-1.50/hr per node DBU+EC2 combined, Photon $0.60/DBU. AWS EMR $0.27/hr per m5.xlarge + EC2 $0.192/hr = $0.462/hr. Glue free 1M objects + 1 DPU-hr/mo then $0.44/DPU-hr.

Checkpoint Storage (High-speed NVMe + S3)

Pricing & free-tier limitsFree: Instance NVMe 8x3.84TB free with p4d.24xlarge/p5. JuiceFS OSS free + S3 cost only. Cheaper: RunPod Volume $0.10/GB/mo = $100/TB/mo, Vast included. JuiceFS Cloud Free 10GB then $0.02/GB metadata + S3. Paid: FSx Lustre $0.14/GB/mo = $140/TB + throughput $3.90/MBps-month. EFS $300/TB/mo. Weka $0.25/GB = $250/TB/mo + IOPs. S3 Express $0.16/GB = $160/TB/mo + $0.0025/1k GET. 70B model ckpt 140GB x500 =70TB ~$7k-21k/mo.

Experiment Tracking / MLOps Platform (W&B, MLflow)

Pricing & free-tier limitsFree: W&B Free 100GB storage, 1 team, personal unlimited experiments, 1h system metrics. MLflow OSS unlimited self-host free. Cheaper: Neptune Free 100GB, 10GB monthly tracking. Comet Free 5000h/mo + 2GB. Paid: W&B Team $50/u/mo 5-seat min $250/mo 250GB storage. Enterprise $100/u/mo 1TB + SSO. Overage: $0.50/GB storage, $0.10/1k API calls. MLflow Managed free w/ Databricks.

Model Registry & Versioning

Pricing & free-tier limitsFree: HF Hub free unlimited public models, 10 private repos 100GB bandwidth. MLflow OSS free unlimited. Cheaper: HF Pro $9/mo 1 private org, 1TB bandwidth, 300GB storage. DVC + S3 $23/TB. Paid: HF Enterprise Cloud $20/user/mo + infra. JFrog Pro $98/mo 5 users 25GB. Databricks UC free with compute. Model size 70B 140GB x20 versions 2.8TB ~ $64/mo S3.

Tokenization / Data Processing Service

Pricing & free-tier limitsFree: Tokenizers OSS 1M+ tok/sec per core, unlimited free, tiktoken OSS free, SentencePiece OSS free local. Cheaper: Modal CPU $0.000046/sec => tokenize 1B tokens ~$0.50. Self-host free. Paid: Cohere $0.10/1M tokens tokenization API, 15T tokens = $1500 API cost, better OSS $2000 GPU compute 500 H100-hrs Ray Data $4/hr = $2000 one-time for 15T.

Evaluation Harness / Benchmarking (HELM, LM-Eval)

Pricing & free-tier limitsFree: LM-Eval Harness free self-host unlimited benchmarks, HELM free. HF Evaluate OSS free. Cheaper: HF Inference free 30k req/mo then $0.06/1k req. Modal eval $2.10/hr H100. Paid: Scale AI eval $2-5k per full run (100 tasks x 70B). LangSmith Plus $59/mo 10k eval traces, $0.50/1k over. Anyscale evaluation endpoints: Llama 70B $0.25/1M input, $1/1M output. Full benchmark suite 500M tokens ~ $200-500.

Container Orchestration (Kubernetes for Training)

Pricing & free-tier limitsFree: k3s OSS free self-host, Kind/MicroK8s free local unlimited. Cheaper: EKS free credit $144? Actually GKE free 1 zonal cluster free, Autopilot $0.10/vCPU/hr. EKS $0.10/hr per cluster = $72/mo. Paid: EKS $144/mo with extended support + $0.10/hr. Karpenter free. 1000 GPU cluster control plane $144/mo + 1000 x EC2 $4/hr = $4k/hr compute. Anyscale $2000/mo platform + compute.

Artifact Registry / Container Registry

Pricing & free-tier limitsFree: GHCR 500MB private storage + 2k mins Actions free, unlimited public. ECR Public free pulls. Docker Hub free 1 private repo 200 pulls/6h. Cheaper: ECR private $0.10/GB/mo storage, data transfer free within region to EC2. GitLab free 10GB registry. Paid: Docker Hub Pro $9/mo 5 private repos unlimited pulls. JFrog Pro $98/mo 5 users 25GB + $0.50/GB over. Artifactory storage: training images 50GB x20 =1TB ~$100/mo ECR.

Vector Database for Evaluation / RAG Eval

Pricing & free-tier limitsFree: Pinecone Free 100k vectors up to 2GB 1 index, 100 QPS. Qdrant OSS free unlimited self-host. Chroma OSS free. Cheaper: Qdrant Cloud Free 1GB RAM 500k vectors, $25/mo 1GB cluster 1M vectors. Pinecone Starter $0 with 100k free then $70/mo. Paid: Pinecone Serverless $0.33/GB/mo storage + $0.08/1M reads + $2/1M writes. Qdrant Cloud $75/mo 4GB RAM 4M vectors. Over 10M vectors 768d = 30GB RAM = $300/mo Qdrant Pro.

Inference Preview / Base Model Serving (for eval)

Pricing & free-tier limitsFree: vLLM OSS self-host free + GPU $1.5-4/hr. Ollama local free. HF Inference free 30k req/mo. Cheaper: Modal free $30 credit/mo then $2.10/hr H100 + $0.000046/sec GPU idle, $0/min scale-to-zero. RunPod Serverless free idle, $0.0002/sec H100 => 70B inference $0.72/hr active. Paid: Anyscale $10 free credits then Llama 70B $0.25/1M input $1/1M output. Fireworks $0.20/1M input. Replicate $5 free then $0.0014/sec H100. Preview cluster for 70B vLLM 2xH100 $8/hr.

Distributed File System (Lustre, JuiceFS, Alluxio)

Pricing & free-tier limitsFree: JuiceFS OSS free + S3 cost, CephFS OSS free self-host, Lustre community free. Cheaper: JuiceFS Cloud Free 10GB then $10/mo 100GB metadata + S3 $23/TB. Alluxio OSS free self-host $100/mo VM. Paid: FSx Lustre Scratch 1.2TB min $168/mo @ $0.14/GB, Persistent $0.24/GB = $240/TB/mo + $3.90/MBps-month throughput. Weka min $50k/mo PB-scale $0.35/GB = $350/TB.

Queuing / Job Scheduler for Training

Pricing & free-tier limitsFree: Slurm OSS free unlimited nodes, Kueue free, Volcano free, Airflow OSS free 1k DAGs. Ray Jobs free. Cheaper: AWS Batch free scheduler, you pay EC2 $1.5-12/GPU/hr only. GCP Batch same. Paid: LSF $5000/node/yr Enterprise, PBS Pro $10k/cluster $1200/node. AWS Batch Enterprise Support $15k/mo. Queue for 1000 GPUs free with Slurm.

Dataset Versioning / Data Catalog (DVC, LakeFS)

Pricing & free-tier limitsFree: DVC OSS free S3 backend cost only $23/TB. LakeFS OSS free self-host $100/mo VM. HF Datasets free private datasets with Pro plan, public unlimited. Cheaper: DVC Studio Free 20GB remote cache, 500 API calls/day. LakeFS Cloud Free 3GB branches. Paid: DVC Studio Team $20/u/mo 3 seats min $60/mo + cache $0.25/GB = $250/TB/mo. LakeFS Cloud $50/mo 100GB + $0.20/GB over. 10PB dataset versioning metadata ~1TB ~$250/mo.

Log Management for Training Logs

Pricing & free-tier limitsFree: Loki OSS free unlimited self-host + S3 $23/TB. OpenSearch OSS free. Grafana Cloud Free 50GB logs, 50GB traces, 10k series metrics. Cheaper: Axiom Free 500GB/mo logs, $25/mo 10GB. Better Stack Free 3GB logs 3d retention. Paid: Grafana Pro $8/user/mo + $0.50/GB logs over 50GB. Datadog Logs $0.10/GB ingest (100GB = $10) + $1.70/1M scanned, retention $0.02/GB/mo. Splunk $0.80/GB/day = $24/GB/mo. Training generates 10TB/mo logs = $5k Datadog, $500 Loki.

Secrets Management

Pricing & free-tier limitsFree: Vault OSS free unlimited secrets self-host + VM $20/mo. Infisical OSS free self-host. Sealed Secrets OSS free. Cheaper: Infisical Cloud Free 1000 secrets 2 envs, Doppler Free 100 secrets 5 users. Paid: AWS Secrets $0.40/secret/mo + $0.05/10k API calls. Example 500 API keys $200/mo + $50 calls. Vault Cloud $0.03/hr small = $22/mo + $0.10/secret-hour large. Azure KeyVault $0.03/10k ops + $0.30/cert/mo.

Security Scanning & Data Filtering (PII, Toxic, CSAM)

Pricing & free-tier limitsFree: Presidio free self-host CPU $0.05/hr, NeMo Curator free GPU 1 H100 $4/hr for PB filtering, Llama Guard 8B free OSS $0.20/1M tokens self-host. ClamAV free. Cheaper: AWS Macie Free trial 30d 1TB free. Paid: Macie $0.10/GB first TB = $100/TB, then $0.05/GB = $50/TB. 10PB training data scan = $500k one-time. Datadog Sensitive $0.25/GB = $2500/TB. Scale to 15T tokens ~ 60TB text = $3-15k scan OSS + compute.

GPU Monitoring & Observability (DCGM, Prometheus)

Pricing & free-tier limitsFree: DCGM + DCGM Exporter free self-host, Prometheus free unlimited, Grafana OSS free. Cheaper: Grafana Cloud Free 10k series metrics, 50GB logs, 50GB traces. W&B free system metrics 100GB. Paid: Datadog Pro $23/host/mo + container $0.002/hr. GPU Monitoring $0.10/hr per GPU. 1000 H100 cluster => GPU $0.10x1000x730 = $73k/mo + hosts 125x$23 = $2875/mo. Chronosphere $750/mo cluster + $0.30/1k metrics.

Cost Tracking / FinOps for GPU (OpenCost, Kubecost)

Pricing & free-tier limitsFree: OpenCost OSS free unlimited clusters self-host. Kubecost Free 1 cluster 15d retention free. Cheaper: Infracost Free up to 1000 runs/mo, $25/mo 5000 runs. Paid: Kubecost Enterprise $999/mo per cluster unlimited retention, Kubecost Cloud Pro $499/mo 5 clusters. CloudHealth $15k/mo for $2M spend enterprise. Saves 40-60% via spot tracking for GPU $3/hr saved x1000 GPUs = $2M/mo savings.

High-Speed Interconnect / Networking (EFA, InfiniBand, RoCE)

Pricing & free-tier limitsFree: AWS EFA driver + fabric free with instance, IB included in GPU hourly price at all neoclouds. GPUDirect RDMA free. Cheaper: Same-AZ traffic free, cross-AZ $0.01/GB. CoreWeave/Lambda fabric inclusive. Paid: Equinix Fabric Fabric port $500/mo 10G + $2k/mo interconnect + $0.05/GB. AWS inter-AZ $0.01/GB = $10/TB. All-reduce for 70B training 15T tokens ~500TB internal traffic = $5k inter-AZ if misplaced, $0 if same-AZ placement group.

Cache / In-Memory Store for Data Loading

Pricing & free-tier limitsFree: Redis OSS free unlimited self-host on VM $50/mo 32GB. Dragonfly OSS free 25x faster. Cheaper: Upstash Free 10k cmds/day, 256MB, $0.20/100k cmds over. Dragonfly Cloud $10/mo 500MB. Paid: ElastiCache Serverless $0.125/GB/hr = $90/GB/mo RAM + $0.20/M ECPUs + $0.125/GB/hr backup. Example 500GB cache for dataloader = $45k/mo serverless, $5k/mo self-host. Redis Enterprise Cloud $35/mo 100MB + $0.50/100k ops.

Metadata Catalog / Data Lineage

Pricing & free-tier limitsFree: DataHub OSS free self-host VM $100/mo 4 vCPU 16GB, OpenLineage free, Marquez free. Cheaper: DataHub Cloud free trial 30d then Team $500/mo 10 users. Atlan free trial then growth $50k/yr. Paid: Acryl Team $500/mo 10 users 1TB metadata, Enterprise $2000-5000/mo 100 users. Atlan $40k/yr starter, $80k/yr growth. Databricks Unity Catalog free with DBR, storage $23/TB lineage logs.

CI/CD for Training Jobs (GitHub Actions, Argo)

Pricing & free-tier limitsFree: GH Actions Free 2000 min/mo Linux, 500MB storage, self-hosted GPU runners free but EC2 $4/hr. Argo OSS free unlimited workflows. Cheaper: GitLab Free 400 compute min/mo + self-host. Buildkite Free 1 agent. Paid: GH Actions Team $4/u/mo + Linux $0.008/min ($0.48/hr), GPU self-host still EC2 cost + time. Large training CI 100 jobs/day 60min H100 self-host: $480/day compute + $20/mo platform. AWS CodePipeline $1/pipeline/mo + $0.002/action over 100.

Model Optimization & Compilation (TensorRT-LLM, torch.compile)

Pricing & free-tier limitsFree: TensorRT-LLM OSS free + H100 $4/hr, torch.compile free unlimited, ONNX Runtime free, llama.cpp 4-bit quant free local. Cheaper: Modal $30 free then H100 $2.10/hr optimization experiments 100hr trial $210. Paid: Deci $0 trial then Developer $1500/mo 3 models, Enterprise $20k/mo unlimited. OctoML Starter $500/mo 2 models, Scale $15k/mo. Optimization reduces inference preview cost 40-70%: 70B FP16 $4/hr -> INT4 $1.2/hr.

Backup & Disaster Recovery for Checkpoints

Pricing & free-tier limitsFree: Restic OSS free + S3 $23/TB/mo, Velero free + S3 storage only. MinIO OSS free. Cheaper: B2 $6/TB/mo = $6000/PB/mo backup replica, Wasabi $6.99/TB zero egress. Paid: AWS Backup $0.10/GB/mo vault storage = $100/TB/mo + CRR transfer $0.02/GB = $20/TB transfer. 1PB checkpoint backups cross-region = $20k transfer + $100k/mo storage S3 $23k = $43k-123k/mo. Glacier Deep Archive $1/TB/mo long-term = $1k/PB/mo for compliance.

Rate Limiting / API Gateway for Eval API

Pricing & free-tier limitsFree: Kong OSS free self-host VM $50/mo unlimited req, Cloudflare Free 100k req/day, 1000 Gateway rules, AWS Free Tier 1M calls/mo 12mo. Cheaper: Cloudflare Workers Paid $5/mo + $0.50/M requests includes 10M = $5/M after. AWS API GW $3.50/M req + $1/300M cache. Paid: Kong Konnect Team $250/mo includes 7.5M requests 3 services. AWS $3.50/M 100M eval requests = $350/mo + data 1TB $90. Cloudflare Enterprise Gateway $5k/mo 100M req.

Identity & Access Management (IAM for GPU Clusters)

Pricing & free-tier limitsFree: AWS IAM free no cost, Teleport Community free unlimited 100 resources self-host VM $50/mo. Keycloak OSS free. Cheaper: Authentik OSS free self-host unlimited users, $50/mo infra. Paid: Teleport Enterprise $15/user/mo + $20/server/mo = 500 users + 1000 nodes = $27.5k/mo list. Okta Free 100 MAU then $2/user/mo + addons Workplace $8/u/mo. Auth0 Free 25k MAU then $0.07/MAU over = 10k extra = $700/mo. Cluster 500 eng.

Documentation / Knowledge Base for Infra

Pricing & free-tier limitsFree: MkDocs free + GitHub Pages free hosting unlimited bandwidth, Docusaurus OSS free + Vercel free hobby 100GB bandwidth. Cheaper: GitBook Free 1 private space unlimited public projects, ReadTheDocs free 1GB storage. Paid: GitBook Pro $40/u/mo 10GB bandwidth $0.50/GB over. Notion Plus $10/u/mo 1000 blocks? Actually unlimited blocks $10/u/mo min $80. Confluence Standard $6.05/u x100 = $605/mo + storage 250GB free. Mintlify Pro $120/mo 200k page views + $0.50/1k over.

Options and prices come straight from our research sheets for a professional foundation model project. Prices are estimates and change often — always confirm on the provider's page before committing.

How this professional foundation model checklist works.

Each requirement below is something a professional foundation model build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.

Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.