Stack Cost AI

Foundation Modelsmall scale stack.

For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.

GPU Supercluster / Training Compute (A100/H100/B200)

Pricing & free-tier limitsFree: Colab T4 12GB 5h/day limited, Kaggle 30h P100/wk, Lightning 22h T4/mo. Cheaper: Vast.ai 4090 $0.69/hr, 2xH100 $3.98/hr; RunPod Secure H100 $2.89/hr, Community $1.99/hr. Paid: CoreWeave H100 $3.96-4.76/hr per GPU ($31.68-38.08/8x), AWS p5.48xlarge $98.32/hr ($12.29/GPU), Lambda 8xH100 $32.08/hr ($4.01/GPU), Azure ND96 $99.76/hr. Spot 40-60% off. B200 ~$7-9/hr/GPU in 2026. Overage: inter-node egress $0.01/GB.

Distributed Training Framework

Pricing & free-tier limitsFree: PyTorch DDP, DeepSpeed OSS, Accelerate, Colossal-AI free self-host unlimited. Cheaper: Self-host on RunPod/Vast at $0.69/hr infra, no license. Paid: Anyscale free 20 core-hrs/mo then $0.55/vCPU/hr + $2000/mo platform. Databricks Mosaic training overhead $0.00065/token (~$9750 for 15T tokens). Deterministic AI $2k/mo.

Data Lake / Object Storage (Petabyte Scale)

Pricing & free-tier limitsFree: MinIO OSS unlimited self-host, R2 Free 10GB, 10M Class A ops, 10M Class B, zero egress. S3 Free Tier 5GB 12mo. Cheaper: B2 $6/TB/mo storage, free downloads 3x stored/day, $0.01/GB over. Wasabi $6.99/TB zero egress, min 90d retention. Paid: S3 $23/TB/mo + $0.005/1k PUT + $0.09/GB egress. GCS $20/TB + egress $0.12/GB. 100PB dataset = $2.3M/mo S3. R2 $15/TB + $4.50/M operations.

Data Pipeline / ETL for Training Data (Spark, Ray Data)

Pricing & free-tier limitsFree: Spark OSS self-host on K8s free, Ray Data OSS free, DuckDB free single node. Cheaper: Modal free $30 credit/mo then CPU $0.000046/sec, Coiled free 5k OCPU-hrs. Dagster OSS free. Paid: Databricks $0.40/DBU + EC2 = $0.75-1.50/hr per node DBU+EC2 combined, Photon $0.60/DBU. AWS EMR $0.27/hr per m5.xlarge + EC2 $0.192/hr = $0.462/hr. Glue free 1M objects + 1 DPU-hr/mo then $0.44/DPU-hr.

Checkpoint Storage (High-speed NVMe + S3)

Pricing & free-tier limitsFree: Instance NVMe 8x3.84TB free with p4d.24xlarge/p5. JuiceFS OSS free + S3 cost only. Cheaper: RunPod Volume $0.10/GB/mo = $100/TB/mo, Vast included. JuiceFS Cloud Free 10GB then $0.02/GB metadata + S3. Paid: FSx Lustre $0.14/GB/mo = $140/TB + throughput $3.90/MBps-month. EFS $300/TB/mo. Weka $0.25/GB = $250/TB/mo + IOPs. S3 Express $0.16/GB = $160/TB/mo + $0.0025/1k GET. 70B model ckpt 140GB x500 =70TB ~$7k-21k/mo.

Experiment Tracking / MLOps Platform (W&B, MLflow)

Pricing & free-tier limitsFree: W&B Free 100GB storage, 1 team, personal unlimited experiments, 1h system metrics. MLflow OSS unlimited self-host free. Cheaper: Neptune Free 100GB, 10GB monthly tracking. Comet Free 5000h/mo + 2GB. Paid: W&B Team $50/u/mo 5-seat min $250/mo 250GB storage. Enterprise $100/u/mo 1TB + SSO. Overage: $0.50/GB storage, $0.10/1k API calls. MLflow Managed free w/ Databricks.

Model Registry & Versioning

Pricing & free-tier limitsFree: HF Hub free unlimited public models, 10 private repos 100GB bandwidth. MLflow OSS free unlimited. Cheaper: HF Pro $9/mo 1 private org, 1TB bandwidth, 300GB storage. DVC + S3 $23/TB. Paid: HF Enterprise Cloud $20/user/mo + infra. JFrog Pro $98/mo 5 users 25GB. Databricks UC free with compute. Model size 70B 140GB x20 versions 2.8TB ~ $64/mo S3.

Tokenization / Data Processing Service

Pricing & free-tier limitsFree: Tokenizers OSS 1M+ tok/sec per core, unlimited free, tiktoken OSS free, SentencePiece OSS free local. Cheaper: Modal CPU $0.000046/sec => tokenize 1B tokens ~$0.50. Self-host free. Paid: Cohere $0.10/1M tokens tokenization API, 15T tokens = $1500 API cost, better OSS $2000 GPU compute 500 H100-hrs Ray Data $4/hr = $2000 one-time for 15T.

Evaluation Harness / Benchmarking (HELM, LM-Eval)

Pricing & free-tier limitsFree: LM-Eval Harness free self-host unlimited benchmarks, HELM free. HF Evaluate OSS free. Cheaper: HF Inference free 30k req/mo then $0.06/1k req. Modal eval $2.10/hr H100. Paid: Scale AI eval $2-5k per full run (100 tasks x 70B). LangSmith Plus $59/mo 10k eval traces, $0.50/1k over. Anyscale evaluation endpoints: Llama 70B $0.25/1M input, $1/1M output. Full benchmark suite 500M tokens ~ $200-500.

Container Orchestration (Kubernetes for Training)

Pricing & free-tier limitsFree: k3s OSS free self-host, Kind/MicroK8s free local unlimited. Cheaper: EKS free credit $144? Actually GKE free 1 zonal cluster free, Autopilot $0.10/vCPU/hr. EKS $0.10/hr per cluster = $72/mo. Paid: EKS $144/mo with extended support + $0.10/hr. Karpenter free. 1000 GPU cluster control plane $144/mo + 1000 x EC2 $4/hr = $4k/hr compute. Anyscale $2000/mo platform + compute.

Artifact Registry / Container Registry

Pricing & free-tier limitsFree: GHCR 500MB private storage + 2k mins Actions free, unlimited public. ECR Public free pulls. Docker Hub free 1 private repo 200 pulls/6h. Cheaper: ECR private $0.10/GB/mo storage, data transfer free within region to EC2. GitLab free 10GB registry. Paid: Docker Hub Pro $9/mo 5 private repos unlimited pulls. JFrog Pro $98/mo 5 users 25GB + $0.50/GB over. Artifactory storage: training images 50GB x20 =1TB ~$100/mo ECR.

Vector Database for Evaluation / RAG Eval

Pricing & free-tier limitsFree: Pinecone Free 100k vectors up to 2GB 1 index, 100 QPS. Qdrant OSS free unlimited self-host. Chroma OSS free. Cheaper: Qdrant Cloud Free 1GB RAM 500k vectors, $25/mo 1GB cluster 1M vectors. Pinecone Starter $0 with 100k free then $70/mo. Paid: Pinecone Serverless $0.33/GB/mo storage + $0.08/1M reads + $2/1M writes. Qdrant Cloud $75/mo 4GB RAM 4M vectors. Over 10M vectors 768d = 30GB RAM = $300/mo Qdrant Pro.

Inference Preview / Base Model Serving (for eval)

Pricing & free-tier limitsFree: vLLM OSS self-host free + GPU $1.5-4/hr. Ollama local free. HF Inference free 30k req/mo. Cheaper: Modal free $30 credit/mo then $2.10/hr H100 + $0.000046/sec GPU idle, $0/min scale-to-zero. RunPod Serverless free idle, $0.0002/sec H100 => 70B inference $0.72/hr active. Paid: Anyscale $10 free credits then Llama 70B $0.25/1M input $1/1M output. Fireworks $0.20/1M input. Replicate $5 free then $0.0014/sec H100. Preview cluster for 70B vLLM 2xH100 $8/hr.

Distributed File System (Lustre, JuiceFS, Alluxio)

Pricing & free-tier limitsFree: JuiceFS OSS free + S3 cost, CephFS OSS free self-host, Lustre community free. Cheaper: JuiceFS Cloud Free 10GB then $10/mo 100GB metadata + S3 $23/TB. Alluxio OSS free self-host $100/mo VM. Paid: FSx Lustre Scratch 1.2TB min $168/mo @ $0.14/GB, Persistent $0.24/GB = $240/TB/mo + $3.90/MBps-month throughput. Weka min $50k/mo PB-scale $0.35/GB = $350/TB.

Queuing / Job Scheduler for Training

Pricing & free-tier limitsFree: Slurm OSS free unlimited nodes, Kueue free, Volcano free, Airflow OSS free 1k DAGs. Ray Jobs free. Cheaper: AWS Batch free scheduler, you pay EC2 $1.5-12/GPU/hr only. GCP Batch same. Paid: LSF $5000/node/yr Enterprise, PBS Pro $10k/cluster $1200/node. AWS Batch Enterprise Support $15k/mo. Queue for 1000 GPUs free with Slurm.

Dataset Versioning / Data Catalog (DVC, LakeFS)

Pricing & free-tier limitsFree: DVC OSS free S3 backend cost only $23/TB. LakeFS OSS free self-host $100/mo VM. HF Datasets free private datasets with Pro plan, public unlimited. Cheaper: DVC Studio Free 20GB remote cache, 500 API calls/day. LakeFS Cloud Free 3GB branches. Paid: DVC Studio Team $20/u/mo 3 seats min $60/mo + cache $0.25/GB = $250/TB/mo. LakeFS Cloud $50/mo 100GB + $0.20/GB over. 10PB dataset versioning metadata ~1TB ~$250/mo.

Log Management for Training Logs

Pricing & free-tier limitsFree: Loki OSS free unlimited self-host + S3 $23/TB. OpenSearch OSS free. Grafana Cloud Free 50GB logs, 50GB traces, 10k series metrics. Cheaper: Axiom Free 500GB/mo logs, $25/mo 10GB. Better Stack Free 3GB logs 3d retention. Paid: Grafana Pro $8/user/mo + $0.50/GB logs over 50GB. Datadog Logs $0.10/GB ingest (100GB = $10) + $1.70/1M scanned, retention $0.02/GB/mo. Splunk $0.80/GB/day = $24/GB/mo. Training generates 10TB/mo logs = $5k Datadog, $500 Loki.

Secrets Management

Pricing & free-tier limitsFree: Vault OSS free unlimited secrets self-host + VM $20/mo. Infisical OSS free self-host. Sealed Secrets OSS free. Cheaper: Infisical Cloud Free 1000 secrets 2 envs, Doppler Free 100 secrets 5 users. Paid: AWS Secrets $0.40/secret/mo + $0.05/10k API calls. Example 500 API keys $200/mo + $50 calls. Vault Cloud $0.03/hr small = $22/mo + $0.10/secret-hour large. Azure KeyVault $0.03/10k ops + $0.30/cert/mo.

Security Scanning & Data Filtering (PII, Toxic, CSAM)

Pricing & free-tier limitsFree: Presidio free self-host CPU $0.05/hr, NeMo Curator free GPU 1 H100 $4/hr for PB filtering, Llama Guard 8B free OSS $0.20/1M tokens self-host. ClamAV free. Cheaper: AWS Macie Free trial 30d 1TB free. Paid: Macie $0.10/GB first TB = $100/TB, then $0.05/GB = $50/TB. 10PB training data scan = $500k one-time. Datadog Sensitive $0.25/GB = $2500/TB. Scale to 15T tokens ~ 60TB text = $3-15k scan OSS + compute.

Options and prices come straight from our research sheets for a small scale foundation model project. Prices are estimates and change often — always confirm on the provider's page before committing.

How this small scale foundation model checklist works.

Each requirement below is something a small scale foundation model build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.

Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.