Foundation Model — enterprise stack.
For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.
GPU Supercluster / Training Compute (A100/H100/B200)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: Colab T4 12GB 5h/day limited, Kaggle 30h P100/wk, Lightning 22h T4/mo. Cheaper: Vast.ai 4090 $0.69/hr, 2xH100 $3.98/hr; RunPod Secure H100 $2.89/hr, Community $1.99/hr. Paid: CoreWeave H100 $3.96-4.76/hr per GPU ($31.68-38.08/8x), AWS p5.48xlarge $98.32/hr ($12.29/GPU), Lambda 8xH100 $32.08/hr ($4.01/GPU), Azure ND96 $99.76/hr. Spot 40-60% off. B200 ~$7-9/hr/GPU in 2026. Overage: inter-node egress $0.01/GB.
Distributed Training Framework
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: PyTorch DDP, DeepSpeed OSS, Accelerate, Colossal-AI free self-host unlimited. Cheaper: Self-host on RunPod/Vast at $0.69/hr infra, no license. Paid: Anyscale free 20 core-hrs/mo then $0.55/vCPU/hr + $2000/mo platform. Databricks Mosaic training overhead $0.00065/token (~$9750 for 15T tokens). Deterministic AI $2k/mo.
Data Lake / Object Storage (Petabyte Scale)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: MinIO OSS unlimited self-host, R2 Free 10GB, 10M Class A ops, 10M Class B, zero egress. S3 Free Tier 5GB 12mo. Cheaper: B2 $6/TB/mo storage, free downloads 3x stored/day, $0.01/GB over. Wasabi $6.99/TB zero egress, min 90d retention. Paid: S3 $23/TB/mo + $0.005/1k PUT + $0.09/GB egress. GCS $20/TB + egress $0.12/GB. 100PB dataset = $2.3M/mo S3. R2 $15/TB + $4.50/M operations.
Data Pipeline / ETL for Training Data (Spark, Ray Data)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Spark OSS self-host on K8s free, Ray Data OSS free, DuckDB free single node. Cheaper: Modal free $30 credit/mo then CPU $0.000046/sec, Coiled free 5k OCPU-hrs. Dagster OSS free. Paid: Databricks $0.40/DBU + EC2 = $0.75-1.50/hr per node DBU+EC2 combined, Photon $0.60/DBU. AWS EMR $0.27/hr per m5.xlarge + EC2 $0.192/hr = $0.462/hr. Glue free 1M objects + 1 DPU-hr/mo then $0.44/DPU-hr.
Checkpoint Storage (High-speed NVMe + S3)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: Instance NVMe 8x3.84TB free with p4d.24xlarge/p5. JuiceFS OSS free + S3 cost only. Cheaper: RunPod Volume $0.10/GB/mo = $100/TB/mo, Vast included. JuiceFS Cloud Free 10GB then $0.02/GB metadata + S3. Paid: FSx Lustre $0.14/GB/mo = $140/TB + throughput $3.90/MBps-month. EFS $300/TB/mo. Weka $0.25/GB = $250/TB/mo + IOPs. S3 Express $0.16/GB = $160/TB/mo + $0.0025/1k GET. 70B model ckpt 140GB x500 =70TB ~$7k-21k/mo.
Experiment Tracking / MLOps Platform (W&B, MLflow)
Ships features to a subset of users or toggles them without redeploying. De-risks releases and enables experiments.Pricing & free-tier limitsFree: W&B Free 100GB storage, 1 team, personal unlimited experiments, 1h system metrics. MLflow OSS unlimited self-host free. Cheaper: Neptune Free 100GB, 10GB monthly tracking. Comet Free 5000h/mo + 2GB. Paid: W&B Team $50/u/mo 5-seat min $250/mo 250GB storage. Enterprise $100/u/mo 1TB + SSO. Overage: $0.50/GB storage, $0.10/1k API calls. MLflow Managed free w/ Databricks.
Model Registry & Versioning
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: HF Hub free unlimited public models, 10 private repos 100GB bandwidth. MLflow OSS free unlimited. Cheaper: HF Pro $9/mo 1 private org, 1TB bandwidth, 300GB storage. DVC + S3 $23/TB. Paid: HF Enterprise Cloud $20/user/mo + infra. JFrog Pro $98/mo 5 users 25GB. Databricks UC free with compute. Model size 70B 140GB x20 versions 2.8TB ~ $64/mo S3.
Tokenization / Data Processing Service
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Tokenizers OSS 1M+ tok/sec per core, unlimited free, tiktoken OSS free, SentencePiece OSS free local. Cheaper: Modal CPU $0.000046/sec => tokenize 1B tokens ~$0.50. Self-host free. Paid: Cohere $0.10/1M tokens tokenization API, 15T tokens = $1500 API cost, better OSS $2000 GPU compute 500 H100-hrs Ray Data $4/hr = $2000 one-time for 15T.
Evaluation Harness / Benchmarking (HELM, LM-Eval)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: LM-Eval Harness free self-host unlimited benchmarks, HELM free. HF Evaluate OSS free. Cheaper: HF Inference free 30k req/mo then $0.06/1k req. Modal eval $2.10/hr H100. Paid: Scale AI eval $2-5k per full run (100 tasks x 70B). LangSmith Plus $59/mo 10k eval traces, $0.50/1k over. Anyscale evaluation endpoints: Llama 70B $0.25/1M input, $1/1M output. Full benchmark suite 500M tokens ~ $200-500.
Container Orchestration (Kubernetes for Training)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: k3s OSS free self-host, Kind/MicroK8s free local unlimited. Cheaper: EKS free credit $144? Actually GKE free 1 zonal cluster free, Autopilot $0.10/vCPU/hr. EKS $0.10/hr per cluster = $72/mo. Paid: EKS $144/mo with extended support + $0.10/hr. Karpenter free. 1000 GPU cluster control plane $144/mo + 1000 x EC2 $4/hr = $4k/hr compute. Anyscale $2000/mo platform + compute.
Artifact Registry / Container Registry
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: GHCR 500MB private storage + 2k mins Actions free, unlimited public. ECR Public free pulls. Docker Hub free 1 private repo 200 pulls/6h. Cheaper: ECR private $0.10/GB/mo storage, data transfer free within region to EC2. GitLab free 10GB registry. Paid: Docker Hub Pro $9/mo 5 private repos unlimited pulls. JFrog Pro $98/mo 5 users 25GB + $0.50/GB over. Artifactory storage: training images 50GB x20 =1TB ~$100/mo ECR.
Vector Database for Evaluation / RAG Eval
Persistent storage for your app’s data — users, products, orders. The single most important architectural decision for most projects.Pricing & free-tier limitsFree: Pinecone Free 100k vectors up to 2GB 1 index, 100 QPS. Qdrant OSS free unlimited self-host. Chroma OSS free. Cheaper: Qdrant Cloud Free 1GB RAM 500k vectors, $25/mo 1GB cluster 1M vectors. Pinecone Starter $0 with 100k free then $70/mo. Paid: Pinecone Serverless $0.33/GB/mo storage + $0.08/1M reads + $2/1M writes. Qdrant Cloud $75/mo 4GB RAM 4M vectors. Over 10M vectors 768d = 30GB RAM = $300/mo Qdrant Pro.
Inference Preview / Base Model Serving (for eval)
An isolated place to test changes before real users see them. Catches bugs without embarrassing your production users.Pricing & free-tier limitsFree: vLLM OSS self-host free + GPU $1.5-4/hr. Ollama local free. HF Inference free 30k req/mo. Cheaper: Modal free $30 credit/mo then $2.10/hr H100 + $0.000046/sec GPU idle, $0/min scale-to-zero. RunPod Serverless free idle, $0.0002/sec H100 => 70B inference $0.72/hr active. Paid: Anyscale $10 free credits then Llama 70B $0.25/1M input $1/1M output. Fireworks $0.20/1M input. Replicate $5 free then $0.0014/sec H100. Preview cluster for 70B vLLM 2xH100 $8/hr.
Distributed File System (Lustre, JuiceFS, Alluxio)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: JuiceFS OSS free + S3 cost, CephFS OSS free self-host, Lustre community free. Cheaper: JuiceFS Cloud Free 10GB then $10/mo 100GB metadata + S3 $23/TB. Alluxio OSS free self-host $100/mo VM. Paid: FSx Lustre Scratch 1.2TB min $168/mo @ $0.14/GB, Persistent $0.24/GB = $240/TB/mo + $3.90/MBps-month throughput. Weka min $50k/mo PB-scale $0.35/GB = $350/TB.
Queuing / Job Scheduler for Training
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Slurm OSS free unlimited nodes, Kueue free, Volcano free, Airflow OSS free 1k DAGs. Ray Jobs free. Cheaper: AWS Batch free scheduler, you pay EC2 $1.5-12/GPU/hr only. GCP Batch same. Paid: LSF $5000/node/yr Enterprise, PBS Pro $10k/cluster $1200/node. AWS Batch Enterprise Support $15k/mo. Queue for 1000 GPUs free with Slurm.
Dataset Versioning / Data Catalog (DVC, LakeFS)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: DVC OSS free S3 backend cost only $23/TB. LakeFS OSS free self-host $100/mo VM. HF Datasets free private datasets with Pro plan, public unlimited. Cheaper: DVC Studio Free 20GB remote cache, 500 API calls/day. LakeFS Cloud Free 3GB branches. Paid: DVC Studio Team $20/u/mo 3 seats min $60/mo + cache $0.25/GB = $250/TB/mo. LakeFS Cloud $50/mo 100GB + $0.20/GB over. 10PB dataset versioning metadata ~1TB ~$250/mo.
Log Management for Training Logs
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsFree: Loki OSS free unlimited self-host + S3 $23/TB. OpenSearch OSS free. Grafana Cloud Free 50GB logs, 50GB traces, 10k series metrics. Cheaper: Axiom Free 500GB/mo logs, $25/mo 10GB. Better Stack Free 3GB logs 3d retention. Paid: Grafana Pro $8/user/mo + $0.50/GB logs over 50GB. Datadog Logs $0.10/GB ingest (100GB = $10) + $1.70/1M scanned, retention $0.02/GB/mo. Splunk $0.80/GB/day = $24/GB/mo. Training generates 10TB/mo logs = $5k Datadog, $500 Loki.
Secrets Management
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Vault OSS free unlimited secrets self-host + VM $20/mo. Infisical OSS free self-host. Sealed Secrets OSS free. Cheaper: Infisical Cloud Free 1000 secrets 2 envs, Doppler Free 100 secrets 5 users. Paid: AWS Secrets $0.40/secret/mo + $0.05/10k API calls. Example 500 API keys $200/mo + $50 calls. Vault Cloud $0.03/hr small = $22/mo + $0.10/secret-hour large. Azure KeyVault $0.03/10k ops + $0.30/cert/mo.
Security Scanning & Data Filtering (PII, Toxic, CSAM)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Presidio free self-host CPU $0.05/hr, NeMo Curator free GPU 1 H100 $4/hr for PB filtering, Llama Guard 8B free OSS $0.20/1M tokens self-host. ClamAV free. Cheaper: AWS Macie Free trial 30d 1TB free. Paid: Macie $0.10/GB first TB = $100/TB, then $0.05/GB = $50/TB. 10PB training data scan = $500k one-time. Datadog Sensitive $0.25/GB = $2500/TB. Scale to 15T tokens ~ 60TB text = $3-15k scan OSS + compute.
GPU Monitoring & Observability (DCGM, Prometheus)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsFree: DCGM + DCGM Exporter free self-host, Prometheus free unlimited, Grafana OSS free. Cheaper: Grafana Cloud Free 10k series metrics, 50GB logs, 50GB traces. W&B free system metrics 100GB. Paid: Datadog Pro $23/host/mo + container $0.002/hr. GPU Monitoring $0.10/hr per GPU. 1000 H100 cluster => GPU $0.10x1000x730 = $73k/mo + hosts 125x$23 = $2875/mo. Chronosphere $750/mo cluster + $0.30/1k metrics.
Cost Tracking / FinOps for GPU (OpenCost, Kubecost)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: OpenCost OSS free unlimited clusters self-host. Kubecost Free 1 cluster 15d retention free. Cheaper: Infracost Free up to 1000 runs/mo, $25/mo 5000 runs. Paid: Kubecost Enterprise $999/mo per cluster unlimited retention, Kubecost Cloud Pro $499/mo 5 clusters. CloudHealth $15k/mo for $2M spend enterprise. Saves 40-60% via spot tracking for GPU $3/hr saved x1000 GPUs = $2M/mo savings.
High-Speed Interconnect / Networking (EFA, InfiniBand, RoCE)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: AWS EFA driver + fabric free with instance, IB included in GPU hourly price at all neoclouds. GPUDirect RDMA free. Cheaper: Same-AZ traffic free, cross-AZ $0.01/GB. CoreWeave/Lambda fabric inclusive. Paid: Equinix Fabric Fabric port $500/mo 10G + $2k/mo interconnect + $0.05/GB. AWS inter-AZ $0.01/GB = $10/TB. All-reduce for 70B training 15T tokens ~500TB internal traffic = $5k inter-AZ if misplaced, $0 if same-AZ placement group.
Cache / In-Memory Store for Data Loading
Stores frequent results in fast memory so you serve less from the database. Big lever for speed and cost.Pricing & free-tier limitsFree: Redis OSS free unlimited self-host on VM $50/mo 32GB. Dragonfly OSS free 25x faster. Cheaper: Upstash Free 10k cmds/day, 256MB, $0.20/100k cmds over. Dragonfly Cloud $10/mo 500MB. Paid: ElastiCache Serverless $0.125/GB/hr = $90/GB/mo RAM + $0.20/M ECPUs + $0.125/GB/hr backup. Example 500GB cache for dataloader = $45k/mo serverless, $5k/mo self-host. Redis Enterprise Cloud $35/mo 100MB + $0.50/100k ops.
Metadata Catalog / Data Lineage
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: DataHub OSS free self-host VM $100/mo 4 vCPU 16GB, OpenLineage free, Marquez free. Cheaper: DataHub Cloud free trial 30d then Team $500/mo 10 users. Atlan free trial then growth $50k/yr. Paid: Acryl Team $500/mo 10 users 1TB metadata, Enterprise $2000-5000/mo 100 users. Atlan $40k/yr starter, $80k/yr growth. Databricks Unity Catalog free with DBR, storage $23/TB lineage logs.
CI/CD for Training Jobs (GitHub Actions, Argo)
Automates testing and deploying your code. Saves enormous time and prevents “works on my machine” releases.Pricing & free-tier limitsFree: GH Actions Free 2000 min/mo Linux, 500MB storage, self-hosted GPU runners free but EC2 $4/hr. Argo OSS free unlimited workflows. Cheaper: GitLab Free 400 compute min/mo + self-host. Buildkite Free 1 agent. Paid: GH Actions Team $4/u/mo + Linux $0.008/min ($0.48/hr), GPU self-host still EC2 cost + time. Large training CI 100 jobs/day 60min H100 self-host: $480/day compute + $20/mo platform. AWS CodePipeline $1/pipeline/mo + $0.002/action over 100.
Model Optimization & Compilation (TensorRT-LLM, torch.compile)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: TensorRT-LLM OSS free + H100 $4/hr, torch.compile free unlimited, ONNX Runtime free, llama.cpp 4-bit quant free local. Cheaper: Modal $30 free then H100 $2.10/hr optimization experiments 100hr trial $210. Paid: Deci $0 trial then Developer $1500/mo 3 models, Enterprise $20k/mo unlimited. OctoML Starter $500/mo 2 models, Scale $15k/mo. Optimization reduces inference preview cost 40-70%: 70B FP16 $4/hr -> INT4 $1.2/hr.
Backup & Disaster Recovery for Checkpoints
Copies of your data so a bug, hack, or outage doesn’t become permanent loss. Cheap insurance every serious project needs.Pricing & free-tier limitsFree: Restic OSS free + S3 $23/TB/mo, Velero free + S3 storage only. MinIO OSS free. Cheaper: B2 $6/TB/mo = $6000/PB/mo backup replica, Wasabi $6.99/TB zero egress. Paid: AWS Backup $0.10/GB/mo vault storage = $100/TB/mo + CRR transfer $0.02/GB = $20/TB transfer. 1PB checkpoint backups cross-region = $20k transfer + $100k/mo storage S3 $23k = $43k-123k/mo. Glacier Deep Archive $1/TB/mo long-term = $1k/PB/mo for compliance.
Rate Limiting / API Gateway for Eval API
Controls and protects your APIs — quotas, abuse prevention, and firewalls. Important once you have real traffic or many clients.Pricing & free-tier limitsFree: Kong OSS free self-host VM $50/mo unlimited req, Cloudflare Free 100k req/day, 1000 Gateway rules, AWS Free Tier 1M calls/mo 12mo. Cheaper: Cloudflare Workers Paid $5/mo + $0.50/M requests includes 10M = $5/M after. AWS API GW $3.50/M req + $1/300M cache. Paid: Kong Konnect Team $250/mo includes 7.5M requests 3 services. AWS $3.50/M 100M eval requests = $350/mo + data 1TB $90. Cloudflare Enterprise Gateway $5k/mo 100M req.
Identity & Access Management (IAM for GPU Clusters)
How users sign up, log in, and are authorized. Getting roles and access control right early prevents painful rewrites.Pricing & free-tier limitsFree: AWS IAM free no cost, Teleport Community free unlimited 100 resources self-host VM $50/mo. Keycloak OSS free. Cheaper: Authentik OSS free self-host unlimited users, $50/mo infra. Paid: Teleport Enterprise $15/user/mo + $20/server/mo = 500 users + 1000 nodes = $27.5k/mo list. Okta Free 100 MAU then $2/user/mo + addons Workplace $8/u/mo. Auth0 Free 25k MAU then $0.07/MAU over = 10k extra = $700/mo. Cluster 500 eng.
Documentation / Knowledge Base for Infra
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: MkDocs free + GitHub Pages free hosting unlimited bandwidth, Docusaurus OSS free + Vercel free hobby 100GB bandwidth. Cheaper: GitBook Free 1 private space unlimited public projects, ReadTheDocs free 1GB storage. Paid: GitBook Pro $40/u/mo 10GB bandwidth $0.50/GB over. Notion Plus $10/u/mo 1000 blocks? Actually unlimited blocks $10/u/mo min $80. Confluence Standard $6.05/u x100 = $605/mo + storage 250GB free. Mintlify Pro $120/mo 200k page views + $0.50/1k over.
Compliance / WORM Storage for Training Logs (AI Act, SOC2)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: MinIO Object Lock WORM OSS free + disk $23/TB self-host, OPA free policy, Wazuh free SIEM. Cheaper: B2 with Object Lock $6/TB/mo min 90d retention, Wasabi $6.99/TB compliance. Paid: S3 Object Lock $0.023/GB = $23/TB/mo retention + API $0.005/1k PUT. Glacier Vault Lock $0.004/GB = $4/TB/mo, retrieval $0.03/GB. Vanta $10k/yr starter, $25k/yr growth audit. Training logs 100TB 7yr = $2300/mo WORM + audit $10k/yr.
Service Mesh / Secure Communication (Istio, Cilium)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Istio OSS free self-host mgmt VM $100/mo, Linkerd free <10ms sidecar latency, Cilium eBPF free. Cheaper: Gloo Mesh Free 5 workloads. Paid: Tetrate Enterprise $75k/yr starter 50 services + support. Solo Enterprise $100k/yr 100 services. AWS App Mesh $0.011/envoy/hr = $8/mo small proxy + $0.0015/GB data = $1.5/TB. 1000 pods mesh = $8k/mo App Mesh.
Network Security / VPC / Firewall / WAF
Controls and protects your APIs — quotas, abuse prevention, and firewalls. Important once you have real traffic or many clients.Pricing & free-tier limitsFree: VPC free, SG free unlimited rules, Cloudflare Free WAF 5 custom rules, DDoS unmetered mitigation. Cheaper: Cloudflare Pro $20/mo 20 WAF rules, $0.05/10k filtered requests? Actually included. Paid: AWS WAF $5/rule/mo + $1/M req WCU. 10 rules $50/mo + 100M req $100 = $150/mo. Network Firewall endpoint $0.395/hr = $288/mo AZ + data $0.065/GB. 100TB filtered = $6500/mo + $288 endpoint. Palo Alto VM $1.5/hr = $1080/mo each.
Load Balancer for Inference Preview (vLLM clusters)
An isolated place to test changes before real users see them. Catches bugs without embarrassing your production users.Pricing & free-tier limitsFree: NGINX OSS free self-host VM $20/mo, HAProxy free, Envoy free unlimited RPS. Cloudflare Free 1 LB pool free health checks. Cheaper: Cloudflare LB $5/mo first pool 500k queries, $10/mo extra pools. DigitalOcean LB $12/mo 10k rps. Paid: AWS NLB $0.0225/hr = $16.43/mo + LCU $0.006 per LCU-hour (new conns, active conns, bytes). Example 10k rps eval: 20 LCUs = $0.12/hr = $87/mo + base $16 = $103/mo. ALB similar $0.0225/hr + $0.008/LCU. F5 $3000/mo VE.
Incident Management / On-call (PagerDuty, Firefighting GPU cluster)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: Grafana OnCall OSS free unlimited on-call schedules self-host + VM $20/mo. PagerDuty Free 5 users, 1 escalation, email/push. Opsgenie Free 5 users. Cheaper: Better Stack Free 10 monitors 100 SMS. Incident.io Free 30 incidents/mo. Paid: PagerDuty Pro $41/u/mo (5 users $205/mo) + $5/phone call. Opsgenie $23/u/mo Std = $115/mo 5 users + $0.15/SMS. Incident.io Growth $49/u/mo = $245/mo 5 users + Slack sync. GPU cluster outage $10k/hr idle for 1000 H100, so on-call essential. $500/hr MTTR impact.
Data Anonymization / PII Scrubbing for Training Data
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Presidio OSS free self-host CPU $0.05/hr = $36/mo per pipeline + spaCy models free. Faker free unlimited fake data gen. Scrubadub free. Cheaper: Private AI Free 100k chars/mo then Starter $50/mo 1M chars + $5/100k over. Gretel Free 50GB synthetic. Paid: Comprehend PII $0.0001/unit = $0.20 per 1k docs 500 chars each. 10B docs = $2M Comprehend. Better OSS Presidio 15T token dataset (60T chars) = 120B units = 120B*100x cheaper OSS compute ~$10k GPU+CPU time 500 H100 hrs. Private AI $0.50/10k chars 60T = $3M.
Carbon / Power Usage Tracking + Sustainability Reporting
Rates, labels, and tracking for physical goods. Needed once you actually ship products.Pricing & free-tier limitsFree: CodeCarbon OSS free pip package unlimited tracking per training job, CCF OSS free self-host VM $50/mo tracks AWS/GCP/Azure. Cheaper: Climatiq Free 1000 estimates/mo, Electricity Maps free 100/day. Paid: Climatiq Pro $200/mo 100k req, $0.002/over = $2/1k. Watershed $30k/yr enterprise carbon accounting. Training 70B model 500 MWh ~ $60k electricity 200t CO2 x $50/t offset = $10k offsets. Tracking platform $200-1000/mo negligible vs compute $400k job.
Feature / Embedding Store (for RLHF, Alignment, Post-training)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: Feast OSS free + Redis OSS self-host $50/mo infra 32GB, Qdrant OSS free self-host unlimited embeddings. Cheaper: Upstash Vector Free 100k vectors, Feast self-host Redis $50/mo. Tecton free trial 30 days. Paid: Tecton Team $2500/mo platform + $0.50/1k feature retrievals, Growth $5000/mo. Pinecone $0.33/GB storage + $0.08/1M reads = 10M RLHF embeddings 768d 30GB = $10/mo storage + $8/mo reads. SageMaker Feature Store $0.30/GB/mo = 10TB embeddings = $3000/mo + $0.20/M writes.
How this enterprise foundation model checklist works.
Each requirement below is something a enterprise foundation model build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.
Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.