Machine Learning Platforms — enterprise stack.
For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.
GPU Training Cluster (A100/H100 - Core Compute)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: Colab T4 12h/week, 15GB GPU RAM. Lambda: A100 $1.10/hr, H100 SXM $2.49/hr. AWS p4d 8xA100 $32.77/hr ($4.10/GPU-hr). H100 8x $98.32/hr. Free tier $0, overage $1.10-4.50/GPU-hr, Spot 60-70% off.
Distributed Training Orchestration (Ray, DeepSpeed, Horovod)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: OSS $0 self-hosted. Anyscale: Free tier $0 for 5 hrs trial, then $1.20/hr head node + workers. DeepSpeed OSS free, managed $0.50/GPU-hr overhead. Overage: data transfer $0.09/GB.
Experiment Tracking & Visualization (W&B, MLflow)
Ships features to a subset of users or toggles them without redeploying. De-risks releases and enables experiments.Pricing & free-tier limitsW&B Free: 100GB storage, 1 team, unlimited experiments personal. Team $50/user/mo. MLflow OSS $0 + EC2 ~$20/mo. Neptune Free 1 user, 100h monitoring. Overage: W&B storage $0.60/GB/mo.
Model Registry & Versioning
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: MLflow OSS $0, HF Free: unlimited public, 1 private model limit 2026. Databricks: $0.40/DBU + storage. SageMaker Registry: $0 listing, storage $0.023/GB/mo. HF Pro $9/user/mo unlimited private.
Artifact Storage for Models & Datasets (S3/GCS)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: R2 10GB storage, 10M Class A ops, zero egress fees. S3 Free 5GB 12mo. Wasabi $6.99/TB no egress. S3 Standard $0.023/GB/mo, $0.0004/1k GET, $0.09/GB egress. Overage: $0.023/GB.
Data Versioning & Lakehouse (DVC, LakeFS, Delta Lake)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: DVC OSS $0, lakeFS Community $0 self-hosted. lakeFS Cloud Free: 3 repos, 2M objects/mo, 10GB cache. Paid: $0.10/GB scanned + $300/mo base. Databricks DLT $0.20/DBU. Overage: $0.05/GB versioned.
Feature Store (Feast, Tecton, Hopsworks)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Feast OSS $0, Hopsworks Community $0 (3 users). Tecton Free trial $0, then $2k/mo base. SageMaker FS: $0.05/1k writes, $0.01/1k reads, storage $0.023/GB. Feast hosting ~$50/mo infra.
Data Pipeline / ETL Orchestration (Airflow, Spark, Ray Data)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Airflow OSS $0. Astronomer Free: 1 deployment, $0 trial. MWAA: $0.49/env-hr ~$350/mo. Databricks Jobs $0.15/DBU. Spark on EKS: on-demand $0.10/vCPU-hr. Overage: $0.49/hr orchestration.
Vector Database for Evaluation & Retrieval (RAG)
Persistent storage for your app’s data — users, products, orders. The single most important architectural decision for most projects.Pricing & free-tier limitsFree: Pinecone Free 100k vectors (1 index). Qdrant Free 1GB RAM, 1 cluster. Chroma $0. Paid: Pinecone Starter $70/mo 1M vectors. Serverless $0.33/1M writes, $0.08/1M reads, storage $0.33/GB. Overage: 100% over = auto upgrade.
Model Serving / Inference Framework (Seldon, KServe, BentoML, vLLM)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: OSS $0 + GPU cost. BentoCloud Free: $0 trial 200hrs. Anyscale Endpoints: $0.50/1M tokens Llama2-70B + GPU hr. Seldon Enterprise $3k/mo base. vLLM self-hosted $0 + A10G $0.75/hr. Overage: $0.75-4/GPU-hr.
Serverless GPU Inference Hosting (Modal, Replicate, Anyscale)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: Modal $30/mo credit (A10G ~40hrs). Replicate $5 free credit. Paid: Modal A10G $0.75/hr, A100 $1.95/hr, billed per second (0.01s). Replicate $0.000725/s A100 ($2.61/hr). Anyscale Llama3-70B $0.59/1M input, $0.79/1M output tokens. Overage: per-second GPU $0.0002-0.0012/s.
Hyperparameter Tuning (Optuna, Ray Tune, W&B Sweeps)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Optuna, Ray Tune OSS $0. W&B Sweeps Free 3 sweeps concurrent. Paid: W&B Team $50/user/mo includes sweeps. Determined AI Free trial, then $2k/mo cluster. Self-host cost ~ $0 + GPU. Overage: GPU $1-4/hr.
CI/CD for ML Pipelines (GitHub Actions, Metaflow, MLflow Pipelines)
Automates testing and deploying your code. Saves enormous time and prevents “works on my machine” releases.Pricing & free-tier limitsFree: GH Actions 2000 min/mo Linux, Metaflow $0. Paid: GH Team $4/user/mo 3000 min, larger runners $0.008/min 2-core. AWS CodePipeline $1/active pipeline/mo. Self-hosted runner $0 + EC2. Overage: $0.008-0.016/min.
Notebook / IDE Infrastructure (JupyterHub, SageMaker Studio)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Kaggle 30h GPU/week, Colab 12h. JupyterHub self-host $0 + EC2 ~$30/mo. SageMaker Studio: ml.t3.medium $0.05/hr, ml.g5.xlarge $1.41/hr. Deepnote Free 750h standard. Overage: $0.05-3.5/hr instance.
Model Monitoring & Drift Detection (Evidently, WhyLabs, Arize)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsFree: Evidently OSS $0, WhyLabs Free 1M profiles/mo. Arize Free 1k predictions/day. Paid: WhyLabs Pro $250/mo 5M profiles, Arize $500/mo base. Evidently Cloud $200/mo. Overage: $0.00001/prediction, $0.15/1k inferences.
LLM Evaluation & Tracing (LangSmith, Langfuse, Helicone)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsFree: LangSmith Free 5k traces/mo. Langfuse Cloud Free 50k observations/mo. Helicone Free 10k requests/mo. Paid: LangSmith Plus $79/mo 10k traces, Pro $399/mo 50k. Langfuse Pro $39/mo 100k obs. Overage: $0.005/1k observations, LLM token log $0.001/1k tokens.
Logging & Observability (Datadog, Grafana Loki, OpenTelemetry)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsFree: Grafana Cloud Free 50GB logs, 50GB traces, 10k metrics. SigNoz OSS $0. Datadog Free 14d trial then $15/host/mo + logs $0.10/GB. Grafana Pro $29/mo. Overage: $0.50/GB logs ingested, $0.10/1k metrics.
Secrets Management (HashiCorp Vault, AWS Secrets Manager)
Stores API keys and credentials safely, separate from your code. Leaking secrets is one of the most common breaches.Pricing & free-tier limitsFree: Vault OSS $0, Doppler Free 5 projects. AWS Secrets Manager Free 40 secrets 30d. Paid: AWS $0.40/secret/mo + $0.05/10k API calls. Vault Cloud $0.03/hr small ~$22/mo. Doppler $18/user/mo. Overage: $0.40/secret/mo.
Container Registry & Orchestration (Kubernetes - EKS/GKE)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: GHCR Free 500MB storage, Docker Hub Free 1 private repo. EKS $73/cluster/mo + EC2. ECR Free 500MB/mo 12mo, then $0.10/GB-mo + $0.20/GB transfer. GKE Standard $0.10/hr cluster (~$73/mo). DOKS $12/mo. Overage: $0.10/GB registry + $73 cluster.
Data Warehouse / Analytical Store for Training Data
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: BigQuery Free 10GB storage, 1TB queries/mo. ClickHouse Cloud Free 5GB. Snowflake Free trial $400 credits. Paid: Snowflake $2/credit, storage $23/TB/mo compressed. BQ $5/TB scanned, storage $0.02/GB. Databricks SQL $0.22/DBU. Overage: $5/TB query.
Streaming Data Ingestion (Kafka, Redpanda, Kinesis)
Delivers and processes audio/video to users. Bandwidth-heavy, so plan the cost.Pricing & free-tier limitsFree: Upstash Kafka Free 10k msg/day. Redpanda Serverless Free 30-day. Confluent Free $0 first $200 usage. Paid: Confluent $0.12/GB ingress, $0.12/GB egress, storage $0.10/GB-mo. Kinesis On-Demand $0.023/GB ingested. Redpanda $0.10/GB. Overage: $0.12/GB.
Distributed / High-Performance Storage (FSx Lustre, JuiceFS, Weka)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: JuiceFS OSS $0 + S3 costs, MinIO OSS $0. FSx Lustre Free tier none, minimum $3.06/TB-mo SSD 1200 MB/s. Paid: FSx $0.14/GB-mo scratch, persistent $0.20/GB-mo. Throughput $0.02/MBps-mo. EFS $0.30/GB-mo. Overage: $0.14/GB FSx.
Model Optimization & Compression (ONNX Runtime, TensorRT, Quantization)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: ONNX Runtime, TensorRT, Optimum $0 OSS self-hosted. OctoML Free trial, then pay per optimized model $500/mo base. Deci $0 trial. SaaS optimization $0.10/model-hr optimization compute. GPU for TensorRT conversion ~ $1-2/hr. Overage: $0/mo OSS, Enterprise $2500/mo.
GPU Cost Management, Billing & Quotas (FinOps)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: OpenCost OSS $0, Kubecost Free 1 cluster. Vantage Free trial $0. Paid: Kubecost Enterprise $499/cluster/mo. Vantage $0 Pro plan $0, Team $100/mo + 1% savings. Spot.io $0 base + 20% of savings. Overage: per cluster $349-599/mo.
Authentication & SSO for ML Teams (Auth0, Keycloak, Okta)
How users sign up, log in, and are authorized. Getting roles and access control right early prevents painful rewrites.Pricing & free-tier limitsFree: Auth0 Free 7.5k MAU, unlimited social logins. Keycloak OSS $0. Clerk Free 10k MAU. Paid: Auth0 Pro $35/mo 1k MAU then $0.07/MAU. Okta $2/user/mo. WorkOS $150/mo base + SSO. Overage: $0.07/MAU or $0.20/extra SSO connection.
Multi-tenant Isolation & Workspaces (Namespaces, Determined AI)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: K8s namespaces $0, Kubeflow multi-tenancy OSS $0. vCluster Free 3. Paid: Run:ai $30k/year base + per GPU $200/yr. Determined Enterprise $5k/mo. SageMaker Domain $0 domain + user per-hr $0.05. Overage: $200/GPU/yr management.
Audit Logs WORM for Experiments & Compliance
Keeps you secure and compliant with regulations. Required for enterprise customers and handling sensitive data.Pricing & free-tier limitsFree: CloudTrail 1 free trail, S3 Object Lock $0 feature + storage $0.023/GB. Axiom Free 500GB logs. Paid: CloudTrail Lake $0.75/GB ingested, storage $0.023/GB. Datadog Audit $0.20/GB. Axiom Pro $100/mo 1TB. WORM retention min 1 day to 100 years.
Backup & Disaster Recovery for Models (Velero + S3 Cross-Region)
Copies of your data so a bug, hack, or outage doesn’t become permanent loss. Cheap insurance every serious project needs.Pricing & free-tier limitsFree: Velero OSS $0, S3 Free 5GB. Kasten Free up to 5 nodes free. Paid: Kasten $2k/node/yr. AWS Backup $0.05/GB-mo warm, $0.01 cold, restore $0.02/GB. S3 Cross-Region $0.02/GB transfer. Overage: $0.05/GB backup.
Compliance & Governance (SOC2, HIPAA, Model Cards)
Keeps you secure and compliant with regulations. Required for enterprise customers and handling sensitive data.Pricing & free-tier limitsFree: Drata Free compliance check $0. Vanta Free trial $0. Paid: Vanta $7.5k/yr base + $3k/framework, Growth $15k/yr. Drata $10k/yr base. SOC2 Audit itself $20k-50k one-time external. Overage: per vendor $100/mo, per framework $3000/yr.
Security Scanning for Models & Containers (Snyk, Trivy, HiddenLayer)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Trivy OSS $0, Snyk Free 200 container tests/mo, 1 user. HiddenLayer OSS model scan $0. Paid: Snyk Team $98/dev/mo includes 10k tests. Wiz $50k/yr base. HiddenLayer Enterprise $40k/yr. JFrog Xray $30/mo. Overage: $0.10/scan, Snyk $0.20/container overage.
Edge Deployment & Model Sync (TF Lite, ONNX Edge, Fleet)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: TFLite OSS $0, ONNX Mobile $0. Edge Impulse Free 2 edge devices. Fleet OSS $0 self-hosted. Paid: Greengrass $0.16/device/mo, Fleet Premium $22/100 devices/mo. OctoML Edge $500/mo base. NVIDIA Fleet Command $35/device/mo. Overage: $0.16/device.
Autoscaling & Spot Instance Management (Karpenter, Spot.io)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Karpenter OSS $0 + EC2 Spot 60-90% savings (A100 spot $1.5 vs $4 on-demand). Spot.io Free trial 14d. Paid: Spot.io 20% of savings fee. Cast.ai Free trial, then $500/mo cluster + 20% savings share. Self-host Karpenter $0 infra only. Overage: EC2 on-demand $0.10-4/GPU-hr.
Metadata Store & Lineage Tracking (OpenLineage, Amundsen, MLMD)
Rates, labels, and tracking for physical goods. Needed once you actually ship products.Pricing & free-tier limitsFree: MLMD OSS $0, OpenLineage $0, DataHub OSS $0 self-hosted. DataHub Cloud Free 30d trial. Paid: Atlan $25k/yr base 10 users, DataHub Pro $1500/mo. Unity Catalog $0 + DBU. Managed OpenLineage $300/mo. Overage: $0.05/1k lineage events.
Cache & Feature Caching Layer (Redis, Dragonfly, Upstash)
Stores frequent results in fast memory so you serve less from the database. Big lever for speed and cost.Pricing & free-tier limitsFree: Upstash Free 10k commands/day, 256MB. Redis Cloud Free 30MB. Paid: Redis Cloud Essentials $12/mo 250MB, Pro $100/mo 10GB $0.80/GB-hr. ElastiCache Serverless $0.125/GB-hr storage + $0.20/ECPU. Dragonfly Cloud $25/mo 1GB. Overage: $0.20/ECPU + $0.125/GB-hr.
Notification & Alerting for Training Jobs (PagerDuty, Slack, OpsGenie)
Reaches users on their phone via SMS or push. Useful for verification codes and re-engagement; watch the per-message cost.Pricing & free-tier limitsFree: Slack Free webhooks $0, PagerDuty Free 1 user 30 alerts/mo. Grafana Alerts Free unlimited. Paid: PagerDuty Team $34/user/mo, Business $64/user/mo. OpsGenie $29/user/mo. Slack Pro $8.75/user/mo. Overage: $0.20/SMS, $0.01/email over quota.
Service Mesh & API Gateway for Model APIs (Istio, Kong, AWS API GW)
Controls and protects your APIs — quotas, abuse prevention, and firewalls. Important once you have real traffic or many clients.Pricing & free-tier limitsFree: Istio OSS $0, Kong OSS $0, API GW Free 1M calls 12mo. Paid: API GW $3.50/1M calls + $0.09/GB transfer. Kong Konnect $250/mo base. Tetrate Istio $500/mo cluster. AWS ALB $0.0225/hr + LCU. Overage: $3.50/1M API calls.
Private Networking / VPC for Training (AWS VPC, Tailscale, ZTNA)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Tailscale Free 3 users/100 devices, Cloudflare ZT Free 50 users. AWS VPC $0 (no charge) + NAT GW $0.045/hr + $0.045/GB. Paid: Tailscale Business $6/user/mo, Enterprise $15/user. Cloudflare Std $7/user/mo. TGW $0.05/hr attachment + $0.02/GB. Overage: NAT $0.045/GB, Tailscale $6/user/mo.
Data Labeling & Human-in-the-Loop Infra (Label Studio, Scale AI, Argilla)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: Label Studio OSS $0, Argilla OSS $0, CVAT Free. Label Studio Cloud Free 3 seats, 10k data rows. Paid: Label Studio Enterprise $800/mo 5 seats unlimited. Scale AI $0.05-0.15/label image, $0.08/token text. Labelbox Starter $500/mo. Overage: $0.06/label.
How this enterprise machine learning platforms checklist works.
Each requirement below is something a enterprise machine learning platforms build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.
Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.