Machine Learning Platforms — hobby stack.
For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.
GPU Training Cluster (A100/H100 - Core Compute)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: Colab T4 12h/week, 15GB GPU RAM. Lambda: A100 $1.10/hr, H100 SXM $2.49/hr. AWS p4d 8xA100 $32.77/hr ($4.10/GPU-hr). H100 8x $98.32/hr. Free tier $0, overage $1.10-4.50/GPU-hr, Spot 60-70% off.
Distributed Training Orchestration (Ray, DeepSpeed, Horovod)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: OSS $0 self-hosted. Anyscale: Free tier $0 for 5 hrs trial, then $1.20/hr head node + workers. DeepSpeed OSS free, managed $0.50/GPU-hr overhead. Overage: data transfer $0.09/GB.
Experiment Tracking & Visualization (W&B, MLflow)
Ships features to a subset of users or toggles them without redeploying. De-risks releases and enables experiments.Pricing & free-tier limitsW&B Free: 100GB storage, 1 team, unlimited experiments personal. Team $50/user/mo. MLflow OSS $0 + EC2 ~$20/mo. Neptune Free 1 user, 100h monitoring. Overage: W&B storage $0.60/GB/mo.
Model Registry & Versioning
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: MLflow OSS $0, HF Free: unlimited public, 1 private model limit 2026. Databricks: $0.40/DBU + storage. SageMaker Registry: $0 listing, storage $0.023/GB/mo. HF Pro $9/user/mo unlimited private.
Artifact Storage for Models & Datasets (S3/GCS)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: R2 10GB storage, 10M Class A ops, zero egress fees. S3 Free 5GB 12mo. Wasabi $6.99/TB no egress. S3 Standard $0.023/GB/mo, $0.0004/1k GET, $0.09/GB egress. Overage: $0.023/GB.
Data Versioning & Lakehouse (DVC, LakeFS, Delta Lake)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: DVC OSS $0, lakeFS Community $0 self-hosted. lakeFS Cloud Free: 3 repos, 2M objects/mo, 10GB cache. Paid: $0.10/GB scanned + $300/mo base. Databricks DLT $0.20/DBU. Overage: $0.05/GB versioned.
Feature Store (Feast, Tecton, Hopsworks)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Feast OSS $0, Hopsworks Community $0 (3 users). Tecton Free trial $0, then $2k/mo base. SageMaker FS: $0.05/1k writes, $0.01/1k reads, storage $0.023/GB. Feast hosting ~$50/mo infra.
Data Pipeline / ETL Orchestration (Airflow, Spark, Ray Data)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: Airflow OSS $0. Astronomer Free: 1 deployment, $0 trial. MWAA: $0.49/env-hr ~$350/mo. Databricks Jobs $0.15/DBU. Spark on EKS: on-demand $0.10/vCPU-hr. Overage: $0.49/hr orchestration.
Vector Database for Evaluation & Retrieval (RAG)
Persistent storage for your app’s data — users, products, orders. The single most important architectural decision for most projects.Pricing & free-tier limitsFree: Pinecone Free 100k vectors (1 index). Qdrant Free 1GB RAM, 1 cluster. Chroma $0. Paid: Pinecone Starter $70/mo 1M vectors. Serverless $0.33/1M writes, $0.08/1M reads, storage $0.33/GB. Overage: 100% over = auto upgrade.
Model Serving / Inference Framework (Seldon, KServe, BentoML, vLLM)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: OSS $0 + GPU cost. BentoCloud Free: $0 trial 200hrs. Anyscale Endpoints: $0.50/1M tokens Llama2-70B + GPU hr. Seldon Enterprise $3k/mo base. vLLM self-hosted $0 + A10G $0.75/hr. Overage: $0.75-4/GPU-hr.
Serverless GPU Inference Hosting (Modal, Replicate, Anyscale)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: Modal $30/mo credit (A10G ~40hrs). Replicate $5 free credit. Paid: Modal A10G $0.75/hr, A100 $1.95/hr, billed per second (0.01s). Replicate $0.000725/s A100 ($2.61/hr). Anyscale Llama3-70B $0.59/1M input, $0.79/1M output tokens. Overage: per-second GPU $0.0002-0.0012/s.
How this hobby machine learning platforms checklist works.
Each requirement below is something a hobby machine learning platforms build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.
Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.