Infrastructure — small scale stack.
For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.
GPU Compute Rental (A100 / H100 / B200 Fleet)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: Modal $30/mo credits, 1x T4/A10G dev; Colab T4 12h session. Paid: RunPod T4 $0.22/hr, A100 40GB $1.19/hr, H100 $3.49/hr; Lambda A100 $1.50/hr, H100 $3.99/hr (on-demand), $2.49/hr 1yr reserved; CoreWeave H100 $4.25/hr OD, $2.99/hr reserved, B200 $7.50/hr OD $4.99/hr 3yr; Overage: Idle keep-alive $0.05/hr, egress $0.08/GB
Bare Metal Provisioning & Fleet Management
Enterprise single sign-on and user provisioning. Usually required to sell to larger organizations.Pricing & free-tier limitsFree: MaaS self-hosted $0, Ironic open-source. Cheaper: Vultr BM from $120/mo (1x RTX 4000), Latitude.sh c3.large $0.75/hr. Paid: Equinix Metal c3.medium $0.85/hr ~$550/mo, GPU BM g2.large A100 $2,200/mo + $1.5/hr metering; OVH A100 BM €1,800/mo. Overage: Provisioning API $0.01/call after 10k/mo, support $500/mo
GPU Virtualization / Fractional GPU (MIG, vGPU)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: KVM+MIG self-host $0, up to 7x MIG slices per A100/H100. Open tools free. Cheaper: Modal 0.1 GPU $0.15/hr fractional. Paid: Run:ai $0.45/GPU/hr license ($350/GPU/mo), VMware Bitfusion $500/GPU/mo. Overage: Extra MIG profile $0.10/slice/hr, Run:ai over-quota $0.55/GPU/hr
Kubernetes for GPU Workloads (GPU Operator & Scheduler)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: K3s OSS $0, GPU Operator OSS $0 (self-manage). Cheaper: DOKS GPU $1.10/node/hr includes control plane $0. EKS $0.10/hr cluster + EC2 p4d H100 $4.56/hr. Paid: Run:ai $450/GPU/mo platform fee, OpenShift AI $0.50/GPU/hr. Overage: Karpenter scale-up 30s, extra control plane $0.10/hr, GPU Operator enterprise support $2k/mo
Serverless GPU Platform / Model Hosting Platform
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: Modal $30/mo credits, 100 GPU sec free; Replicate free $5 credits, 500ms cold start. Cheaper: Baseten $0.0005/sec per A100 ($1.80/hr effective), Banana $0.00008/sec. Paid: Anyscale $1.00/hr worker + $1.50/A100/hr, Replicate $0.000725/s A100 40GB, Together $0.90/hr L4 + token pricing. Overage: Scale-to-zero wake $0.02, concurrent limit $5 per extra concurrency over 10
Inference Server Engine (vLLM, TGI, Triton)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsFree: vLLM/TGI OSS $0, self-host on existing GPU. Cheaper: Ollama $0 + $10/mo managed. Paid: NVIDIA NIM $0.45/GPU/hr license ($1/hr with H100 included via NGC), Baseten $1.80/hr managed vLLM. Overage: 8k RPS license +$0.10/GPU/hr, enterprise support $3k/mo for Triton
Inference Autoscaling (Scale-to-Zero, KEDA)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: KEDA/Knative OSS $0, scale-to-zero <2s on self-host. Cheaper: Fly.io scale-to-zero $0.10/hr min + $1.50/A100/hr active. Paid: Anyscale $50/mo autoscaler addon + GPU costs, Run:ai $100/GPU/mo autoscale license. Overage: KEDA queue lag scaling $0.001/s metric poll, excess scale events $0.02 per 1k over 100k/mo
Vector Database (Embedding Store for RAG)
Persistent storage for your app’s data — users, products, orders. The single most important architectural decision for most projects.Pricing & free-tier limitsFree: Qdrant OSS $0 unlimited self-host; Qdrant Cloud Free 1GB, 1M vectors, 1k QPS; Pinecone Free 100k vectors, 1 pod p1.x1. Cheaper: Qdrant Cloud $25/mo 5GB, Chroma $30/mo 5M vectors. Paid: Pinecone Starter $70/mo 1M vectors + $0.40/1M queries, $0.20/1M writes; Zilliz Standard $0.25/hr cluster. Overage: Pinecone $0.20 per 1M vectors over, Qdrant $5/GB over, $0.20/1M searches over 10M
Model Registry & Versioning (HuggingFace-like)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: HF Hub Free: public models unlimited, private 1 model, 100GB bandwidth; MLflow OSS $0. Cheaper: W&B Free 100GB artifact. HF Pro $9/mo, 50 private models. Paid: HF Enterprise $50/user/mo, private unlimited, 2TB/mo bandwidth, SSO; CoreWeave registry $0.20/GB/mo stored. Overage: HF $0.10/GB bandwidth over 2TB, $0.05/GB storage over 500GB
Object Storage for Multi-GB Model Weights
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: MinIO self-host $0, B2 Free 10GB storage. R2 Free 10GB store, 10M Class B ops/mo, zero egress. Cheaper: B2 $0.006/GB/mo, R2 $0.015/GB/mo store, $0 egress. Paid: S3 Standard $0.023/GB/mo store, $0.09/GB egress, $0.005/1k PUT; R2 $15/mo 500GB included. Wasabi $7.99/TB/mo no egress. Overage: S3 egress $0.09/GB, R2 $0 egress, ops $0.36/M, Wasabi over 1TB $7.99/TB
CDN / Edge Distribution for Model Weights
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: Cloudflare Free: unlimited egress, 100GB R2 fetch free. Cheaper: Bunny CDN $0.01/GB, $1/mo min; CF Pro $20/mo 20% more edge. Paid: Fastly $0.12/GB first 10TB, $0.08/GB next, $50/mo min; Cloudflare Enterprise $5k/mo + custom bandwidth. Overage: Fastly $0.12/GB, Bunny $0.02/GB over 10TB, CF no overage on free but $0.05/GB cache reserve
Model Training Orchestration (Jobs & Queues)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: Kubeflow OSS $0 self-host, Modal Jobs $30 credits. Cheaper: RunPod Jobs $0.20 setup + GPU $1.19/hr. Paid: Anyscale Jobs $0.25/hr orchestrator + GPU cost $1.50-3.99/hr, CoreWeave SUNK $0.10/hr management + GPU. Overage: Job retry $0.005/job, queue wait $0, failed job logs $0.20/GB, priority queue +25% GPU cost
Experiment Tracking & ML Metadata (W&B, MLflow)
Ships features to a subset of users or toggles them without redeploying. De-risks releases and enables experiments.Pricing & free-tier limitsFree: MLflow OSS $0 unlimited; W&B Free: 100GB, 3 team members, 10M logged steps. Cheaper: Neptune Free 200GB, W&B Teams Starter $50/mo 100GB + 10 users. Paid: W&B Teams $25/user/mo (min $250) + $0.50/GB over 250GB; Growth $50/user/mo. Databricks MLflow $0.15/DBU. Overage: W&B storage $0.50/GB/mo over, tracked hours $0.10/hr
Distributed Training Framework (Ray, NCCL, Deepspeed)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFree: DDP/DeepSpeed/Ray OSS $0 self-host. SkyPilot OSS free launcher. Cheaper: Modal distributed jobs $30 credits included then $1.50/hr. Paid: Anyscale Enterprise $400/mo + $0.40/hr worker overhead, CoreWeave IB fabric $0.30/hr per GPU IB addon. Overage: Cross-node NCCL traffic on-demand $0.02/GB on cloud, Anyscale extra workers $0.45/hr
Container Registry (GPU Optimized, Multi-arch)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsFree: GHCR 500MB storage free, 1GB transfer free public unlimited; Docker Hub free 1 private repo. Harbor OSS $0 self-host. Cheaper: Quay free 1 org, ECR free 500MB. Paid: ECR $0.10/GB/mo storage, $0.09/GB egress outside AWS; GAR $0.10/GB/mo; Docker Pro $24/mo 50GB. Overage: ECR storage $0.10/GB/mo, egress $0.09/GB over 1GB free, API $0.60/M
Container Security & Image Scanning
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsFree: Trivy/Grype OSS unlimited scans $0, Scout free 3 repos. Cheaper: Anchore free 10 images. Paid: Snyk $25/dev/mo container scanning, 200 images/mo; Wiz $0.30/container/mo min $500/mo; Aqua $1/container/mo. Overage: Snyk $0.15/image scan over 200, Wiz $0.40/container over, vuln DB $20/mo
CI/CD for Model & Infra Deployment (GPU builds)
Automates testing and deploying your code. Saves enormous time and prevents “works on my machine” releases.Pricing & free-tier limitsFree: GHA Free 2000 min/mo (2-core), ArgoCD $0 self-host. Cheaper: GitLab Free 400 min GPU (shared runners). Paid: GitHub Team $4/user/mo + GPU runner self-host cost; Buildkite $15/user/mo unlimited minutes + Own GPU runner $50/mo; CircleCI GPU $0.08/credit (A100 ~2c/min). Overage: GHA $0.008/min Linux, GPU self-host $1.50/hr build, ArgoCD support $2k/mo
GPU Monitoring (Utilization, DCGM, Health, Xid)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsFree: DCGM/Prom/Grafana OSS $0 self-host unlimited metrics. Cheaper: Netdata Cloud free 5 nodes. Paid: Datadog Infrastructure Pro $15/host/mo + custom metrics $0.05/100 metrics, GPU monitoring included; New Relic $0.30/GB ingest; Run:ai $150/GPU/mo monitoring. Overage: Datadog $0.30/GB logs, $0.05/100 custom metrics over 100, APM $31/host/mo
Inference Observability (Latency, Tokens/sec, Traces)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsFree: Langfuse OSS $0 self-host unlimited; Cloud Free 50k traces/mo. Helicone Free 10k requests/mo, 30d retention. Cheaper: Langfuse Cloud $39/mo 100k traces, Helicone Pro $20/mo 250k req. Paid: Langfuse Pro $99/mo 500k traces, $0.30/1k over; Helicone Teams $199/mo 5M req, $0.05/1k over. Overage: Langfuse $0.40/1k traces over quota, 90d retention +$50/mo
How this small scale infrastructure checklist works.
Each requirement below is something a small scale infrastructure build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.
Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.