Stack Cost AI

Generative AIsmall scale stack.

For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.

GPU Inference Compute (A10G / A100 / H100)

Pricing & free-tier limitsModal: Free $30/mo credit, A10G $0.32/hr ($0.000089/sec), A100 40GB $1.60/hr, H100 $3.95/hr, 0s scale, concurrency auto. Colab: Free 1x T4 (15GB) ~12hr limit, 0$/mo. RunPod: Serverless $0.00021/sec for 24GB (A10G ~ $0.45/hr), $0.00058/sec H100 ~ $2.10/hr, free $5 credit. AWS p5.48xlarge 8x H100 $98.32/hr (~$12.29/hr per GPU). Overage: billed per 100ms.

Inference Hosting Platform (Serverless GPU)

Pricing & free-tier limitsTogether AI: Free $25 credit, Llama 3.1 70B $0.88/1M input $0.88/1M output, FLUX.1 schnell $0.0023/image, autoscale included. HF Free: Rate limit ~1k req/day for public models, no GPU guarantee. Replicate: Pay-per-second, Llama 3 70B $0.65/1M tokens, SDXL $0.0023/sec (~$0.00045/run), free $5 credit. Anyscale: $0/mo control plane + $0.15/1M tokens managed + underlying GPU $1.5-4/hr; free trial $10.

Model Hosting & Model Registry

Pricing & free-tier limitsHF Endpoints: CPU $0.06/hr, T4 $0.60/hr, A10G $0.90/hr, A100 $1.64/hr, H100 $4.1/hr billed per minute; Free Hub: unlimited public models, 100GB storage free. Replicate: free registry, pay per run as above. Baseten: Free $25/mo credit, then Dedicated $0.08/hr overhead + GPUs (A10G $0.85/hr, H100 $5/hr), TRT-LLM optimized, free 100k inference calls tier deprecated, now usage-based.

LLM API Gateway & Multi-Provider Router

Pricing & free-tier limitsPortkey: Free 10k requests/mo, Growth $49/mo includes 50k req + $0.0005/req overage, fallbacks/retries/balancing. LiteLLM: Free OSS self-host, pay infra $5-20/mo VPS, tracks OpenAI $0.005/1k GPT-4o input, $0.015/1k output; Anthropic Claude 3.5 Sonnet $3/1M input $15/1M output. Cloudflare AI Gateway: Free 100k logs/day, $5/mo Workers + $0.50/million requests over. Zuplo: Free 10k req/mo, $250/mo for 1M req, overage $0.0002/req.

Vector Database for RAG (Prompt Context)

Pricing & free-tier limitsQdrant Free: 1GB RAM, up to 4M vectors x768d (approx), 1 cluster free forever; Starter $25/mo 2GB, $50/mo 4GB; Overage $0.20/GB storage. Weaviate Serverless Free: 14-day trial + free sandbox 1M vectors; Paid $25/mo starter. Pinecone Free: 100k vectors (~400MB), 1 index; Starter $70/mo includes writes $0.20/million + reads $2/1M + storage $0.33/GB; gcp-starter p2. Supabase: Free 500MB DB, pgvector extension free, $25/mo Pro 8GB.

File Storage for Generated Media (S3-Compatible)

Pricing & free-tier limitsR2 Free: 10GB storage, 10M Class A + 10M Class B ops/mo, 0 egress fee. Paid $0.015/GB-mo storage, $4.50/million Class A, $0.36/million Class B, zero egress. B2 Free 10GB, $0.006/GB-mo, download $0.01/GB, first 1GB/day free egress. S3 Free 5GB for 12mo only, then $0.023/GB-mo, PUT $0.005/1k, GET $0.0004/1k, egress $0.09/GB first 10TB.

Media Delivery CDN (Images/Video/Audio)

Pricing & free-tier limitsCloudflare Free: Unlimited bandwidth, 100k cache purge/day, global 300+ PoPs, $0/mo. Pro $20/mo adds WAF + image optimization 5k. BunnyCDN Free 14-day trial then $1/mo min, $0.01/GB NA/EU egress, $0.03 SG, free SSL. Fastly: $0/mo dev $50/mo minimum, $0.12/GB first 10TB, $0.02/10k req overage, free $500 credit first month.

Background Jobs & Async Generation Queue

Pricing & free-tier limitsInngest Free: 50k step runs/mo, 1k concurrency, 7-day history. Growth $49/mo 250k runs + $0.16/1k overage, 1yr retention. Trigger.dev Free: 10k tasks/mo self-host unlimited, Cloud $29/mo 50k tasks $0.0006/task over. QStash Free 500 msgs/day, $10/mo 2k/day + $0.02/100 msgs over. SQS: 1M free/mo forever, then $0.40/million req + Step State $0.025/1k transitions.

API Gateway + Rate Limiting for Gen-API

Pricing & free-tier limitsUpstash Rate Limit Free: 10k requests/day (Global). Paid $10/mo 100k/day, overage $0.20/100k. Unkey: Free 2.5k verifications/day, Pro $25/mo 150k + $0.05/1k over. Cloudflare API Shield: Free 1M gateway req, $20/mo Pro + $0.03/10k over. AWS API Gateway: Free 1M calls first 12mo, then $3.50/million + $0.09/GB egress.

Usage Metering (Tokens/Images/Seconds)

Pricing & free-tier limitsOpenMeter OSS: Free self-host, unlimited events, metered billing aggregation. Cloud Free 1M events/mo, Pro $250/mo 10M events, $25/million over. Lago OSS Free 100M events self-host, Cloud Free 250k events, Starter $199/mo 1M events + $0.15/1k metered. Metronome: $2k/mo platform fee minimum + 0.5% of metered revenue, free proof-of-concept tier.

Billing & Subscriptions (Usage-Based)

Pricing & free-tier limitsStripe Billing: Free 0.7% of recurring volume + 0.5% metered billing fee; Stripe fees 2.9% + $0.30 per transaction. Free tier: first $1M billing waived? Billing itself free until $100k processed. Lago OSS free, Cloud Free dev, Pro $199/mo as above. Polar: 4% + $0.30 per trans + billing included, free tier unlimited products. Stripe Tax +0.5% per transaction.

Caching (Prompt Cache + KV + Semantic Cache)

Pricing & free-tier limitsUpstash Free: 10k commands/day, 256MB, TLS, global replication. Pay-as-you-go $0.20/100k commands + $0.25/GB storage; $10/mo fixed 100MB. Redis Cloud Free 30MB; Essentials $5/mo 250MB + $0.20/100k ops. ElastiCache Serverless: Free trial 1mo, then $0.125/GB-hour + $0.14/million ECPUs. Semantic cache hit saves ~80% LLM cost $0.005/1k tokens.

Webhooks for Generation Completion

Pricing & free-tier limitsSvix Cloud Free: 50 messages/day, 5 apps, 7-day retention. Growth $59/mo 50k messages + $0.0002/msg over, 90-day retention, retries 72h. Hookdeck Free 500 events/day, 3d retention; Team $99/mo 50k events + 15d, over $0.001/event. Self-host Svix OSS free unlimited, infra cost ~$10/mo.

Content Moderation API (Generated Content)

Pricing & free-tier limitsOpenAI Moderation: Free unlimited up to rate limit, 6 categories. Hive Free 2k images/mo, then $0.0008/image ($0.80/1k) for visual, $0.0005/1k chars text. Sightengine Free 2k ops/mo, Starter $29/mo 10k ops + $0.002/op over. AWS Rekognition Image Moderation: Free 5k/mo first yr, then $1/1k images over 1M $0.80/1k.

NSFW Detection & Filtering (Self-Host Option)

Pricing & free-tier limitsNudeNet OSS: Free Apache2, run on CPU ~100ms/image, infra cost $0 (self-host on existing GPU box). LAION Safety Checker free, ~0.04s CLIP based. Sightengine $29/mo bundle includes NSFW 10k imgs, over $0.002/image. Hive NSFW $0.0006/image volume. Self-host savings: ~$800/1M images vs API.

Image/Video Optimization & Transform CDN

Pricing & free-tier limitsCloudflare Images Free: 5k delivered/mo? Actually $5/mo includes 100k stored + 100k delivered, $5/100k extra, transforms $0.50/1k. Bunny Optimizer $9/mo min + $20/1M transformations. Imgproxy OSS free self-host $5-15/mo compute. Cloudinary Free 25k transformations/mo, 25GB storage, bandwidth 25GB; Growth $89/mo 30k transforms + $0.14/1k over, storage $0.20/GB.

Monitoring for Inference Latency & Errors (APM)

Pricing & free-tier limitsHelicone Free: 10k LLM requests/mo, 7-day retention, caching included. Pro $25/mo 1M req, $0.01/1k over, 30-day log. Grafana Cloud Free: 50GB logs, 50GB traces, 10k metrics forever. Datadog APM: $31/host/mo, LLM Observability addon $0.002/trace ($2/1k traces), free 14-day trial. Langfuse Free Cloud 50k obs/mo.

Logging & Request Tracing (LLM Observability)

Pricing & free-tier limitsLangfuse Free: 50k observations/mo (~50k LLM calls), 2 members, 30-day retention. Cloud Pro $39/mo 100k obs + $0.08/1k over, unlimited members. Axiom Free 500GB ingest/mo? Actually 500MB free? Correct: Free 500MB/mo ingest, $25/mo 10GB + $0.25/GB over. LangSmith: Free 5k traces/mo, Plus $39/mo 10k traces + $0.50/1k over.

Prompt Management & Versioning (PromptOps)

Pricing & free-tier limitsLangfuse Prompts: Included free 50k obs, versioning, rollout AB. Pro same as above. Portkey Prompts: Free 10k requests includes prompts, Growth $49/mo 50k requests includes 100 prompt versions. Humanloop Free 5 users 10k logs/mo, Pro $199/mo 100k logs. LangSmith Hub: Free public prompts, private requires Plus $39/mo.

Options and prices come straight from our research sheets for a small scale generative ai project. Prices are estimates and change often — always confirm on the provider's page before committing.

How this small scale generative ai checklist works.

Each requirement below is something a small scale generative ai build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.

Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.