Generative AI — small scale stack.
For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.
GPU Inference Compute (A10G / A100 / H100)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsModal: Free $30/mo credit, A10G $0.32/hr ($0.000089/sec), A100 40GB $1.60/hr, H100 $3.95/hr, 0s scale, concurrency auto. Colab: Free 1x T4 (15GB) ~12hr limit, 0$/mo. RunPod: Serverless $0.00021/sec for 24GB (A10G ~ $0.45/hr), $0.00058/sec H100 ~ $2.10/hr, free $5 credit. AWS p5.48xlarge 8x H100 $98.32/hr (~$12.29/hr per GPU). Overage: billed per 100ms.
Inference Hosting Platform (Serverless GPU)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsTogether AI: Free $25 credit, Llama 3.1 70B $0.88/1M input $0.88/1M output, FLUX.1 schnell $0.0023/image, autoscale included. HF Free: Rate limit ~1k req/day for public models, no GPU guarantee. Replicate: Pay-per-second, Llama 3 70B $0.65/1M tokens, SDXL $0.0023/sec (~$0.00045/run), free $5 credit. Anyscale: $0/mo control plane + $0.15/1M tokens managed + underlying GPU $1.5-4/hr; free trial $10.
Model Hosting & Model Registry
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsHF Endpoints: CPU $0.06/hr, T4 $0.60/hr, A10G $0.90/hr, A100 $1.64/hr, H100 $4.1/hr billed per minute; Free Hub: unlimited public models, 100GB storage free. Replicate: free registry, pay per run as above. Baseten: Free $25/mo credit, then Dedicated $0.08/hr overhead + GPUs (A10G $0.85/hr, H100 $5/hr), TRT-LLM optimized, free 100k inference calls tier deprecated, now usage-based.
LLM API Gateway & Multi-Provider Router
Controls and protects your APIs — quotas, abuse prevention, and firewalls. Important once you have real traffic or many clients.Pricing & free-tier limitsPortkey: Free 10k requests/mo, Growth $49/mo includes 50k req + $0.0005/req overage, fallbacks/retries/balancing. LiteLLM: Free OSS self-host, pay infra $5-20/mo VPS, tracks OpenAI $0.005/1k GPT-4o input, $0.015/1k output; Anthropic Claude 3.5 Sonnet $3/1M input $15/1M output. Cloudflare AI Gateway: Free 100k logs/day, $5/mo Workers + $0.50/million requests over. Zuplo: Free 10k req/mo, $250/mo for 1M req, overage $0.0002/req.
Vector Database for RAG (Prompt Context)
Persistent storage for your app’s data — users, products, orders. The single most important architectural decision for most projects.Pricing & free-tier limitsQdrant Free: 1GB RAM, up to 4M vectors x768d (approx), 1 cluster free forever; Starter $25/mo 2GB, $50/mo 4GB; Overage $0.20/GB storage. Weaviate Serverless Free: 14-day trial + free sandbox 1M vectors; Paid $25/mo starter. Pinecone Free: 100k vectors (~400MB), 1 index; Starter $70/mo includes writes $0.20/million + reads $2/1M + storage $0.33/GB; gcp-starter p2. Supabase: Free 500MB DB, pgvector extension free, $25/mo Pro 8GB.
File Storage for Generated Media (S3-Compatible)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsR2 Free: 10GB storage, 10M Class A + 10M Class B ops/mo, 0 egress fee. Paid $0.015/GB-mo storage, $4.50/million Class A, $0.36/million Class B, zero egress. B2 Free 10GB, $0.006/GB-mo, download $0.01/GB, first 1GB/day free egress. S3 Free 5GB for 12mo only, then $0.023/GB-mo, PUT $0.005/1k, GET $0.0004/1k, egress $0.09/GB first 10TB.
Media Delivery CDN (Images/Video/Audio)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsCloudflare Free: Unlimited bandwidth, 100k cache purge/day, global 300+ PoPs, $0/mo. Pro $20/mo adds WAF + image optimization 5k. BunnyCDN Free 14-day trial then $1/mo min, $0.01/GB NA/EU egress, $0.03 SG, free SSL. Fastly: $0/mo dev $50/mo minimum, $0.12/GB first 10TB, $0.02/10k req overage, free $500 credit first month.
Background Jobs & Async Generation Queue
Runs slow work (emails, exports, payouts) in the background so users don’t wait. Keeps your app snappy.Pricing & free-tier limitsInngest Free: 50k step runs/mo, 1k concurrency, 7-day history. Growth $49/mo 250k runs + $0.16/1k overage, 1yr retention. Trigger.dev Free: 10k tasks/mo self-host unlimited, Cloud $29/mo 50k tasks $0.0006/task over. QStash Free 500 msgs/day, $10/mo 2k/day + $0.02/100 msgs over. SQS: 1M free/mo forever, then $0.40/million req + Step State $0.025/1k transitions.
API Gateway + Rate Limiting for Gen-API
Controls and protects your APIs — quotas, abuse prevention, and firewalls. Important once you have real traffic or many clients.Pricing & free-tier limitsUpstash Rate Limit Free: 10k requests/day (Global). Paid $10/mo 100k/day, overage $0.20/100k. Unkey: Free 2.5k verifications/day, Pro $25/mo 150k + $0.05/1k over. Cloudflare API Shield: Free 1M gateway req, $20/mo Pro + $0.03/10k over. AWS API Gateway: Free 1M calls first 12mo, then $3.50/million + $0.09/GB egress.
Usage Metering (Tokens/Images/Seconds)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsOpenMeter OSS: Free self-host, unlimited events, metered billing aggregation. Cloud Free 1M events/mo, Pro $250/mo 10M events, $25/million over. Lago OSS Free 100M events self-host, Cloud Free 250k events, Starter $199/mo 1M events + $0.15/1k metered. Metronome: $2k/mo platform fee minimum + 0.5% of metered revenue, free proof-of-concept tier.
Billing & Subscriptions (Usage-Based)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsStripe Billing: Free 0.7% of recurring volume + 0.5% metered billing fee; Stripe fees 2.9% + $0.30 per transaction. Free tier: first $1M billing waived? Billing itself free until $100k processed. Lago OSS free, Cloud Free dev, Pro $199/mo as above. Polar: 4% + $0.30 per trans + billing included, free tier unlimited products. Stripe Tax +0.5% per transaction.
Caching (Prompt Cache + KV + Semantic Cache)
Stores frequent results in fast memory so you serve less from the database. Big lever for speed and cost.Pricing & free-tier limitsUpstash Free: 10k commands/day, 256MB, TLS, global replication. Pay-as-you-go $0.20/100k commands + $0.25/GB storage; $10/mo fixed 100MB. Redis Cloud Free 30MB; Essentials $5/mo 250MB + $0.20/100k ops. ElastiCache Serverless: Free trial 1mo, then $0.125/GB-hour + $0.14/million ECPUs. Semantic cache hit saves ~80% LLM cost $0.005/1k tokens.
Webhooks for Generation Completion
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsSvix Cloud Free: 50 messages/day, 5 apps, 7-day retention. Growth $59/mo 50k messages + $0.0002/msg over, 90-day retention, retries 72h. Hookdeck Free 500 events/day, 3d retention; Team $99/mo 50k events + 15d, over $0.001/event. Self-host Svix OSS free unlimited, infra cost ~$10/mo.
Content Moderation API (Generated Content)
Manages the content and assets you publish without editing code. Essential for sites with frequent updates.Pricing & free-tier limitsOpenAI Moderation: Free unlimited up to rate limit, 6 categories. Hive Free 2k images/mo, then $0.0008/image ($0.80/1k) for visual, $0.0005/1k chars text. Sightengine Free 2k ops/mo, Starter $29/mo 10k ops + $0.002/op over. AWS Rekognition Image Moderation: Free 5k/mo first yr, then $1/1k images over 1M $0.80/1k.
NSFW Detection & Filtering (Self-Host Option)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsNudeNet OSS: Free Apache2, run on CPU ~100ms/image, infra cost $0 (self-host on existing GPU box). LAION Safety Checker free, ~0.04s CLIP based. Sightengine $29/mo bundle includes NSFW 10k imgs, over $0.002/image. Hive NSFW $0.0006/image volume. Self-host savings: ~$800/1M images vs API.
Image/Video Optimization & Transform CDN
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsCloudflare Images Free: 5k delivered/mo? Actually $5/mo includes 100k stored + 100k delivered, $5/100k extra, transforms $0.50/1k. Bunny Optimizer $9/mo min + $20/1M transformations. Imgproxy OSS free self-host $5-15/mo compute. Cloudinary Free 25k transformations/mo, 25GB storage, bandwidth 25GB; Growth $89/mo 30k transforms + $0.14/1k over, storage $0.20/GB.
Monitoring for Inference Latency & Errors (APM)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsHelicone Free: 10k LLM requests/mo, 7-day retention, caching included. Pro $25/mo 1M req, $0.01/1k over, 30-day log. Grafana Cloud Free: 50GB logs, 50GB traces, 10k metrics forever. Datadog APM: $31/host/mo, LLM Observability addon $0.002/trace ($2/1k traces), free 14-day trial. Langfuse Free Cloud 50k obs/mo.
Logging & Request Tracing (LLM Observability)
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsLangfuse Free: 50k observations/mo (~50k LLM calls), 2 members, 30-day retention. Cloud Pro $39/mo 100k obs + $0.08/1k over, unlimited members. Axiom Free 500GB ingest/mo? Actually 500MB free? Correct: Free 500MB/mo ingest, $25/mo 10GB + $0.25/GB over. LangSmith: Free 5k traces/mo, Plus $39/mo 10k traces + $0.50/1k over.
Prompt Management & Versioning (PromptOps)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsLangfuse Prompts: Included free 50k obs, versioning, rollout AB. Pro same as above. Portkey Prompts: Free 10k requests includes prompts, Growth $49/mo 50k requests includes 100 prompt versions. Humanloop Free 5 users 10k logs/mo, Pro $199/mo 100k logs. LangSmith Hub: Free public prompts, private requires Plus $39/mo.
How this small scale generative ai checklist works.
Each requirement below is something a small scale generative ai build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.
Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.