Speech And Conversational — professional stack.
For each requirement below, pick the option that fits your build — recommended first, then free and cheaper alternatives — or skip what your project doesn't need. Tap the info icon next to any requirement to see why it matters.
Audio Object Storage (Recordings, Prompts, Datasets)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsS3: Free 5GB 12mo, then $0.023/GB/mo, $0.0004/1k PUT, $0.0004/10k GET, egress $0.09/GB. R2: Free 10GB storage, 10M Class A, 10M Class B, ZERO egress, then $0.015/GB. B2: Free 10GB, then $0.006/GB/mo, egress $0.01/GB (3x free). Overage: S3 requests $0.005/1k.
Speech-to-Text (STT) API / Hosting
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsDeepgram: $200 free credits (~46k mins), then $0.0043/min pre-recorded, $0.0048/min streaming, $0.004/min Nova-2. Whisper API: $0.006/min. AssemblyAI: Free 5 hrs ($50), then $0.00025/sec ($0.015/min), $0.015/min extra speaker labels. Self-hosted: ~$0.0008/min on A10G ($0.75/hr).
Text-to-Speech (TTS) API / Hosting
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsElevenLabs: Free 10k chars/mo (~10 mins), Starter $5/mo 30k chars, Creator $22/mo 100k, Pro $99/mo 500k, $0.30/1k overage. AWS Polly: Free 12mo 5M chars, then $16/1M neural chars. Self-hosted Coqui: $0 GPU if local, A10G $0.75/hr = ~$0.001/1k chars. Cartesia: $0.015/1k chars ultra-low latency 40ms.
GPU Compute for Speech Inference (Whisper/TTS)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsModal: Free $30/mo credits, A10G $0.000306/sec ($1.10/hr), A100 40GB $0.00064/sec ($2.30/hr), H100 $0.00122/sec ($4.39/hr) autoscale to 0. RunPod: Free $5 trial, A10G $0.38/hr, A100 $1.19/hr, H100 $2.29/hr spot. Vast.ai: A10G $0.26/hr, A100 $0.99/hr. AWS p4d: A100 $3.06/hr on-demand, H100 p5 $5.98/hr. Overages billed per-second.
Real-time Voice Infra (WebRTC, Voice Agents)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsLiveKit Cloud: Free 50GB/mo + 50k mins, then $0.004/min audio SFU + $0.015/participant/min AI agent. Daily.co: Free 10k mins, then $0.004/min + $0.01/min transcription. Cloudflare Calls: Free 1k mins, then $0.01/min + $5/1k MAU. Agora: Free 10k mins, then $0.00099/min voice, $0.015/min AI agent. Self-hosted LiveKit: $5-20/mo VM cost.
Conversation State Management (Dialog Memory)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsUpstash: Free 10k cmds/day, 256MB, then $0.20/100k cmds, $0.25/GB storage. Redis Cloud: Free 30MB, then Essentials $5/mo 250MB, $0.15/GB. ElastiCache Serverless: Free tier none, $0.125/GB-hr storage + $0.20/1M ECPUs. Self-hosted: $0 + $6/mo VPS.
Vector Database for Conversation Memory & RAG
Persistent storage for your app’s data — users, products, orders. The single most important architectural decision for most projects.Pricing & free-tier limitsPinecone: Free 100k vectors (2GB), then Starter $70/mo 2M vectors, $0.33/1M reads, $0.07/1k writes. Qdrant: Free 1GB cluster, then $0.25/GB + $0.05/M op, 25$/mo for 4GB. Supabase: Free 500MB pgvector, then $25/mo 8GB. Self-hosted Qdrant: $0 + $10/mo VM for 2M vectors.
LLM Inference for Dialogue Management / Agent Reasoning
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsGroq: Free 14.4k tokens/sec limit, $0.59/M input, $0.79/M output (Llama 70B). Together: Free $25 credits, $0.88/M in/out Llama 70B. OpenAI GPT-4o: $0.0025/1k input, $0.01/1k output, Free $5 trial. Claude 3.5 Sonnet: $0.003/1k in, $0.015/1k out. Anyscale: $0.15/M input Llama 70B, $1/min overage. Overage: Groq 5k req/min limit.
Background Jobs for Async Transcription & TTS
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsInngest: Free 50k steps/mo, 30k events, then $20/mo 500k steps, $0.10/1k overage. Trigger.dev: Free 10k tasks, then $30/mo 100k, $0.20/1k overage. QStash: Free 500 msg/day, then $10/mo 40k msgs, $0.30/10k. AWS Step Functions: Free 4k transitions, then $0.025/1k transitions.
Queue for Audio Processing Jobs (Durable)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsSQS: Free 1M req/mo, then $0.40/1M req, FIFO $0.50/1M. Cloudflare Queues: Free 1M ops, then $0.50/M ops. QStash: Free 500/day, $10/mo 40k. Confluent Kafka: Free $400 trial, then $0.20/GB in/out, $0.001/vCPU-hr. RabbitMQ Cloud: Free 100 queues, then Little Lemur $19/mo.
API Gateway + Rate Limiting for Voice API
Controls and protects your APIs — quotas, abuse prevention, and firewalls. Important once you have real traffic or many clients.Pricing & free-tier limitsCloudflare Workers: Free 100k req/day, 10ms CPU, then $5/mo 10M req inclusive, $0.30/M overage. Upstash Rate Limit: Free 10k/day, $0.20/100k. AWS API Gateway: Free 12mo 1M calls, then $3.50/M REST, $1/M HTTP API. Kong Konnect: $250/mo platform + $0.02/1k req.
Telephony API (PSTN / SIP for Voice Agents)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsTwilio: Free $15 trial, US inbound $0.0085/min, outbound $0.013/min, number $1/mo, SIP $0.004/min. Telnyx: Free $10 credit, $0.005/min inbound/outbound US, numbers $1/mo, $0.003/min SIP. Plivo: Free trial, $0.005/min in/out. Overages: Twilio transcription add-on $0.05/min.
Usage Metering (Audio Minutes, Characters, Tokens)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsOpenMeter Cloud: Free 1M events/mo, then $25/mo 10M events, $0.02/10k overage. Lago: Free self-hosted, Cloud $49/mo 50k events. Stripe: 0.8% revenue + $0.15/bill. Orb: Free $10k revenue, then 0.5-1% usage. Metronome: Custom 0.5% + $500/mo min. Self-hosted OpenMeter: $0 + Postgres cost.
Monitoring & Observability for ASR/TTS Latency
Tells you when things break and why. You cannot fix what you cannot see — this is how you keep downtime short.Pricing & free-tier limitsLangfuse: Free 50k observations/mo self-host, Cloud Free 50k/mo, Pro $39/mo 250k, $0.25/1k overage. Helicone: Free 10k req/mo, Pro $20/mo 100k, $15/100k overage. Grafana Cloud: Free 10k metrics, 50GB logs, then $8/mo user + $0.50/1k metrics. Datadog: Free 5 hosts, APM $31/host/mo, logs $0.10/GB, overage $0.15/GB.
Inference Hosting Platform (Serverless GPU for Whisper/TTS/LLM)
Where your code actually runs and serves requests. Picking the right host affects speed, scaling, and how much ops work you do.Pricing & free-tier limitsModal: $30 free, per-second billing, A100 $2.30/hr. Replicate: Free $10 credits, then RTX 4090 $0.000725/sec ($2.61/hr), A100 $0.0013/sec. RunPod Serverless: Free $5, $0.00029/sec ($1.04/hr) active, $0.0001/sec idle. Baseten: $30/mo + GPU $1.5-4/hr. Anyscale: Free $10, $1/mo + $0.15/M tokens Llama.
CDN for Audio Delivery (TTS playback, recordings)
Where you keep files users upload or you serve — images, videos, documents — and how fast they reach visitors around the world.Pricing & free-tier limitsCloudflare: Free unlimited egress from R2, Workers 100k/day. Bunny.net: Free 14-day, then $0.01/GB EU/US, $0.03/GB Asia, storage $0.01/GB. CloudFront: Free 1TB 12mo, then $0.085/GB first 10TB, $0.02/10k req. Fastly: Free $50 trial, $0.12/GB + $0.009/10k req min $50/mo. R2 overage: $0.015/GB after 10GB free.
Authentication & Authorization for Voice API
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsClerk: Free 10k MAU, then $25/mo + $0.02/MAU beyond 10k, $0.50/1k verifications. Supabase Auth: Free 50k MAU, then $25/mo 100k MAU. Auth0: Free 25k MAU, Essentials $35/mo 500 external, $0.07/MAU overage. WorkOS: Free 100 MAU, then $125/mo + $0.05/MAU.
Logging & Analytics for Conversations
Understands your users and supports them. Drives product decisions and retention.Pricing & free-tier limitsAxiom: Free 500MB/mo ingest, Pro $25/mo 10GB, $0.25/GB overage, 1yr retention. PostHog: Free 1M events/mo, 5k survey responses, then $0.000248/event. Better Stack Logtail: Free 1GB/mo, then $24/mo 20GB, $0.25/GB. Datadog Logs: $0.10/GB ingest, $0.70/GB indexed, overage $0.12/GB.
Secrets Management (API Keys, Voice Credentials)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsDoppler: Free 25 secrets, 5 users, then $20/mo/user Pro, Team $10/user. Infisical: Free 10k secrets, unlimited users self-host, Cloud Starter $19/mo. AWS Secrets: Free 30 days, then $0.40/secret/mo + $0.05/10k API. Vault Cloud: Free 25 secrets, $0.03/hr + $0.60/secret. Overages: Doppler $0.50/extra secret beyond.
Billing & Monetization (Per-minute, per-char billing)
How you monetize with advertising. Relevant for content and media businesses.Pricing & free-tier limitsStripe Billing: Free 0.8% on recurring + $0.15/invoice, metered $0.15/report. Lago: OSS free, Cloud Starter $49/mo 1k customers + $0.01/invoice. Orb: Free <10k revenue/mo, then 0.5-0.8% + $550/mo platform. Paddle: 5% + $0.50/transaction includes tax. Overages: Stripe $0.50/additional metered event >100k.
Model Registry & Versioning (Custom STT/TTS models)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsHF Hub: Free unlimited public, Free 1 private model 10GB, Pro $9/mo 100GB private, Enterprise $20/user/mo. MLflow self-hosted: $0 + S3 $0.023/GB. W&B Artifacts: Free 100GB, Pro $50/mo 500GB, $0.15/GB overage. Neptune: Free 100GB-hr, $49/mo 500GB-hr.
Evaluation & Testing for Speech (WER, MOS, LLM-as-judge)
Powers AI features — model access, embeddings, and inference. Costs scale with usage, so watch the meter.Pricing & free-tier limitsLangSmith: Free 5k traces/mo, Starter $39/mo 10k traces, Plus $149/mo 100k. Braintrust: Free 1M tokens evaluated/mo, Pro $50/mo 5M, $0.008/1k eval tokens overage. Ragas OSS self-host: $0 + LLM eval cost ~$0.01/eval (GPT-4). Weave W&B: Free 10k traces, $25/mo 100k.
Speaker Diarization & Identification
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsPyannote AI API: Free 10 hrs, then $0.05/min ($3/hr) diarization, $99/mo 100 hrs. Deepgram Diarization: +$0.001/min add-on ( $0.0053/min total). AssemblyAI: Speaker labels +$0.003/min. Self-hosted pyannote: $0 + A10G $0.38/hr = $0.006/min. WhisperX: OSS free.
Webhooks Engine for Transcription Completion & Events
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsSvix: Free 50 messages/day + 7 day retention, $99/mo 1M messages ($0.0001/msg), $0.10/1k overage. Hookdeck: Free 100 events/day, Pro $35/mo 100k events, $0.25/10k overage. QStash: Free 500/day, $10/mo 40k. EventBridge: Free 100 custom events/mo, $1/M events.
Content Moderation for Voice (Toxicity, PII)
Manages the content and assets you publish without editing code. Essential for sites with frequent updates.Pricing & free-tier limitsOpenAI Moderation: Free unlimited (rate-limited), no token cost. Hive: Free 1k calls/mo, then $0.001/audio sec ($0.06/min). AWS Comprehend: Free 50k units (12mo), $0.0001/unit + PII $0.0001/unit. Clarifai: Free 1k ops, $30/mo 5k ops, $0.003/op overage. Hume AI: Free $10 credit, $0.015/min moderation.
Voice Activity Detection (VAD) & Noise Suppression
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsSilero VAD OSS: $0, 0.5ms latency CPU, MIT license. LiveKit Krisp: Free trial, $0.003/min/participant noise suppression. Krisp SDK: $0.02/min per stream enterprise, Free eval 30 days. Dolby.io: Free 50 hrs, $0.05/min enhance, $0.003/min VAD. RNNoise: Free OSS.
Container Orchestration (Voice microservices)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsFly.io: Free 3x shared-cpu, 3GB vol, then $1.94/mo/GB + $0.0000016/sec CPU, $5 min. AWS Fargate: $0.040/vCPU/hr + $0.00438/GB/hr. Cloud Run: Free 2M req, then $0.000024/vCPU-sec + $0.0000025/GB-sec, 360k vCPU-sec free. Render: Free 750 hrs, Starter $7/mo. GKE Autopilot: $0.10/hr cluster + $0.044/vCPU/hr.
Training & Fine-tuning Infra (Custom STT/TTS)
A piece of your stack you may or may not need, depending on scope. Pick the option that fits — or skip it if your project doesn’t require this capability yet.Pricing & free-tier limitsLambda Cloud: Free $2 credit, A10 $0.75/hr, A100 40GB $1.29/hr, H100 $2.49/hr (8x $19.92/hr). Colab Pro: Free T4, Pro $9.99/mo 100 compute units, Pro+ $49.99/mo 500 units. Kaggle: Free 30 hrs T4x2/week. SageMaker: ml.p4d $7.42/hr A100 40GB, ml.p5 $20.7/hr H100. Modal Training: Same GPU + per-sec, $30 free.
Experiment Tracking (WER, CER, MOS experiments)
Ships features to a subset of users or toggles them without redeploying. De-risks releases and enables experiments.Pricing & free-tier limitsW&B: Free 100GB storage, 1 team member, 10k traces, then Team $50/mo 500GB + $0.15/GB, $20/user overage. MLflow self-host: $0 + $5/mo VPS + S3 $0.023/GB. Comet: Free 1 project 100GB, Community $49/mo. Neptune: Free 100GB-hr, $49/mo 500GB-hr.
CI/CD for Models & Voice APIs
Automates testing and deploying your code. Saves enormous time and prevents “works on my machine” releases.Pricing & free-tier limitsGitHub Actions: Free 2k Linux mins/mo, 500MB storage, then $0.008/min, $0.25/GB storage overage. CircleCI: Free 6k credits/mo, then $15/mo 30k credits, $0.001/credit overage. Buildkite: Free 100 mins, $15/agent/mo. Harness CI: Free 1k builds, then $60/mo 2k builds.
How this professional speech and conversational checklist works.
Each requirement below is something a professional speech and conversational build typically needs. Pick one of the four researched options — recommended, free, cheaper or paid — add your own with "Other", or skip the requirement if your project doesn't need it. Nothing is mandatory; the plan on the right tracks what you've decided so nothing gets forgotten.
Your picks are saved in this browser automatically, so you can come back anytime. Options are researched per build level and refreshed as vendors change their plans — always verify details on the provider's page before committing.