Learn

Real-Time AI Voice API Pricing Explained

Real-time AI voice API pricing explained: compare native multimodal audio tokens against modular STT-LLM-TTS streaming pipelines, evaluate telephony SIP costs, and optimize live agent budgets.

Clarify the spend threshold before you commit. Use this page when the core product is familiar and the real question is whether to stay free, upgrade, or switch pricing tracks.

UpdatedSeptember 27, 2026
Browse tool profiles

Editorial guide

Guide

Start with the spend threshold and the conditions that change the pricing decision.

Real-time AI voice API pricing cannot be evaluated through a single per-character or per-minute rate. When moving from static text-to-speech batch rendering to interactive, conversational audio pipelines, pricing structures fracture across multiple decoupled meters: speech-to-text transcription latency, large language model inference and reasoning tokens, audio generation streaming rates, WebSocket and WebRTC session duration, concurrency reservations, telephony SIP interconnects, and managed agent orchestration fees. Evaluating vendors requires modeling the complete runtime conversation path rather than relying on isolated headline rates.

Engineering teams building customer service voicebots, sales assistants, interactive characters, and in-app voice agents must choose between three distinct architectural paradigms: integrated multimodal end-to-end APIs (such as OpenAI Realtime API), modular streaming pipelines (combining Deepgram or Whisper for transcription, Claude or GPT for reasoning, and Cartesia or ElevenLabs for synthesis), or fully managed voice agent orchestration platforms (such as Retell AI or Vapi). Each architecture shifts cost burdens across different meters, dramatically altering monthly operational expenses at scale.

The Decoupled Meters of Real-Time Voice Architectures

In a traditional batch TTS pipeline, pricing is straightforward: a developer sends a string of characters and pays a fixed rate per thousand or million characters. In real-time conversational voice, five distinct cost categories accumulate simultaneously during an ongoing session:

1. Inbound Audio Processing (Speech-to-Text / Audio Input)

When a user speaks into a microphone or telephone handset, their voice stream must be digitized, streamed over low-latency protocols, and converted into linguistic tokens. Providers bill this either by audio duration (for example, Deepgram Nova-2 charging $0.0043 per minute of incoming audio) or by raw multimodal input audio tokens (for example, OpenAI's gpt-realtime billing $32.00 per million audio input tokens, which translates to roughly $0.06 per minute of user speech).

2. Reasoning and Linguistic Generation (LLM Core)

Once transcribed, the conversation context is processed by an LLM to generate the assistant's reply. In modular architectures, developers pay standard text token rates (such as $2.00 per million input tokens and $10.00 per million output tokens for Claude Sonnet 5, or $0.10 / $0.50 for GPT-6 Luna). In native multimodal realtime models, audio tokens consume reasoning capacity directly, bypassing textual transcription but incurring higher base token unit costs.

3. Outbound Speech Synthesis (Text-to-Speech Streaming)

The generated response must be converted to audio with sub-second time-to-first-byte (TTFB) latency. Providers meter outbound voice through distinct units:

  • Cartesia Sonic: Charges by output audio duration or model credits, translating to approximately $0.015 to $0.020 per minute of synthesized speech with 90ms TTFB.
  • ElevenLabs Conversational AI & Flash v2.5: Billed per character generated, where Flash v2.5 provides a 50% credit discount compared to flagship Multilingual v2, averaging $0.020 to $0.035 per speech minute.
  • Deepgram Aura: Billed per thousand characters ($0.015 per 1,000 characters), averaging roughly $0.012 per minute of synthesized speech.

4. Transport, Connection State, and Concurrency

Maintaining continuous, bi-directional audio streaming requires WebSocket or WebRTC connections. Managed platforms frequently charge session-runtime fees (such as $0.05 to $0.08 per minute) to handle turn-taking arbitration, ambient noise suppression, endpoint detection, and speech interruption handling (barge-in). Furthermore, enterprise contracts frequently levy concurrency charges for reserved channels during peak call surges.

5. Telephony Ingestion and Termination

If voice agents operate over traditional public switched telephone networks (PSTN) rather than in-app WebRTC, telephony carrier surcharges apply. Services like Twilio, Telnyx, or Vonage bill between $0.0040 and $0.0130 per minute for inbound SIP trunking and $0.0140 to $0.0220 per minute for outbound calling, plus carrier connection fees and phone number leasing.

Architectural Component

Primary Metering Unit

Typical Unit Rate Range

Key Latency Drivers

Primary Cost Scaling Vector

Speech-to-Text (STT)

Audio stream minutes or seconds

$0.0043 – $0.0100 / min

Chunk size, VAD threshold, VAD model

Total call listening duration

LLM Reasoning (Text)

Input and output text tokens

$0.10 – $10.00 / 1M tokens

Context window size, system prompt

Turn count and conversational depth

LLM Reasoning (Audio)

Native input/output audio tokens

$32.00 in / $64.00 out / 1M

Multimodal attention over raw audio

Total session audio duration

Text-to-Speech (TTS)

Characters or synthesized audio seconds

$0.0120 – $0.0350 / min spoken

Model size, streaming buffer depth

Total agent speaking duration

Agent Orchestration

Active session connection minutes

$0.0500 – $0.0800 / min

Turn detection, interruption handling

Wall-clock session duration

Telephony SIP Trunk

Inbound/outbound carrier minutes

$0.0040 – $0.0220 / min

Carrier routing, codec negotiation

Wall-clock telephone connection

Multivendor Architectural Comparison: Three Core Approaches

Understanding how vendor pricing models translate into operational budgets requires analyzing the three primary architectural designs deployed in production:

Approach A: Integrated Native Multimodal Realtime API

OpenAI Realtime API represents the integrated approach. Audio is streamed directly into the model without an intermediate textual transcription layer, and the model synthesizes conversational audio responses natively. This eliminates transcription serialization latency, resulting in fluid human-like turn-taking. However, audio tokens are expensive: $32 per million input audio tokens and $64 per million output audio tokens on gpt-realtime. While text tokens cost $4 input / $16 output per million, streaming continuous background audio quickly consumes high-cost audio tokens unless aggressive Voice Activity Detection (VAD) suppresses silence.

Approach B: Modular Low-Latency Streaming Pipeline

In a modular pipeline, developers compose best-in-breed specialized microservices: Deepgram Nova-2 for ultra-fast STT (~200ms latency), a small fast LLM (such as GPT-6 Luna, Claude Haiku 4.5, or a fine-tuned open-source model) for reasoning, and Cartesia Sonic or ElevenLabs Flash v2.5 for streaming TTS. The developer manages orchestration and WebSocket transport via self-hosted infrastructure (LiveKit, Asterisk, or Node.js). This approach yields the lowest raw infrastructure unit cost per minute, but requires significant internal engineering to handle interruption state machines, jitter buffers, and telephony synchronization.

Approach C: Managed Voice Agent Platform (Vapi / Retell AI)

Managed platforms provide turnkey agent infrastructure. Developers configure prompt instructions, voice IDs, and tool-calling webhooks, while the platform coordinates STT, LLM, TTS, telephony, and interruption logic. These platforms bill an orchestration fee ($0.05 to $0.08 per call minute) on top of underlying vendor costs (or allow developers to bring their own API keys). While the per-minute cost is highest, development velocity is accelerated and infrastructure maintenance is offloaded.

Feature & Metric

Native Multimodal (OpenAI Realtime)

Modular Custom Pipeline (Self-Hosted)

Managed Platform (Vapi / Retell)

Architecture

Single unified end-to-end audio model

Chained microservices (STT + LLM + TTS)

Managed cloud orchestration layer

End-to-End Latency

300ms – 450ms

500ms – 750ms

600ms – 900ms

Base Orchestration Fee

$0.00 / min (included in token cost)

$0.00 / min (internal server cost)

$0.05 – $0.08 / min platform fee

Speech-to-Text Cost

Included in audio input token rate

$0.0043 / min (Deepgram Nova-2)

Passthrough or bundled ($0.005/min)

LLM Reasoning Cost

Multimodal tokens ($32/$64 per 1M)

Standard text tokens ($0.10/$0.50 per 1M on GPT-6 Luna)

BYO API key or bundled token rate

Text-to-Speech Cost

Included in audio output token rate

$0.0150 / min (Cartesia Sonic)

Passthrough ($0.015 – $0.030/min)

Telephony Integration

SIP bridging via third-party required

Twilio/Telnyx SIP trunk configured

Built-in one-click phone numbers

Interruption Handling

Native inside neural architecture

Custom client-side VAD state machine

Server-side barge-in algorithms

End-to-End Cost per Minute Simulation

To establish realistic budgetary benchmarks, we simulate a standard customer service phone call lasting exactly 5 minutes (300 seconds). In this typical call profile:

  • The human customer speaks for 2.0 minutes (40% of duration).
  • The AI assistant speaks for 2.0 minutes (40% of duration, generating approximately 300 words / 1,500 characters).
  • Silence, pauses, and turn-taking latency account for 1.0 minute (20% of duration).
  • Telephony operates over inbound SIP trunking ($0.005/min).

Modular Pipeline Cost Breakdown (5-Minute Call):

  • STT (Deepgram Nova-2, listening across 5 minutes): 5 min × $0.0043 = $0.0215
  • LLM Reasoning (GPT-6 Luna, 6 turns, ~4,000 input tokens + 600 output tokens): ~$0.0007
  • TTS (Cartesia Sonic, 2.0 minutes synthesized audio): 2 min × $0.0150 = $0.0300
  • Telephony Inbound (Twilio SIP): 5 min × $0.0050 = $0.0250
  • Self-Hosted Compute & WebSocket Egress: ~$0.0050
  • Total Call Cost: $0.0822 (Effective rate: $0.0164 per minute).

Managed Platform Breakdown (5-Minute Call via Vapi/Retell):

  • Platform Orchestration Fee: 5 min × $0.0500 = $0.2500
  • Underlying STT + LLM + TTS costs: ~$0.0525
  • Telephony Inbound: 5 min × $0.0050 = $0.0250
  • Total Call Cost: $0.3275 (Effective rate: $0.0655 per minute).

OpenAI Realtime API Breakdown (5-Minute Call):

  • Input Audio Tokens (including ambient background noise, ~25,000 audio tokens at $32 per million): ~$0.8000
  • Output Audio Tokens (~15,000 generated audio tokens at $64 per million): ~$0.9600
  • Context Window & Text Tokens: ~$0.0150
  • Third-Party WebRTC-to-SIP Bridge & Telephony: 5 min × $0.0100 = $0.0500
  • Total Call Cost: $1.8250 (Effective rate: $0.3650 per minute).

OpenAI's newer GPT-Live-1 voice model offers a different meter for integrated voice: $0.05 per minute of voice session, billed per second, plus the backend model and tool usage it delegates to.

Monthly Call Volume (5-min calls)

Modular Pipeline ($0.0164/min)

OpenAI Realtime API ($0.3650/min)

Managed Platform ($0.0655/min)

Variance (Modular vs Managed)

500 calls (2,500 minutes)

$41.00

$912.50

$163.75

Save $122.75 / month

5,000 calls (25,000 minutes)

$410.00

$9,125.00

$1,637.50

Save $1,227.50 / month

25,000 calls (125,000 minutes)

$2,050.00

$45,625.00

$8,187.50

Save $6,137.50 / month

100,000 calls (500,000 minutes)

$8,200.00

$182,500.00

$32,750.00

Save $24,550.00 / month

At 100,000 calls per month, the modular pipeline saves over $24,000 monthly compared to managed platforms, providing clear financial justification for engineering teams to invest in self-hosted LiveKit and custom pipeline orchestration.

Production Optimization Strategies

Organizations deploying real-time voice APIs can implement four engineering practices to significantly reduce unit costs without harming caller experience:

  1. Client-Side Voice Activity Detection (VAD): Never stream continuous unmuted audio to raw audio-token endpoints. Implementing aggressive local VAD cuts inbound audio transmission during customer listening pauses, eliminating up to 50% of unnecessary audio token billing.
  2. Streaming Text Buffering: In modular pipelines, do not wait for the LLM to complete its entire response before initiating TTS synthesis. Buffer the first 5 to 8 words and immediately initiate WebSocket audio streaming. This cuts TTFB below 500ms while keeping model generation on low-cost text models.
  3. Model Tier Switching by Interaction Stage: Use lightweight fast models (like GPT-6 Luna or Claude Haiku 4.5) for initial greeting and intent classification, escalating to heavier reasoning models only when complex transactional logic is required.
  4. SIP Trunk Direct Peering: Route telephone calls directly through private SIP interconnects rather than public carrier aggregators, lowering telephony ingestion expenses from $0.015 to under $0.004 per minute.

Evidence boundary

Official sources

Editorial guidance grounded in official product sources.

FAQ

Common questions

Why can per-character TTS pricing be misleading for real-time voice apps?

Per-character pricing only covers the text that becomes speech. A real-time voice app can also pay for speech input, call runtime, telephony, concurrency, retries, LLM usage, no-answer calls, testing traffic, logs, and evaluations.

Should Cartesia credits, ElevenLabs characters, and Fish Audio UTF-8 bytes be compared directly?

No. Convert each vendor unit into the same sample workload first. Estimate submitted text, generated audio, listening audio, call minutes, concurrent sessions, and retries, then map that workload to each official meter.

When do agent minutes matter more than TTS rates?

Agent minutes matter more when the vendor hosts the conversational layer or call path. Support lines, outbound calls, IVR replacements, and sales agents can be shaped more by total call time, telephony, and concurrency than by the amount of speech generated.

What is the difference between streaming TTS and a voice-agent platform?

Streaming TTS turns text into audio quickly, often through HTTP, SSE, or WebSocket routes. A voice-agent platform usually adds listening, turn-taking, LLM routing, tools, phone handling, logs, evaluations, and operational controls around the conversation.

How should telephony be included in a voice API pricing model?

Treat telephony as its own line item. Include phone-number ownership, provider pass-through, inbound and outbound minutes, failed or unanswered calls, testing traffic, and whether the voice vendor charges extra for using its provisioned numbers.

Where do OpenAI and Deepgram fit if they are not in this comparison?

Keep them as deferred context unless the buying question expands to realtime speech-to-speech or STT-first stacks. They need their own official pricing sources and route assumptions before they can be compared fairly with this voice API set.

Next steps

Take the next buying step

Use these next pages to confirm the plan, tool, or alternate route that fits once the spend boundary is clear.

View all tools