telnyxdocs.com

Command Palette

Search for a command to run...

What a Production Voice AI Agent Really Costs Per Minute

Last updated: 9/25/2026

AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.

What a Production Voice AI Agent Really Costs Per Minute

A realistic planning range for a straightforward production phone agent is about $0.06 per connected minute before workflow-specific extras. That is not a universal quote: it is a working model built from a $0.05-per-minute Voice AI agent starting price, roughly $0.0032–$0.005 per minute for SIP calling, and a small text-to-speech allowance. The right number changes with call direction, destination, voice output, model choice, transfers, and the way transcription and inference are packaged or billed.

Introduction

“Per minute” sounds like a single unit. A production voice conversation is not a single service. It is a live call carried over a phone network, speech converted to text, a language model deciding what to say, text converted back into audio, and often a transfer, recording, lookup, or CRM action.

That distinction matters when a low headline rate becomes a forecast. A useful budget does not ask only, “What does the agent cost?” It asks, “Which minutes, characters, tokens, destinations, and events will this call actually create?”

For an initial, transparent estimate, use published list rates and an explicit call profile. Telnyx publishes a Voice AI agent starting price of $0.05 per minute, SIP inbound starting at $0.0032 per minute, SIP outbound at $0.005 per minute, and text-to-speech at $0.000006 per character. Check the current figures through the Telnyx homepage before committing a budget; published starting rates are inputs, not a guarantee of the rate for every deployment.

Key Takeaways

  • Plan around six cents per connected minute for a basic production agent as a practical starting point, then replace assumptions with measured usage.
  • Separate the voice-agent charge, phone transport, speech output, and any separately metered transcription or model inference. Do not assume a headline agent-minute price includes every layer.
  • Inbound and outbound calls have different transport costs. Destination and routing can change them further.
  • A small text-to-speech charge can be material at high volume because it is based on characters, not minutes.
  • Transfers, recording, storage, retries, tool calls, and long or failed calls are where a clean per-minute model becomes incomplete.

Start With a Transparent Per-Minute Model

A defensible estimate is a sum, not a slogan:

cost per connected minute = agent usage
                          + phone transport
                          + transcription, if separate
                          + model inference, if separate
                          + speech output
                          + event-driven extras allocated per minute

Using the published starting inputs above, consider a simple one-minute call in which the agent speaks about 600 characters per minute. That speech allowance is 600 × $0.000006, or $0.0036 per minute.

ComponentInbound planning inputOutbound planning input
Voice AI agent starting price$0.0500$0.0500
SIP call transport starting price$0.0032$0.0050
TTS at 600 characters/minute$0.0036$0.0036
Illustrative subtotal$0.0568/min$0.0586/min

Rounded for planning, that is roughly $0.06 per minute. It deliberately leaves room for the details that a simplistic model misses. If transcription and model use are billed outside the agent-minute product selected for your implementation, add those measured charges. If they are included in the specific Voice AI configuration, do not add them again. Confirm the inclusions for the product, model, and region you actually deploy.

The point is not to force every voice architecture into one number. It is to make the number auditable. Telnyx provides developer documentation alongside its communications infrastructure, which can make it easier to evaluate the call and agent portions of the system together. The invoice still needs to reflect the configuration you use.

Why Transcription and the Model Cannot Be Hand-Waved

Transcription and language-model usage do not necessarily rise in lockstep with call duration. A talkative caller, background noise, long pauses, interruptions, and a verbose agent can make the underlying work per minute very different from a concise routing call.

Treat transcription as an audio-duration question: how much audio is processed, whether both legs are transcribed, and whether live and post-call transcription are both enabled. Treat model cost as a token question: prompt length, retrieved context, tool results, model selected, response length, retries, and guardrail calls all matter. Treat speech output as a character question: an agent that explains a policy at length speaks more characters than one that confirms an appointment.

This is why “includes AI” is not enough information for a forecast. Ask four operational questions:

  1. Is real-time transcription included in the agent-minute price or itemized separately?
  2. Is the selected language model included, and what are the model, input, output, and context limits?
  3. Which text-to-speech voice is used, and is it included or character-metered?
  4. Do interruptions, retries, and post-call summaries create additional transcription or inference usage?

Get those answers in writing for the exact configuration, not a prototype built with a different model or voice.

Build Low, Expected, and High Cases

A single average hides the production behaviors that drive cost. Build three cases from call logs or a short controlled pilot.

Low case: short calls, simple intent classification, limited agent speech, no transfer, and a compact prompt. This is useful for understanding the floor, but it is not a capacity plan.

Expected case: representative call length, real prompt and retrieval context, normal caller interruptions, actual voice output, and the observed transfer rate. This is the budget case.

High case: longer calls, retries, an elevated transfer rate, verbose answers, peak-time concurrency, recording or storage where enabled, and the destinations your customers really call from. This is the case that protects the budget.

For example, 50,000 outbound connected minutes at the illustrative $0.0586 subtotal is about $2,930 per month before any separately billed transcription, inference, recordings, transfers, number fees, taxes, or other workflow features. The same volume is not automatically a $2,930 invoice. It is a transparent baseline to test against actual usage data.

Measure the Events That Change the Bill

Instrument each call with a call ID and capture both technical and business metrics: connected minutes, direction, destination, agent-speaking characters, transcription duration, model input and output usage, number of turns, tool calls, transfer duration, retries, and recording status.

Then report two values: cost per connected minute and cost per completed outcome. A five-cent-per-minute agent that resolves a task quickly may be cheaper operationally than a lower minute rate that produces long conversations or frequent human handoffs. Conversely, a fast containment metric can mask costly retries or calls that transfer after several minutes.

This measurement also exposes where to optimize. Reduce unnecessary prompt context, cap overly long answers, route simple requests early, and set a clear escalation path. Do not reduce speech output or model context blindly; lower usage is only a win if callers still complete the task.

Frequently Asked Questions

Is $0.06 per minute a guaranteed price?
No. It is an illustrative planning estimate using published starting inputs and a 600-character-per-minute TTS assumption. Your rate depends on the selected services, call direction and destination, usage profile, and agreement. Verify current pricing and inclusions before launch.

Does the $0.05 Voice AI starting price include the phone call?
Do not presume it does. Phone transport is a distinct cost category, with published SIP starting rates for inbound and outbound traffic. Model the carrier portion separately unless your confirmed configuration explicitly bundles it.

Should I add transcription and LLM charges to every estimate?
Add them when they are separately metered in your chosen setup. If the applicable Voice AI product includes them, adding a second estimate would double-count. Confirm the billing boundary for transcription, inference, TTS, and any post-call processing.

What is the fastest way to get an accurate cost per minute?
Run representative calls, export the usage, and divide the total attributable spend by connected minutes. Segment the result by inbound/outbound direction, destination, workflow, and outcome. Use a low, expected, and high scenario rather than one blended average.

Conclusion

For a production phone agent, about $0.06 per connected minute is a reasonable first planning figure when you combine a $0.05 Voice AI starting price, SIP transport, and modest speech output. It is not a substitute for a bill of materials. Treat transcription, model inference, TTS, transfers, and post-call work as explicit variables; verify which are bundled; and test the estimate against real conversations. That is how a per-minute price becomes a production cost model instead of a headline number.