AI voice agent pricing is rarely a single number. A production bill combines a platform or orchestration fee, telephony, speech-to-text, a large language model, text-to-speech, optional add-ons, and the one-time work of integration and implementation.
Knowing how each vendor packages those pieces is the fastest way to compare offers on a like-for-like basis.
- Primary cost drivers Billable minutes, average call length, model and voice choices, telephony type, add-ons, and integration scope.
- Common pricing models Bundled per-minute rates, a platform fee plus pass-through provider costs, bring-your-own-keys with no platform fee, and annual enterprise contracts.
- Control strategies Estimate from real call data, price every component, confirm billing increments and minimums, and review usage monthly.
Buyers meet two kinds of cost. Usage costs scale with every minute a caller is connected: telephony, transcription, model tokens, synthesized speech, and often a platform fee on top. Fixed costs stay flat regardless of volume: phone number rental, plan subscriptions, compliance or concurrency add-ons, support tiers, and the engineering needed to connect the agent to business systems.

The same workload can produce very different invoices under different pricing models. A bundled rate hides the component split, a pass-through model exposes it, and a bring-your-own-keys model moves provider billing onto your own accounts. Map expected call volume and call length first, then translate every offer into the same monthly formula before comparing.
Last reviewed: September 12, 2026. Vendor plans and pricing change often, so confirm current terms with each vendor before you commit.
Break Down the Cost Components
Every connected minute of an AI voice agent call consumes several services at once. Telephony carries the audio, speech-to-text (STT) turns the caller's words into text, a large language model (LLM) decides what to say or which tool to call, and text-to-speech (TTS) speaks the reply.

A platform or orchestration layer coordinates those services. It handles turn-taking and interruptions, connects tools and knowledge sources, and usually provides the dashboard, call logs, and analytics. Its fee is the platform charge that appears on many quotes.
Each component is metered differently. Telephony providers charge phone number rental plus per-minute carrier rates that vary by call direction, local or toll-free numbers, and destination country. STT is typically metered by audio duration, LLM usage by input and output tokens, and TTS by characters or audio generated. Platform fees are usually per minute, per month, or both.
Some architectures use a single speech-to-speech model that bills audio input and output together, folding STT, LLM, and TTS into one usage line. The same work still happens; it simply arrives as one charge.
Around that per-minute core sit costs that do not scale with minutes: add-ons such as HIPAA or other compliance modes and extra concurrent-call capacity, recording and transcript storage, integration and implementation work, and support tiers. Leaving any of them out makes the cheapest-looking quote misleading.
- List telephony separately: number rental, inbound minutes, outbound minutes, and toll-free or international rates
- Record how STT, LLM, and TTS are metered: audio duration, tokens, or characters
- Note whether the platform fee is charged per minute, per month, or both
- Add fixed lines for add-ons, storage, support, and integration work
Price every component on the same per-minute and per-month basis before comparing vendor quotes.
Compare the Four Common Pricing Models
Bundled per-minute pricing quotes one rate that covers the platform and some or all provider costs. It is easy to forecast, but the rate often changes with the voice, model, or telephony options selected, so confirm exactly which components a quoted rate includes.
Platform fee plus pass-through pricing charges a per-minute platform fee and bills STT, LLM, TTS, and telephony at provider cost as separate lines. The component split is visible, which helps optimization, but the all-in rate is known only once provider usage is added.
Bring-your-own-keys (BYOK) pricing lets you connect your own provider accounts, so the telephony and model providers bill you directly. When the platform also charges no platform fee, nothing is added to the per-minute cost above provider usage. You still pay the providers, and you manage those accounts, keys, and usage limits.
Annual enterprise contracts replace list prices with a negotiated commitment that typically bundles usage, onboarding, support, and compliance terms. They suit large, predictable programs, but minimum commitments and overage terms decide the real cost.
Pulastya AI uses the BYOK model with no Pulastya software charge: customers connect their own supported telephony and AI-provider accounts, currently Twilio and OpenAI, and those providers bill the usage directly. That removes the platform fee, but it does not automatically make Pulastya the lowest all-in cost, because telephony and model usage varies with call length, model choice, and volume.
Published price points from other vendors, as of September 2026, map onto these models. Vapi follows the platform fee plus pass-through model: its usage-based plan charges a $0.05 per minute platform fee, with transcriber, model, voice, and telephony costs billed at cost, and it does not bill usage for a provider whose key you supply.
Retell prices pay-as-you-go voice agents per minute at $0.07 to $0.31, depending on the voice infrastructure, TTS, LLM, and telephony components chosen. Bland combines per-minute rates with plan tiers: a Start plan at $0.14 per minute with no platform fee, and a Build plan at $0.12 per minute plus $299 per month.
Bland's pricing and billing pages list plans differently, so confirm current plans with Bland before modeling. Synthflow sells enterprise contracts starting at $30,000 a year. Prices change often, so treat these figures as reference points rather than quotes.
Price is only one part of the decision. The comparison of AI voice agent platforms covers capabilities and buyer fit alongside these pricing approaches.
- Bundled per-minute: confirm which components, voices, and models the rate covers
- Platform fee plus pass-through: add every provider line to the platform fee
- BYOK: price provider usage at your own call length, model, and volume
- Enterprise contract: check minimum commitments, overage rates, and included support
Translate every pricing model into the same monthly total before deciding which is cheaper for your workload.
Pricing models
Four Ways Voice Agent Vendors Bill the Same Call Minutes
| Pricing model | How you are billed | Check before signing |
|---|---|---|
| Bundled per-minute | One rate covers the platform plus some or all telephony, STT, LLM and TTS costs | Which components, voices and models the quoted rate includes |
| Platform fee plus pass-through | A per-minute platform fee, with provider usage billed at cost on separate lines | The all-in rate once every provider line is added to the fee |
| Bring your own keys, no platform fee | No platform fee; telephony and model providers bill your own accounts directly | Provider usage at your own call length, model choice and volume |
| Annual enterprise contract | A negotiated yearly commitment that bundles usage, onboarding and support | Minimum commitments, overage rates and what support includes |
Each model can price the same minutes differently, so convert every quote into one monthly total before comparing.
Understand What Drives Per-Minute Cost
Call length is the first multiplier. Every component bills while a call is connected, so an agent that resolves a request in two minutes instead of four cuts usage cost for that call type substantially.
Holds while the agent waits on a slow backend system, repeated confirmations, and callers who stay on the line after the task is done all add billable minutes.
LLM cost can grow faster than call length. Each conversational turn typically resends the system prompt, the conversation so far, and any retrieved document context as input tokens, so long calls and large prompts raise the cost of every later turn.
Model tier matters too: a larger model can cost substantially more per token than a smaller one that handles routine requests well.
Speech costs depend on provider and tier. STT rates differ between standard and higher-accuracy models, and TTS rates differ by provider and voice quality, so price the specific voice you plan to use rather than a default. Telephony rates depend on direction, number type, and destination, and outbound, toll-free, and international calls are often priced differently from local inbound calls.
Match model and voice choices to the job. Routine questions answered from approved documents rarely need the largest model, while complex or sensitive conversations may justify a stronger one or a transfer to a person. Test candidate combinations on scripted or recorded calls and compare both quality and per-minute cost before committing.
- Measure average call length per call type, including hold and wait time
- Track input and output tokens per call, not just minutes
- Price the exact STT model and TTS voice you plan to use
- Check telephony rates for outbound, toll-free, and international calls
Shorter calls and right-sized models can matter as much to per-minute cost as the platform fee.
Budget for Implementation, Integrations, and Support
One-time costs often decide whether a low per-minute rate becomes a low total cost. Implementation covers conversation and prompt design, preparing the documents the agent answers from, configuring transfer rules, and testing. A focused use case with good source documents needs far less work than a multi-department rollout with authentication and transactions.
Integrations add engineering effort whenever the agent must read or write business data, such as looking up an order, creating a ticket, updating a CRM record, or checking availability.
Budget for API work, webhook handling, error cases, and security review, whether your team builds them or a vendor or partner does. Prebuilt connectors reduce this work only when they match your systems and fields.
Support is a recurring line that headline rates often leave out. Self-serve plans may rely on documentation and community help, while enterprise tiers add named contacts, onboarding, and service-level commitments. Decide what response time a production phone line needs and price the tier that provides it.
Fine-tuning and self-hosting are optional costs, not defaults. Most deployments run on hosted models with prompt design and retrieval over approved documents. Consider fine-tuning only when testing shows a gap those methods cannot close, and self-hosting only when data-control requirements rule out hosted APIs; both add labeled data, compute capacity, and engineering to operate.
- Scope prompt design, document preparation, and testing as one-time work
- List every system the agent must read from or write to
- Price the support tier that matches production response needs
- Treat fine-tuning and self-hosting as exceptions that need a tested justification
A low per-minute rate can be outweighed by integration and support costs, so budget them from the start.
Account for Add-Ons, Transfers, and Ongoing Operations
Add-ons change the price of the same minutes. Compliance modes, such as HIPAA handling with a signed business associate agreement, can require a higher plan tier or a paid add-on. Extra concurrent-call capacity, additional phone numbers, longer recording retention, and advanced analytics can also be priced separately, so ask for the add-on list before comparing base rates.
Transfers and fallbacks carry their own costs. When the agent transfers a call to a person, the telephony provider may bill the forwarded leg in addition to the original call, and the staff member's time becomes part of the cost of that interaction. A higher transfer rate raises staff cost while lowering agent minutes, so track the two together.
Operations add steady, smaller lines: reviewing transcripts and failed calls, updating documents when prices or policies change, monitoring dashboards, and storing recordings for the retention period your policies require. Under a BYOK model, finance also needs to watch usage in each provider account, because those charges arrive on separate invoices.
- Request the full add-on list, including compliance modes and concurrency limits
- Include transferred call legs and staff handling time in cost per call
- Set recording and transcript retention to what policy actually requires
- Assign an owner to review platform and provider usage every month
Transfer rate, retention settings, and add-ons can move monthly cost as much as the headline rate.
Example Scenario and a Simple Cost Model
A single formula makes quotes comparable:
Monthly cost = billable minutes x (telephony + STT + LLM + TTS cost per minute) + billable minutes x platform fee per minute + fixed monthly costs + one-time costs spread over the contract term
Billable minutes equal calls per month times average call length, adjusted for each provider's billing increment. Add transferred call legs to the telephony line, since the agent usually stops processing speech once a person takes the call. Fixed monthly costs include phone numbers, plan subscriptions, add-ons, and support.
For a bundled rate, the usage components and platform fee collapse into one number; for BYOK with no platform fee, the platform term is zero and the usage components come from provider invoices.
Run that formula across the whole contract term rather than one month and it becomes the total cost of ownership for the deployment: usage, platform charges, one-time implementation and integration work, and the ongoing operational effort of keeping content and rules current. A quote covers the first of those. The remaining three are where two apparently similar offers usually diverge.
Take component rates from each provider's current price list or, better, from pilot usage. Run test or pilot calls that represent the main call types, then divide each provider's usage charges by the minutes connected. That captures real token counts and call lengths instead of assumptions.
Example: a service business expects 10,000 calls a month with an average length of four minutes, or 40,000 billable minutes. At that volume, every $0.01 added to the all-in per-minute rate adds $400 a month, so small differences in platform fee or model choice add up quickly.
If 15 percent of calls transfer to staff with a three-minute forwarded leg, telephony minutes rise by another 4,500 a month.
Build three scenarios: expected volume, peak season, and a conservative case with longer calls and a higher transfer rate. Fixed fees weigh most at low volume while per-minute fees dominate at high volume, so the cheapest option can change between scenarios. Present the range to finance rather than a single estimate.
- Calculate billable minutes from call count and call length, then add transferred legs to telephony
- Derive component rates from pilot usage, not list prices alone
- Set the platform fee term to zero for BYOK options with no platform charge
- Compare every option across expected, peak, and conservative scenarios
A small change in call length or transfer rate can shift the monthly total more than a small difference in list price.
Procurement and Cost Governance Best Practices
Ask every shortlisted vendor the same pricing questions in writing. Which components does the per-minute rate include? Are provider costs passed through at cost or marked up? What is the billing increment, and are short or failed calls billed? Are there minimum commitments, concurrency limits, or separate charges for numbers, storage, and support?
Clarify which features sit behind higher tiers. Warm transfer, compliance modes, extra concurrency, and dedicated support are sometimes limited to enterprise plans, which changes the effective price of a self-serve rate. Confirm how price changes are communicated and whether committed rates hold for the contract term.
After launch, review cost monthly with the same formula used for the estimate. Compare actual minutes, call length, transfer rate, and provider charges against the forecast, and set budget alerts in the platform and in each provider account. Use the findings at renewal to renegotiate commitments or move routine call types to lower-cost models.
- Send identical written pricing questions to every shortlisted vendor
- Confirm which features require a higher tier or enterprise contract
- Set budget alerts in the platform and every provider account
- Reconcile actual usage against the forecast each month
Cost control depends on comparing actual usage with the forecast, not only on the rate negotiated at signing.
Cost drivers are easier to evaluate alongside software selection criteria, implementation phases, and ongoing operational measurement.
Conclusion
AI voice agent pricing is best compared using total cost, not just a headline per-minute rate. Include telephony, speech-to-text, language model usage, text-to-speech, platform charges, implementation, integrations, support, and add-ons. Estimate costs using realistic call volumes, average call lengths, and transfer rates so bundled, pass-through, bring-your-own-keys, and enterprise offers can be assessed on the same basis.
Before choosing a provider, validate the numbers with pilot calls and confirm billing increments, minimum commitments, overage terms, and features restricted to higher plans. After launch, review actual usage and provider invoices each month, then adjust workflows and model choices as needed. The goal is a reliable voice agent with predictable operating costs and a pricing model that suits your workload.
Frequently Asked Questions
Per-minute cost is the sum of telephony, speech-to-text, the language model, text-to-speech, and any platform fee. It rises with longer calls, larger models, longer prompts and conversation history, premium voices, and outbound, toll-free, or international calling. Transferred call legs and add-ons such as compliance modes add further charges.
Not automatically. Bring-your-own-keys pricing with no platform fee removes that fee, but you still pay the telephony and model providers directly, and those charges vary with call length, model choice, and volume. Compare both options with the same monthly formula and your own call data before deciding.
Multiply calls per month by average call length to get billable minutes, then multiply by the combined per-minute cost of telephony, speech-to-text, the language model, text-to-speech, and any platform fee. Add fixed costs such as phone numbers, subscriptions, add-ons, and support, and validate the component rates with pilot calls.
Ask about billing increments, minimum commitments, concurrency limits, phone number fees, compliance add-ons, recording storage, transferred call legs, support tiers, and features reserved for enterprise plans. Also ask whether provider costs are passed through at cost or marked up, and how price changes are handled during the contract.







