Home / Blogs & Insights / How Much Does an AI Voice Agent Cost Per Minute?

How Much Does an AI Voice Agent Cost Per Minute?

AI voice agent cost per minute infographic showing telephony, speech-to-text, LLM, text-to-speech, platform fees, and cost optimization.

Table of Contents

A single figure per minute looks like the simplest way to compare voice agents, and it is the number most offers lead with. It is also an average of several meters running at once, and some of those meters do not measure time at all.

At A Glance
  • The minute is measured, not given Granularity, rounding and the moment the clock starts and stops change the minute count before any rate is applied to it.
  • Not everything is timed Model and speech amounts are counted in tokens, characters and audio seconds, so they move with how a call goes, not only how long it lasts.
  • Compare the boundary first A rate means something only once you know which parts sit inside it and whose accounts the underlying providers invoice.

A minute is not a fixed unit of work. Two calls of the same length can consume different amounts of the things being counted, because the counters measure audio seconds, tokens and characters rather than conversation.

Pulastya AI voice agent cost view showing how a call's per-minute charge breaks down across telephony, transcription, model and speech usage.

The rest of the difficulty is boundary. Every offer draws a line around what its rate covers, and the lines sit in different places, so the first task is not finding a lower figure. It is making two figures describe the same thing.

What a Minute Means Before Any Rate Is Applied

Before a rate does anything, something has to decide how many minutes happened. That decision is made by counting rules rather than by the conversation, and the rules differ between the parties watching the same call.

  • Granularity and rounding: whether time is recorded in seconds, in six-second blocks, or rounded up to the next whole minute.
  • Where the clock starts: at dial, at ringing, at answer, or at the first spoken word, which are four different moments.
  • Floors: a stated minimum that makes a fifteen-second call charge as though it ran longer.
  • Silence and hold: waiting for a backend system to respond is connected time, and connected time is chargeable time.
  • Call legs: a conversation transferred to a person can be recorded as two connections rather than one.

Rounding matters most where calls are short. An agent that confirms an appointment in forty seconds pays a whole minute under a whole-minute rule, so a rate that looks low can produce an invoice that behaves as though every conversation lasted longer than it did.

Several Clocks, One Conversation

There is rarely a single authoritative duration. The carrier records one, the orchestration layer records another, and the speech and model providers record what they each processed. Those records are taken at different points and rounded under different rules.

That matters at reconciliation rather than at signature. When the invoices arrive and the totals disagree, the difference is usually granularity rather than error, so agree in advance which record is the reference and where a dispute would be settled.

Ask two questions of any rate before comparing it with another: in what granularity is time recorded, and at which moment does the clock start.

Which Parts of a Call Are Counted, and in What Unit

A single figure per minute is a summary of several meters, and only one of them actually watches a clock. The others count units that a minute happens to contain, which is why the same duration can produce different amounts on different calls.

Grid of the six units a voice agent call can be billed in: per second, per audio minute, per token, per character, per month and per call leg, shown as separate meters.
One per-minute figure summarizes meters that count seconds, audio, tokens, characters, months and call legs, which is why the same minute can bill differently.
What is countedThe unit it is counted inWhat makes it move
Carried audioConnected time, by granularityCall length, direction and the destination reached
TranscriptionDuration of audio processedHow much of the conversation contains speech at all
The model stepTokens read and written on each turnPrompt size, history length and the number of turns
The spoken replyCharacters or generated audioHow much the agent chooses to say
OrchestrationTime, month, or both togetherThe counting choice the supplier has made

Read down the middle column and the problem becomes visible. Three of the five rows are counted in units the length of the conversation only loosely predicts, so a figure derived from them is an average across a set of calls rather than a price you were quoted in advance.

This has a practical consequence. A terse agent and a talkative one can run identical durations and still produce different totals, because the talkative one generated more characters and read a longer history on every turn.

It also means the figure moves when nothing commercial has changed. Rewriting a prompt, adding a document to the knowledge set or switching to a more expressive voice all shift the units consumed inside an unchanged duration. When the platform itself is what stopped working rather than the arithmetic, deciding whether to move off the platform you already run works through that separately.

The Amounts That Do Not Move With the Minute

Around the counted core sit standing amounts that stay the same whether the agent takes ten calls or ten thousand. They belong inside a single figure, but only after a volume has been named.

  • A monthly amount for each telephone number the agent answers on or dials from.
  • Any flat subscription that exists whether or not a single conversation happens.
  • Retention of recordings, transcripts and call records for as long as policy requires them.
  • Reserved capacity for simultaneous conversations, and whatever support level was chosen.

Setup work behaves the same way once it is spread across a year, and it is the item most often left outside the arithmetic entirely because it happens once.

Divide a standing monthly amount by minutes and the result falls as volume rises. A blended figure with no stated volume behind it is therefore incomplete rather than wrong, and it will drift as soon as call patterns change. How suppliers package these elements commercially is set out in the way voice agent pricing is usually structured.

Why Two Offers for the Same Minute Rarely Describe the Same Thing

Most of the confusion in a comparison is not arithmetic. It is that each supplier has drawn the boundary of its own rate somewhere different, and the boundary is often implied rather than stated.

The inclusion boundary

Which of the counted parts the figure already contains, and which arrive later as separate lines on a different invoice.

Whose account is invoiced

Whether the telephony and model providers invoice the supplier, who then invoices you, or invoice your own accounts directly.

The counting rule

The granularity, the rounding and the floor, which change the minute count before the rate is even applied.

The assumed call

The duration, voice, model tier and destination the figure was calculated from, which may not resemble your conversations.

The second card is worth isolating, because it changes who receives the invoice rather than only what it says. Pulastya AI works this way: customers connect their own telephony and model provider accounts, those providers invoice the usage directly, and no Pulastya software fee sits above that usage.

Arrangements like that make the counted parts visible, which helps a comparison, but visibility is not the same as lower. The only figure that compares cleanly is a single total built on the same assumed conversation for every supplier.

Normalizing Offers Onto One Number

Normalizing is a short exercise once the pieces are named. The aim is a single monthly total per supplier, built from your own conversations rather than from anyone's example.

  1. Fix a reference callTake a real call reason and record its typical duration, its number of turns and how much the agent has to say.
  2. Define the minuteApply each supplier's granularity, rounding and floor to that reference call before applying any rate.
  3. Convert the unit-counted partsTurn token and character volumes into a per-conversation amount using the reference call, not a generic assumption.
  4. Add the standing layerBring in numbers, subscription, retention, capacity and setup, divided by the volume you actually expect.
  5. Write the assumptions downRecord every value you had to assume, because those are what will move when the comparison is challenged.

Do the exercise once for each distinct call reason rather than once for the whole line. A two-turn appointment confirmation and a six-turn account question sit at opposite ends of the same headline figure, and a blended average of the two predicts neither.

The reference call is the part most worth getting right, and it is not something a supplier can provide. It comes from your own traffic, or from a short trial run on it, as described in a trial designed to predict production behavior.

Per Minute, per Call, or per Outcome

The headline figure answers a narrow question. It says what sixty seconds costs, not what getting something done costs, and those two can move in opposite directions.

  • An agent that answers faster shortens conversations, which lowers the amount per call while leaving the rate untouched.
  • A conversation that ends in a transfer may produce two legs, so the amount per resolved request is higher than the headline view suggests.
  • A caller who has to ring back has produced two charged calls for one piece of work.
  • A cheaper model that needs an extra clarifying turn can raise the total of the conversation it was chosen to make cheaper.

Changing the denominator is what makes the figure comparable to anything outside voice. A request answered is a unit a business already understands; sixty seconds of connected audio is not.

Waiting counts too. Time spent holding while a backend system responds is connected time, which is one reason response delay is a commercial question as well as an experience one, as set out in where delay in a spoken turn actually comes from.

The amount per resolved request is the figure a finance team will eventually ask for, and building it is a different exercise, covered in how to build a cost case a finance team will accept. Model, voice and telephony choices are made during configuration on the AI voice agent platform.

Frequently Asked Questions

It is a reasonable starting point and a poor finishing point. A rate only means something once you know which parts sit inside it, whose accounts the providers invoice, and what call profile it was calculated from. Convert each offer into one monthly total using your own durations and volumes, then divide back down if a single figure is still wanted.

ABOUT THE AUTHOR

Anuj Yadav

Anuj Yadav is the CBO of SDLC Corp, leading business strategy across AI, blockchain, Web3, and digital innovation. He focuses on helping businesses plan and commercialize AI-led products, including generative AI and machine learning, while aligning technology with market fit, implementation, and growth.
PLAN YOUR SOLUTION

More Insights
You Might Find Useful

Explore expert perspectives, practical strategies, and real-world solutions related to this topic.

AI voice agent pricing dashboard showing monthly cost, call usage, per-minute cost breakdown, pricing models, and ROI insights on a laptop screen.

AI Voice Agent Pricing: Cost Drivers and Planning

AI voice agent pricing is rarely a single number. A

AI voice agent for sales dashboard showing inbound leads, qualification, CRM context, meeting booking, follow-up, call outcomes, and sales performance analytics.

AI Voice Agents for Sales Calls and Follow-Up

An AI voice agent for sales is a voice agent

Let’s Talk About Your Product

Get expert guidance on scope, architecture, timelines, and delivery approach so you can move forward with confidence.

What happens next?