A single figure per minute looks like the simplest way to compare voice agents, and it is the number most offers lead with. It is also an average of several meters running at once, and some of those meters do not measure time at all.
- The minute is measured, not given Granularity, rounding and the moment the clock starts and stops change the minute count before any rate is applied to it.
- Not everything is timed Model and speech amounts are counted in tokens, characters and audio seconds, so they move with how a call goes, not only how long it lasts.
- Compare the boundary first A rate means something only once you know which parts sit inside it and whose accounts the underlying providers invoice.
A minute is not a fixed unit of work. Two calls of the same length can consume different amounts of the things being counted, because the counters measure audio seconds, tokens and characters rather than conversation.

The rest of the difficulty is boundary. Every offer draws a line around what its rate covers, and the lines sit in different places, so the first task is not finding a lower figure. It is making two figures describe the same thing.
What a Minute Means Before Any Rate Is Applied
Before a rate does anything, something has to decide how many minutes happened. That decision is made by counting rules rather than by the conversation, and the rules differ between the parties watching the same call.
- Granularity and rounding: whether time is recorded in seconds, in six-second blocks, or rounded up to the next whole minute.
- Where the clock starts: at dial, at ringing, at answer, or at the first spoken word, which are four different moments.
- Floors: a stated minimum that makes a fifteen-second call charge as though it ran longer.
- Silence and hold: waiting for a backend system to respond is connected time, and connected time is chargeable time.
- Call legs: a conversation transferred to a person can be recorded as two connections rather than one.
Rounding matters most where calls are short. An agent that confirms an appointment in forty seconds pays a whole minute under a whole-minute rule, so a rate that looks low can produce an invoice that behaves as though every conversation lasted longer than it did.
Several Clocks, One Conversation
There is rarely a single authoritative duration. The carrier records one, the orchestration layer records another, and the speech and model providers record what they each processed. Those records are taken at different points and rounded under different rules.
That matters at reconciliation rather than at signature. When the invoices arrive and the totals disagree, the difference is usually granularity rather than error, so agree in advance which record is the reference and where a dispute would be settled.
Ask two questions of any rate before comparing it with another: in what granularity is time recorded, and at which moment does the clock start.
Which Parts of a Call Are Counted, and in What Unit
A single figure per minute is a summary of several meters, and only one of them actually watches a clock. The others count units that a minute happens to contain, which is why the same duration can produce different amounts on different calls.

| What is counted | The unit it is counted in | What makes it move |
|---|---|---|
| Carried audio | Connected time, by granularity | Call length, direction and the destination reached |
| Transcription | Duration of audio processed | How much of the conversation contains speech at all |
| The model step | Tokens read and written on each turn | Prompt size, history length and the number of turns |
| The spoken reply | Characters or generated audio | How much the agent chooses to say |
| Orchestration | Time, month, or both together | The counting choice the supplier has made |
Read down the middle column and the problem becomes visible. Three of the five rows are counted in units the length of the conversation only loosely predicts, so a figure derived from them is an average across a set of calls rather than a price you were quoted in advance.
This has a practical consequence. A terse agent and a talkative one can run identical durations and still produce different totals, because the talkative one generated more characters and read a longer history on every turn.
It also means the figure moves when nothing commercial has changed. Rewriting a prompt, adding a document to the knowledge set or switching to a more expressive voice all shift the units consumed inside an unchanged duration. When the platform itself is what stopped working rather than the arithmetic, deciding whether to move off the platform you already run works through that separately.
The Amounts That Do Not Move With the Minute
Around the counted core sit standing amounts that stay the same whether the agent takes ten calls or ten thousand. They belong inside a single figure, but only after a volume has been named.
- A monthly amount for each telephone number the agent answers on or dials from.
- Any flat subscription that exists whether or not a single conversation happens.
- Retention of recordings, transcripts and call records for as long as policy requires them.
- Reserved capacity for simultaneous conversations, and whatever support level was chosen.
Setup work behaves the same way once it is spread across a year, and it is the item most often left outside the arithmetic entirely because it happens once.
Divide a standing monthly amount by minutes and the result falls as volume rises. A blended figure with no stated volume behind it is therefore incomplete rather than wrong, and it will drift as soon as call patterns change. How suppliers package these elements commercially is set out in the way voice agent pricing is usually structured.
Why Two Offers for the Same Minute Rarely Describe the Same Thing
Most of the confusion in a comparison is not arithmetic. It is that each supplier has drawn the boundary of its own rate somewhere different, and the boundary is often implied rather than stated.
Which of the counted parts the figure already contains, and which arrive later as separate lines on a different invoice.
Whether the telephony and model providers invoice the supplier, who then invoices you, or invoice your own accounts directly.
The granularity, the rounding and the floor, which change the minute count before the rate is even applied.
The duration, voice, model tier and destination the figure was calculated from, which may not resemble your conversations.
The second card is worth isolating, because it changes who receives the invoice rather than only what it says. Pulastya AI works this way: customers connect their own telephony and model provider accounts, those providers invoice the usage directly, and no Pulastya software fee sits above that usage.
Arrangements like that make the counted parts visible, which helps a comparison, but visibility is not the same as lower. The only figure that compares cleanly is a single total built on the same assumed conversation for every supplier.
Normalizing Offers Onto One Number
Normalizing is a short exercise once the pieces are named. The aim is a single monthly total per supplier, built from your own conversations rather than from anyone's example.
- Fix a reference callTake a real call reason and record its typical duration, its number of turns and how much the agent has to say.
- Define the minuteApply each supplier's granularity, rounding and floor to that reference call before applying any rate.
- Convert the unit-counted partsTurn token and character volumes into a per-conversation amount using the reference call, not a generic assumption.
- Add the standing layerBring in numbers, subscription, retention, capacity and setup, divided by the volume you actually expect.
- Write the assumptions downRecord every value you had to assume, because those are what will move when the comparison is challenged.
Do the exercise once for each distinct call reason rather than once for the whole line. A two-turn appointment confirmation and a six-turn account question sit at opposite ends of the same headline figure, and a blended average of the two predicts neither.
The reference call is the part most worth getting right, and it is not something a supplier can provide. It comes from your own traffic, or from a short trial run on it, as described in a trial designed to predict production behavior.
Per Minute, per Call, or per Outcome
The headline figure answers a narrow question. It says what sixty seconds costs, not what getting something done costs, and those two can move in opposite directions.
- An agent that answers faster shortens conversations, which lowers the amount per call while leaving the rate untouched.
- A conversation that ends in a transfer may produce two legs, so the amount per resolved request is higher than the headline view suggests.
- A caller who has to ring back has produced two charged calls for one piece of work.
- A cheaper model that needs an extra clarifying turn can raise the total of the conversation it was chosen to make cheaper.
Changing the denominator is what makes the figure comparable to anything outside voice. A request answered is a unit a business already understands; sixty seconds of connected audio is not.
Waiting counts too. Time spent holding while a backend system responds is connected time, which is one reason response delay is a commercial question as well as an experience one, as set out in where delay in a spoken turn actually comes from.
The amount per resolved request is the figure a finance team will eventually ask for, and building it is a different exercise, covered in how to build a cost case a finance team will accept. Model, voice and telephony choices are made during configuration on the AI voice agent platform.
Frequently Asked Questions
It is a reasonable starting point and a poor finishing point. A rate only means something once you know which parts sit inside it, whose accounts the providers invoice, and what call profile it was calculated from. Convert each offer into one monthly total using your own durations and volumes, then divide back down if a single figure is still wanted.
Because most of what is counted is not time. Tokens depend on how long the prompt and conversation history have grown, characters depend on how much the agent says, and transcription depends on how much of the call is speech. A conversation that takes three turns instead of six costs less even when the clock reading is identical.
Usually, but not always. A whole-minute rounding rule or a stated floor can make a forty-second conversation charge the same as a longer one. Shortening calls below the granularity stops producing savings, and a short call that ends in a transfer or a return call can cost more in total than one longer conversation that finished the work.
Often the connection to the person is recorded as a second leg, so the conversation produces two chargeable segments rather than one. Check how each supplier and each telephony arrangement records it, since the difference is material for any call reason that transfers frequently. Include a transferring call reason in the reference conversations used for comparison.
Use the volume you can defend from current traffic, and then test the comparison at a lower and a higher figure. Standing monthly amounts divided by minutes make the headline number fall as volume rises, so a comparison that holds at one volume can reverse at another. Record the volume alongside the result it produced.







