A voice agent looks like one product and behaves like four.
A communications network carries the call. Speech recognition turns audio into text and speech synthesis turns text back into audio. A language model decides what to say. Business systems supply the facts. Voice AI integrations are the joins between those layers.
- Four layers, four markets Communications, speech, models and enterprise systems are bought, priced and swapped separately.
- Recognition is not synthesis Speech-to-text and text-to-speech are distinct components with distinct vendors and opposite failure modes.
- Supported lists are short Every platform supports a specific, named set of providers. Ask for that list in writing during evaluation.
Each layer is a separate market with its own vendors, its own pricing unit and its own ways of being wrong. Treating the stack as a single purchase is the most common mistake in a voice AI evaluation.
Naming the categories, and the vendors who compete in each of them, is what lets a buyer tell which questions belong to a carrier, which belong to a speech or model provider, and which belong to the platform sitting above them. The commercial model layer is compared directly in the major hosted model providers for voice agents .

The Four Layers a Voice Agent Depends On
Naming the layers separately is not pedantry. A call can fail because a carrier dropped the leg, because the transcript was wrong, because the model answered outside policy, or because a record lookup timed out. Four different problems, four different owners.

Above the four layers sits orchestration: the part that holds the conversation together and decides when to retrieve, when to act and when to hand the call to a person. The guide to where calls enter a voice AI architecture covers that coordination in detail.
| Layer | What it does | Who plays in it | What changes if you swap it |
|---|---|---|---|
| Communications | Carries the call over PSTN, SIP or VoIP and moves audio both ways in real time | Twilio, Telnyx, Vonage, Plivo, Sinch, Bandwidth, Azure Communication Services | Numbers, country coverage, per-minute rates, concurrency limits and the call-control model |
| Speech recognition | Turns caller audio into text during the call and detects when a turn has ended | Deepgram, AssemblyAI, Azure AI Speech, Google Cloud Speech, Amazon Transcribe | Accuracy on accents, names and noise, plus how quickly the agent decides you stopped talking |
| Speech synthesis | Turns the agent's text back into spoken audio the caller hears | ElevenLabs, Cartesia, Amazon Polly, Azure AI Speech, Google Cloud speech services | The voice itself, pronunciation control, language coverage and time to first audio |
| Language model | Interprets the transcript, applies instructions and knowledge, and produces the next turn | OpenAI, Azure OpenAI, Anthropic Claude, Google Gemini, Amazon Bedrock, Meta Llama, Mistral, open-weight models | Instruction-following, tool-calling reliability, first-token latency and price per token |
| Enterprise systems | Hold the records an answer depends on and receive what the call produces | Salesforce, HubSpot, Microsoft Dynamics, ServiceNow, Zendesk, SAP, Oracle, Genesys, Five9, Amazon Connect, internal APIs | Which questions can be answered from live data instead of a document, and where outcomes land |
The speech layer occupies two rows on purpose. Recognition and synthesis are routinely collapsed into one word, usually voices. They are separate components, bought separately, and they fail in opposite directions: recognition errors change what the agent hears, synthesis errors change what the caller hears.
Orchestration is absent from the table because it is not a commodity you swap between calls. It is the platform decision, and the layers below it are dependencies that decision inherits. The overview of an enterprise voice AI platform covers what that layer owns.
Account ownership is a design question at every layer, not a billing footnote. Pulastya, the SDLC Corp voice platform for inbound and outbound business calls, takes an explicit position on it: customers connect their own Twilio and OpenAI accounts, and those providers bill usage directly.
Telephony and Communications Providers
The communications layer carries the call. It provisions or ports numbers, connects to the public telephone network, moves audio in both directions with low latency, and tells the agent platform that a call has started. Nothing above it is reachable without it.
What Communications Providers Sell
These vendors package numbers, call control and media handling behind an API. Twilio, Telnyx, Vonage, Plivo, Sinch, Bandwidth and Azure Communication Services all compete here. It is a mature market, and the products are not interchangeable in detail: coverage, rates and call-control models differ.
What separates them for voice AI specifically is media handling: whether audio streams in both directions while the call is live, and how short the path to the first packet is. Voice APIs built for recorded notifications are a poor fit for conversation.
SIP, PSTN and Existing Enterprise Telephony
Larger organizations rarely start from nothing. They already run SIP trunks, an IP PBX, a VoIP estate or a contact center platform, and the numbers are in service and printed on things. The question is not which carrier to buy, but how an agent joins telephony that exists.
- Direct SIP: the agent platform behaves as a SIP endpoint and receives calls from the existing trunk or PBX.
- Webhook redirect: an existing business number keeps its carrier, and its inbound webhook points at the agent platform instead.
- New number: a dedicated number is provisioned for a pilot, so the main line is untouched while the agent is tested.
- Contact center leg: the call reaches the agent as a queue, route or destination inside the contact center platform already in place.
Each route has different consequences for porting, recording and what happens when the agent is unreachable. Failure handling belongs at this layer as much as anywhere, and the article on detecting failure and recovering safely covers detection, fallback and the path back to a human line.
Concurrency is the telephony question that surprises teams most. Automation removes the staffing ceiling that used to cap simultaneous calls, so a campaign or an incident can generate more live legs than the account is provisioned for. Concurrency limits, rate limits and spend caps belong in the evaluation.
Models, Speech Recognition and Speech Synthesis
These three components are usually collapsed into one phrase: the AI. Separating them is the difference between a useful bug report and a shrug, because each has its own vendors, its own latency budget and its own characteristic mistakes.
Speech Recognition Turns Audio Into Text
Speech-to-text runs continuously during the call and decides two things: which words the caller said, and when the caller stopped speaking. Deepgram, AssemblyAI, Azure AI Speech, Google Cloud Speech and Amazon Transcribe compete for this work, and accuracy varies with accent, background noise and domain vocabulary.
End-of-speech detection lives here, and it governs how the conversation feels. Decide too early and the agent interrupts; wait too long and it feels slow. Barge-in, where a caller talks over the agent and is heard, depends on recognition running while audio is still playing.
Speech Synthesis Turns Text Back Into Audio
Text-to-speech is a separate market. ElevenLabs, Cartesia, Amazon Polly, Azure AI Speech and the Google Cloud speech services all sell synthesis, and they differ on voice quality, pronunciation control, language coverage and time to first audio. A natural voice that starts half a second late still sounds wrong on a phone call.
Voice selection is a business decision as much as a technical one, because the voice is what the organization sounds like. Language coverage, tone and the pronunciation of names and product terms are covered in the guide to language, voice and persona settings for AI calling .
The Model Layer Decides What to Say
The language model reads the transcript, applies the instructions and retrieved knowledge it has been given, and produces the next turn. OpenAI, Azure OpenAI, Anthropic Claude, Google Gemini, Amazon Bedrock, Meta Llama and Mistral are the names most often shortlisted, alongside self-hosted open-weight models.
Swapping a model changes more than answer quality. Instruction-following, tool-calling reliability, first-token latency and cost per token all move together, and prompts tuned for one model rarely transfer unchanged. The settings that surround the model are covered in how the same agent settings are configured per call type .
The agent acts on words the caller never said. It shows up as wrong intents and strange follow-up questions.
The words are right and the delivery is wrong: a mispronounced name, an odd emphasis, audio that starts late.
The transcript was correct and the response was still wrong, unsupported by any source, or outside policy.
Audio degrades, latency climbs or the leg drops. No amount of model quality compensates for it.
Every voice AI platform supports a specific, named set of providers at each layer, and that list is usually short. Do not assume a provider is supported because it is well known.
Ask any vendor for the current list of supported telephony, speech and model providers in writing, and ask what happens to your configuration when one is added or retired.
Cost follows the same layering: carriers bill per minute, speech vendors bill per minute or per character, models bill per token. The breakdown is in four ways vendors bill the same call minutes , and the testing sequence in what to test for each evaluation criterion .
Enterprise Systems and Custom APIs
The fourth layer is what makes a voice agent useful rather than merely fluent. It holds the order, the ticket, the policy, the appointment and the account, and it is where the outcome of a call has to land once the caller has hung up.
It is also the broadest market of the four. Salesforce, HubSpot and Microsoft Dynamics in CRM, SAP and Oracle in ERP, ServiceNow in ITSM, Zendesk in support, and Genesys, Five9 and Amazon Connect in the contact center are all systems a voice program is likely to meet.
Which pattern fits depends on the system and the call, not on the length of a connector list. The guide to integration patterns for CRM, ERP and ITSM separates read-only lookups from writes, and both from updates that are queued and completed after the call ends.
Documents or Live Lookups
Two sources answer questions on a call, and they are not interchangeable. Documents answer stable questions: policies, hours, eligibility, product descriptions. Live systems answer questions about one record: where my order is, whether my ticket is open. Knowledge base or real-time API sets out which belongs where.
Grounding also decides what happens at the edge of knowledge. Pulastya answers from documents the organization provides, and when it lacks the context to answer it says so and offers a transfer instead of guessing, which is the behavior a spoken answer needs when nobody can see the source.
Custom Internal APIs
For most enterprises the most important system has no connector at all. It is an internal API, a middleware endpoint or a database view built years ago. That is normal, and it is usually the integration carrying the real business value.
- Authentication: how credentials are stored, rotated and scoped to the least access the agent needs.
- Latency budget: what the agent says while a lookup runs, and what it does when that lookup times out.
- Error handling: a failed call should produce an honest sentence and a transfer, never an invented answer.
- Write safety: retries must not create duplicate tickets, orders or appointments.
- Auditability: which record was read, which was written, and which call did it.
The mechanics of calling a live system mid-conversation, including timeouts and what the caller hears while waiting, are covered in how a voice agent calls a live system without guessing .
This layer also carries access and retention questions, because an agent reading a customer record is a new path to customer data. Those belong with security and governance for enterprise voice AI , and the operating side with voice automation across inbound and outbound programs .
The practical sequence is to settle the phone path first, the speech and model choices second, and the enterprise connections last, because each decision narrows the options beneath it. Scoring layer by layer also keeps a comparison honest when two vendors describe the same stack differently.
Pulastya presents that stack as a four-step setup aimed at business teams rather than developers, with a browser test call before launch and transcripts, call summaries and a call dashboard once calls are running. To see the enterprise voice AI platform end to end, book a guided walkthrough .
Conclusion
Voice AI integrations work best when each part of the stack is evaluated on its own terms. Telephony carries the call, speech recognition and synthesis handle audio, the language model determines the response, and enterprise systems provide the records that make the conversation useful. The orchestration platform connects these dependencies, but it cannot eliminate the limits, latency or failure modes of the services underneath it.
Before choosing a platform, confirm its supported providers, check how it connects to your existing phone system and business applications, and test real calls for accuracy, response time, interruption handling, transfers and safe data updates. Compare provider charges, account ownership and platform fees together. A layer-by-layer evaluation makes it easier to identify problems, assign responsibility and build a voice agent that can operate reliably beyond the pilot.
Frequently Asked Questions
No. Each platform supports a specific, named set of communications providers, and the list is usually short. Twilio, Telnyx, Vonage, Plivo, Sinch, Bandwidth and Azure Communication Services all exist as options in the wider market, but that does not mean a given platform supports them. Ask each vendor for the current supported list in writing during evaluation.
Speech recognition, or speech-to-text, converts what the caller says into text the agent can act on, and detects when a turn has ended. Speech synthesis, or text-to-speech, converts the agent's reply back into audio. They are separate products from separate vendors. Recognition errors change what the agent hears; synthesis errors change what the caller hears.
That depends entirely on the platform. Some support several model providers, some are built around one or two named ones, and some do not expose the choice at all. OpenAI, Azure OpenAI, Anthropic Claude, Google Gemini, Amazon Bedrock, Meta Llama, Mistral and self-hosted open-weight models are all in the market. Treat provider choice as a question to ask, not an assumption.
Usually not. Common patterns are connecting over SIP from an existing trunk or PBX, pointing an existing number's inbound webhook at the agent platform, provisioning a separate number for a pilot, or routing a leg from an existing contact center platform. Which one applies depends on your telephony estate and on what the platform supports.
Often several parties. The communications provider bills call minutes and numbers, speech vendors bill per minute or per character, model providers bill per token, and the platform may add its own fee. Ask which charges arrive on your own provider accounts and which appear on the platform invoice, because the total looks different in each arrangement.







