There is abundant discussion about AI models and architectures, but far less focus on the underlying infrastructure and data integrity that ultimately determine whether AI can be trusted, audited, and scaled in enterprise environments.
In practice, AI is only reliable if the infrastructure and data pipelines that support it are reliable, from data capture through transmission and validation. This is increasingly where the real differentiation will happen.
The reality is that AI shows up in very concrete moments: a call routed through the public telephone network to a patient, a hotel guest, or a customer waiting for service. When such an interaction takes place, the decisive factor is not only the quality of the model, but the reliability of the underlying infrastructure.
That infrastructure is telecommunications, and it is about to become one of the most strategic and contested layers of enterprise technology.
The Problem with Following the Model Wars
Judging by the way the AI industry talks about itself, you might think that only benchmarks, context windows, and a startup’s foundation model matter. This conversation is not entirely wrong; model capability matters. But it is dangerously incomplete, and that gap will surprise many organizations.
Think about what an AI voice agent actually needs to operate in a hospital, a school district, a hotel chain, or an average retail business. It needs a real phone number. It needs a SIP trunk to carry audio from the PSTN to where the model runs. It needs a session border controller to handle NAT traversal, codec negotiation, and fraud prevention. It needs routing logic sophisticated enough to escalate to a human when the agent reaches its limits. All of this must work under 200 milliseconds of round-trip latency, because that is the threshold below which a voice conversation still sounds natural to a human. Exceed that threshold, and the agent sounds broken, no matter how intelligent it is.
None of these things are built by OpenAI, Anthropic, or Google. They build the reasoning. Someone else has to build the pipeline; and building a production voice pipeline that works across every carrier, every device, and every regulatory environment is genuinely difficult. It takes years of operational experience that does not transfer from software engineering. And it turns out that the companies that have already done it are not the ones mentioned in TechCrunch.
The intelligence layer will continue to improve. Models will become cheaper, faster, and more capable over a cycle of about 18 months. The infrastructure they run on does not follow the same curve.
The historical pattern no one wants to apply here
We have seen this before. When the internet grew in the late 1990s, routing infrastructure companies—the Ciscos, bandwidth providers, and network equipment vendors—captured value through every cycle of browser wars and dot-com volatility. When mobile grew, the companies that owned spectrum and towers remained valuable through generations of smartphones. When cloud computing grew, AWS and Azure became some of the most profitable companies in the world, not by building the best application but by owning the layer every application ran on.
The pattern is not subtle: intelligence commoditizes faster than infrastructure. The application at the top changes; the pipeline it passes through is harder to build and replace. Yet every time a new platform transition happens, the conversation focuses on the intelligence layer until it is too late to secure a good position in the infrastructure layer.
We are in that window right now with AI agents and voice.
Where the market really is
The scale of what is taking shape here merits examination, because the numbers are not incremental.
The North American UCaaS market was already $41 B in 2025 and is expected to nearly double by 2032, according to Mobility Foresights. At the same time, SME communications are growing at an annual rate of 27.8 % through 2030 (Mordor Intelligence), while adoption patterns are changing just as rapidly. Telehealth is a clear example, with 86 % of physicians using it in 2021, compared with only 15 % two years earlier (Market.us / Newstrail).
Then there is the figure that deserves more attention than it gets. Organizations see a 53 % increase in revenue when UCaaS, CCaaS, and CPaaS come from a single integrated stack (Mordor Intelligence). That does not mean bundled products are marginally better. It shows that when voice infrastructure, the contact center layer, and the programmable communications layer are designed to work together, the business result is transformative.
An AI agent can then easily hand off from automation to humans, switch from voice to SMS, and retrieve conversation history across multiple channels. This is very different from assembling these components from three separate vendors. Stack integration is what actually makes AI agents work in production.
The four industries that will prove this
Not all verticals are moving at the same pace, and understanding this sequence is as important as understanding the opportunity itself.
| Vertical | Primary Use Case | Where It Stands Today | The Real Friction |
|---|---|---|---|
| Retail & SMB | AI receptionist, order tracking, loyalty follow-up | Now | Essentially none — quick decisions, no compliance burden except PCI, immediate ROI |
| Hospitality | Voice booking, guest services, multi-property communications | Next | Low — GM-level decisions, no regulated data, the ROI story closes quickly |
| Education | Emergency alerts, hybrid classroom, attendance communications | Planning | Annual budget cycles, FERPA certification — slower but near-zero churn once in place |
| Healthcare | Telehealth routing, patient scheduling, clinical call management | Under consideration | HIPAA BAAs, EHR integrations, multi-party purchasing — and a $9.8 B market |
Retail and hospitality are where this will be proven next. An AI voice agent at a restaurant chain is already capturing 30 % of missed inbound calls with a documented annual return on investment of 760 %.
These are production figures, and this is the kind of ROI that makes every other technology look timid by comparison. These deployments are happening now; they work, and they are building the baseline that will make conversations in healthcare and education much easier in 12 months.
Healthcare is a long game and the most important one. The sales cycle is truly painful; HIPAA business associate agreements (BAA), clinical IT security reviews, EHR integration requirements, and purchasing committees that include physicians, administrators, and compliance officers who do not always agree.
This friction is real, but it is also why, once a healthcare organization deploys compliant AI voice infrastructure, it almost never leaves it. The compliance work it has done with the provider and the integrations embedded in clinical workflows are switching costs that compound into something nearly permanent. Healthcare is not the first win. It is the most valuable.
The regulatory complexity that seems like an obstacle from the outside is a moat from the inside. Every HIPAA certification, every FERPA audit, every PCI‑DSS review that a voice infrastructure provider passes is one more thing a competitor must replicate before even being considered. Most don’t bother.
The part of the story engineering teams already know
There is a detail about how AI voice agents are actually built today that isn’t discussed in the business press, but that anyone building in this space encounters immediately.
A common approach to connecting large language models to real phone calls today involves frameworks like Asterisk, a proven and highly adaptable open-source framework. Used with technologies such as the OpenAI Realtime API, Google Gemini Live, and Deepgram, Asterisk offers a flexible and reliable way to connect AI-generated audio to PSTN calls. Its AudioSocket interface, for example, enables real-time streaming between AI models and live calls, making it a solid choice for developers building AI voice systems.
A growing ecosystem of open-source projects and developer communities is building production-ready AI voice agent frameworks on Asterisk. This pattern is gaining traction because it gives developers control over call handling, media routing, and integration logic in ways that many higher-level platforms abstract away.
This is strategically important because of what happens when a developer ecosystem standardizes part of the infrastructure. Enterprise adoption follows developer adoption by about 18 to 24 months. The engineers building AI voice systems on Asterisk today are the architects who will specify the infrastructure for enterprise deployments in 2027. And Sangoma, the company that has led Asterisk for the past two decades, has the deepest relationship with this developer community and the strongest influence on the framework’s future direction; it is not a startup. It is the same organization that built it.
Red Hat built an extraordinary business by doing exactly that with Linux. Elastic did it with search. The open‑source moat is one of the most durable forms of competitive advantage in technology, precisely because it does not rely on customer lock‑in. It is so deeply embedded in how people build that moving on requires starting from scratch.
What organizations should actually do now
If you are a technology or operations leader considering deploying AI agents, the question to ask is not which model to use. It is whether the voice infrastructure you are building on is production-ready, certified for compliance with your industry, and able to withstand a 5× increase in call volume once deployment goes live.
Most AI voice projects that fail in production fail at the infrastructure level. The model was solid. Latency was acceptable in testing. The compliance documentation seemed adequate. Then a real clinical environment, a multi-location retail chain, or a school district with 40 000 students pushed the system in ways the integration was not designed for, and everything fell apart.
There is another dimension that most discussions about AI agents ignore completely. When an AI agent communicates with a human, or increasingly, with another AI agent, each of these interactions generates a trail of sensitive data that must be authenticated, timestamped, and secured across the entire communications stack. Who spoke. When. What was said. Whether the voice on the line was really who it claimed to be. Whether data exchanged between agents was intercepted or altered in transit. These are not edge cases. They are the baseline requirements for any AI voice implementation in a regulated environment, and they become exponentially harder to guarantee when a company assembles voice from one vendor, messaging from another, video from a third, and network security from a fourth.
No one in that setup owns the end-to-end chain. No one can guarantee where the data resides, who has handled it, or whether the audit trail holds up under a HIPAA review or breach investigation. The compliance certificate on a single vendor’s wall covers only that vendor’s slice. The gaps between vendors, the handoffs, the API calls, the data in transit between platforms belong to no one. This accountability gap is where breaches happen, regulatory exposure accumulates, and AI agent deployments silently fail the trust requirements that enterprise and healthcare customers cannot compromise on.
A single trusted provider that owns the entire stack—voice, messaging, video, and network & security—does more than simplify purchasing. It closes the accountability gap entirely, gives enterprises a single chain of custody for every interaction, and makes compliance something that is structurally assured rather than manually assembled through contracts with vendors who each deny responsibility for what happens outside their boundaries. Sangoma is one of the very few providers on the market that can make this guarantee because it built the entire stack itself.
The organizations that succeed are those that have treated voice infrastructure as a first-order architectural decision rather than a convenience purchase. They chose vendors with carrier-grade SLA commitments, existing compliance certifications for their industry, and real operational experience maintaining voice quality at scale. They then built their AI layer on that foundation rather than the other way around.
The window to make this decision thoughtfully, before competitive pressure forces a rushed choice, is probably 12 to 18 months. After that, organizations that got the infrastructure right early will have reference deployments, cumulative compliance investments, and installed bases that their competitors will spend years trying to dislodge.
The AI agent economy is being built now. And as with every platform transition before it, the infrastructure layer will matter more, and for longer, than the intelligence layer everyone is debating right now.
At Sangoma, we have been building this infrastructure for over 40 years, long before anyone called it AI‑ready. What we are seeing now is the market finally catching up with what the voice layer has always been capable of. That does not make us comfortable. It makes us more focused!
