When enterprises come to us at Bandwidth contemplating voice AI, they’re often trying to answer one question: which voice AI platform should I choose? It’s a reasonable question. There are many established and emerging companies that could be the right answer based on use case, geography, price, sophistication and numerous other factors. But, it may not be the most important question.
The more consequential decisions (and harder to change) are architectural. Where does AI fit in your call flow? How does it interact with your contact center? Your CRM? Your security policies? Your company footprint and use cases? What happens to context, compliance recordings, and routing logic when things don’t go according to plan?
A proof of concept (PoC) tuned in a controlled lab environment rarely reflects production economics or performance. After advising Global 2000 enterprises through these transitions, we commonly see three primary architectural patterns. Understanding their trade-offs is the difference between an adaptable AI strategy and an expensive, disruptive migration three years later.
Pattern 1A: AI behind the CCaaS
Placing AI behind an existing CCaaS platform is the most common starting point because it requires minimal upfront telephony changes as the enterprise is already routing calls to the CCaaS. Adding an AI agent behind that platform requires minimal disruption to the existing setup, since you can usually add one of these providers through configuration options on your CCaaS platform and quickly start evaluating how well it handles your use cases.

This can be a very effective deployment model for quickly evaluating how well your use cases can be handled by a particular AI platform considering factors such as your customer’s experience, your relationship experience with the provider, etc. In addition to these things however, it is a good time to consider how this would scale and perform as a commercial deployment as it can unravel your business model and desired customer experience.
Latency Stacking and the Customer Experience
In many cases, the path to the agent platform is not optimized with every call traversing multiple segments creating latency stacking: audio traverses CPaaS -> CCaaS -> AI Inference -> back. Decoding, converting speech-to-text, and re-encoding at multiple hops adds hundreds of milliseconds, causing caller hesitation and ruined conversational rhythm.
The Commercial Model
If they’re making it easy, the CCaaS providers are not likely making it cost effective to choose another vendor AI.
- Agent connection fees: CCaaS providers often charge monthly fees for the agent integration. Many enterprise customers report that these alone can significantly impact their business case.
- Conversation fees: depending upon the integration architecture, the CCaaS may preserve its presence in the end to end call flow and charge the traditional per minute fees. This is compounded with the fees charged by the Agent provider.
- One time charges: CCaaS providers may also provide one time setup or implementation charges, particularly if they involve their professional services teams for configuring routing capabilities.
Enterprises are making it pretty clear that these fees are not exceptions, but the rule, and that they add a significant unanticipated cost to their use case. While this model may be effective in evaluating your AI vendor choices, it quite likely isn’t the model you need for commercially launching and scaling your solution. Although it might be tempting to launch for expediency, keep in mind that migrations are difficult. Moving later has its own set of costs and customer experience challenges.
Pattern 1B: The CCaaS Native Agent Variant
Naturally, this is where many of the established CCaaS providers want an enterprise to be. Some have built their own AI while others have integrated with some of the same providers the enterprise may be looking at. The appeal is simplicity: fewer vendors, fewer handoff points, and a single point of contact for support. The tradeoff is that you’re betting on the CCaaS vendor’s AI roadmap and potentially missing out on selecting the best in breed for your use cases. Also lost is the opportunity to establish the relationship with the AI vendors and directly influence their roadmap. Most enterprise customers have made it clear: they want independence in their choices for CCaaS and voice AI platforms.

Pattern 2: AI in front of the CCaaS
As voice AI evolved, enterprises began placing the AI agent at the front door of the network to intercept PSTN traffic before it hits the CCaaS. The value proposition is straightforward: put the AI agent at the front of the call flow, before the CCaaS. If the agent can contain 70–80% of customer interactions, there is no need to route those calls through a contact center at all. For enterprises expecting high containment rates, this architecture makes a compelling case. But enterprises should scrutinize their first assumptions before using them to make a deployment decision.

In practice, containment rates vary. If you’re getting <35% resolution rather than the 70–80% you planned for, a significant portion of your volume is still flowing to human agents. This creates two key challenges to the business case. For one, the latency stacking model discussed when AI sits behind the CCaaS will often still apply here. Other challenges exist as well including context sharing, compliance and multi-vendor use cases.
Context Sharing Challenges
When a customer spends several minutes working through a voice agent and then gets transferred to a person who has no idea what was just discussed, it degrades customer experience. This “context cliff” fragments the customer experiences, and is one of the most consistent frustrations enterprises report can occur with this approach.
There are ways to share context across systems. One common approach is through CRM logging, but that requires every component in your ecosystem to integrate with the same CRM and write to it consistently. The more vendors involved, the more integration points that can break. In a multinational enterprise with multiple CCaaS vendors, multiple AI platforms, and business units making their own technology decisions, every component integrating with the same CRM is not a safe assumption.
Another method is through passing contextual information through the protocols that interconnect the components. While the protocols may be standard, the context passing often is not. For complex systems as mentioned above, this can create an integration problem at scale.
Compliance
Compliance is another structural issue to consider. Call recordings and transcriptions that have to traverse both the AI platform and the CCaaS either get duplicated across systems, creating correlation headaches, or media is forced through one anchor platform in a way that brings back the stacking latency problem from pattern one. Neither path gives you a clean outcome.
Multi-Vendor
Choosing an AI provider and placing them at the front of your call flow can create vendor lock in. Many enterprises want the flexibility to add alternative providers or use different providers for different use cases. This is difficult to orchestrate and creates a dependency on each provider to integrate in turn to the choices made for other components of the model such as the CCaaS or CRM.
Pattern 3: Side-by-side model with Orchestration
Forward-thinking Global 2000 enterprises—particularly those managing multi-vendor environments, global footprints, or merger and acquisition technology debt—are landing on a side-by-side architecture. In this model, an independent, software-driven carrier/CPaaS orchestration layer sits directly in front of both your AI platform(s) and your CCaaS infrastructure. Bring your own (BYO) integrations, popular in the CCaaS space for allowing enterprises to bring their own carrier, enable the enterprise to quickly and easily add their preferred vendors to the topology. BYO keeps the enterprise in direct control of their AI stack and roadmap. Orchestration, such as Bandwidth Maestro, keeps them in control of their call flows.

Inbound calls hit that orchestration layer first, where the enterprise has complete control over which system gets their calls first and provides an anchor point through which to redirect calls among platforms. This enables seamless handoffs between the AI agent and the CCaaS system for human agents without incurring latency stacking issues while keeping core features like call recording centrally located.
This pattern is especially advantageous for enterprises with complex vendor portfolios. We have customers in service industries who have already concluded that one AI platform performs better for certain use cases while another performs better for others. Without an orchestration layer at the edge, that kind of flexibility is practically impossible to implement at scale. With it, you can route different call types to different AI providers, maintain the same CCaaS for your human agents, and keep adding complexity to your call flows without the whole system becoming fragile.
Context sharing also becomes more trackable. Instead of requiring every vendor in your stack to independently integrate with the same CRM, you have an anchor point in the network layer that can apply consistent context-passing logic across all your call segments and handoffs. One set of integrations, not a growing web of pairwise connections.
The decision for the enterprise now
Even if your current question is “which vendor(s) should I choose?”, avoid commercial launch until you’ve asked the critical question – “what architecture do I need?”. Choosing any of the three patterns may be a good choice under the right conditions.
- Deploying AI behind your CCaaS is a quick and easy method to evaluate how well your AI vendor options support your use cases, satisfy your customers and meet your business objectives .
- Putting AI in front of your contact center can provide similar benefits while improving on the shortcoming of pattern 1 however it is not well positioned to handle growth and multi-vendor evolution.
- When making a long-term architectural commitment, one that needs to scale across business units, geographies, and vendor changes, the side by side – orchestration-first model provides the ultimate in customer experience, cost effectiveness and ease of integration.
The primary risk is committing to an architecture during early trials without vetting its performance at scale or against your future needs. When call volumes grow, unvetted models often trigger unexpected costs, latency issues, and broken customer experiences. Likewise, models which lack flexibility or enterprise control fail to support evolving, multi-vendor solutions and complex use cases.
The enterprises considering architecture now will land in a strong position. Alternatively, difficult and costly migrations may be the future.
David Ress is the Sr. Director, Conversational AI Ecosystem Strategy at Bandwidth. David brings decades of experience in telecom, specializing in architecting innovative voice solutions.