How AI voice agents, multi-provider telephony, CRM, SIP, SMS, and usage-based billing come together in a single multi-tenant platform
How to read this document. Every architectural claim in this case study is labeled using one of three consistent tags, so a reader can tell at a glance what is platform fact versus engineering recommendation:
[Documented Capability] — directly supported by the supplied platform specification. This is what the platform actually does.
[Recommended Architecture] — an implementation approach this case study proposes for something the specification does not fully specify. Not a claim that this exists in production.
[Conceptual Implementation] — one technically reasonable way a documented feature could work internally, offered to make the architecture concrete, not to assert that this is the exact internal implementation.
No performance benchmarks, customer counts, revenue figures, uptime numbers, or compliance certifications are claimed anywhere in this document. Where a claim would require evidence this case study does not have, it is either omitted or explicitly marked as a recommendation.
1. Executive Summary
An AI Cloud Call Center SaaS platform is a multi-tenant system that lets a business stand up AI-driven phone conversations, inbound and outbound, without assembling its own speech pipeline, telephony account, CRM, and billing system by hand. [Documented Capability] The platform covers AI voice agents (built on OpenAI GPT, Deepgram, and ElevenLabs), multi-provider telephony (Twilio, Plivo, and SIP via an ElevenLabs SIP Trunk), a visual call-flow builder, a built-in CRM, SMS campaigns, a phone-number marketplace, multi-tenant administration across three roles (Super Administrator, Business Customer, Call Center/Sales Agent), credit-based usage billing, and KYC and support-ticket workflows.
The reason this class of platform is architecturally interesting has nothing to do with any single component. Speech-to-text, text-to-speech, telephony, and CRM software all exist as mature, independent products. What is hard, and what this document focuses on, is the integration surface between them: getting a phone event, an AI decision, a CRM update, a usage record, and a credit deduction to behave as one coherent flow instead of five loosely-synced systems. That integration discipline is the actual engineering problem a platform like this solves.
At a glance, the platform is organized as five cooperating layers: a presentation layer (dashboards, an agent console with a web dialer, an embeddable website voice widget, and an admin panel), an application layer (AI agent management, the call-flow engine, CRM, campaigns, SMS, phone numbers, billing, KYC, support), an AI processing layer (Deepgram for speech-to-text, OpenAI GPT for reasoning, ElevenLabs for speech synthesis), a communication layer (Twilio, Plivo, and SIP/ElevenLabs SIP Trunk), and a data layer holding contacts, calls, leads, credits, and billing history on a relational (SQL) database. Sections 6 through 26 walk through each layer and the workflows that run across them; Sections 27 through 35 cover implementation, scaling, reliability, security, and engineering trade-offs; the closing sections cover business use cases, the revenue model, and a technical FAQ.
This case study is published by Zipprr, a company that builds white-label, source-code-owned SaaS platforms, including AI-driven products such as Zipprr AI Chat. The architectural patterns examined throughout, provider abstraction, tenant-scoped data isolation, and usage-based billing, apply broadly to AI-powered SaaS platforms of this kind, not only to call center software specifically.
Architecture Scope
| Documented Platform | Recommended Architecture | Conceptual Implementation |
|---|---|---|
| AI voice agents (OpenAI GPT, Deepgram, ElevenLabs); Twilio, Plivo, and SIP (ElevenLabs SIP Trunk) telephony; the visual call-flow builder; CRM; SMS; the phone-number marketplace; multi-tenant administration; credit-based billing; KYC; support tickets; Google/Facebook authentication; a Node.js backend on a relational (SQL) data layer. | Provider-abstraction layers, a usage ledger, tenant-scoped query enforcement, outbound call queueing, idempotency handling, observability signals, and retry/reconciliation strategy — implementation approaches this case study proposes for what the specification leaves unspecified. | Example API routes, a conceptual SQL schema, a per-call session/state model, and a possible browser voice-transport mechanism — illustrations of how a documented feature could plausibly work internally, not statements of the platform's actual internals. |
Every claim in the sections that follow is tagged with one of the three labels above, so the boundary between platform fact and engineering recommendation stays visible throughout.
2. Business & Technical Challenge
Before a platform like this exists, a business trying to run AI-assisted calling has to assemble the equivalent system itself: a telephony account that knows nothing about who is calling, a speech-to-text/text-to-speech vendor with no concept of a lead or a campaign, a CRM with no way to originate or receive a call on its own, and a billing tool that can only meter what it is told about. None of these systems were designed with each other in mind, so wiring them together is bespoke engineering work, repeated by every business that wants the same outcome.
| Challenge | Traditional Approach | Platform Approach | Technical Benefit |
|---|---|---|---|
| Running AI voice conversations | Separately integrate an STT vendor, an LLM API, and a TTS vendor, then write the orchestration code that strings them into a live conversation | [Documented Capability] AI voice agents built on Deepgram (STT), OpenAI GPT (reasoning), and ElevenLabs (TTS) | One conversational capability shared by every calling channel instead of a bespoke integration per channel |
| Placing and receiving calls | Integrate directly against one telephony provider's SDK and rebuild the integration if the provider ever changes | [Documented Capability] Telephony spans Twilio, Plivo, and SIP (including an ElevenLabs SIP Trunk) | Calling logic is not written against a single provider's API shape |
| Tracking leads and customers | A separate CRM product, manually synced (or not synced) with call activity | [Documented Capability] Built-in CRM with contacts, leads, tasks, and call notes | Call and message activity has a direct path into CRM data |
| Designing conversation logic | Hand-code IVR-style call trees or hard-code conversational logic per deployment | [Documented Capability] A visual, node-based call-flow builder | Non-developers can construct and modify call logic |
| Reaching customers across channels | Separate systems for voice and SMS, with no shared contact or campaign model | [Documented Capability] Contact, campaign, calling, and SMS capabilities within the same SaaS environment | Coordinated communication workflows across calling and SMS without separate systems |
| Acquiring phone numbers | Provision numbers directly through a telephony provider's own console | [Documented Capability] A phone-number marketplace for searching, purchasing, and assigning numbers | Number acquisition happens inside the same system that will use the number |
| Billing for usage | A flat subscription unrelated to actual volume, or a separately built metering system | [Documented Capability] Credit-based, usage-based billing | Revenue is tied to a documented usage model rather than a fixed fee |
| Serving many customer businesses | Deploy a separate instance or licensed copy per customer | [Documented Capability] Multi-tenant SaaS architecture | One platform serves many tenants without per-tenant infrastructure |
| Combining AI and human effort | Route every interaction to either a fully automated system or a fully staffed call center, with no structured handoff | [Documented Capability] A hybrid workflow: AI qualifies, a task routes the lead to a human agent | AI and human effort connect through the same CRM task structure |
| Verifying business identity | No structured verification step before higher-risk calling access is granted | [Documented Capability] A KYC submission, review, and approval workflow | Identity verification is a first-class workflow, not an informal check |
Taken together, the technical rationale for a unified platform is not any one row in this table — it is that a call event, a CRM update, an SMS send, and a credit deduction can be expressed as one connected data path instead of five separately maintained ones.
3. Architecture at a Glance
A direct-answer summary for readers, and for AI systems indexing this page, before the deeper technical sections below.
What is an AI Cloud Call Center SaaS platform? It is a multi-tenant cloud platform that lets a business run AI-handled and human-handled phone conversations, inbound and outbound, through one system that also holds the CRM data, billing, and administration those conversations depend on.
What AI technologies does it use? [Documented Capability] OpenAI GPT for conversational reasoning, Deepgram for speech-to-text, and ElevenLabs for text-to-speech.
What telephony is supported? [Documented Capability] Twilio Voice API and Plivo Voice API for standard calling, plus SIP calling through an ElevenLabs SIP Trunk for SIP-based AI agents and SIP bulk calling.
How does the AI voice pipeline work, in one sentence? A caller’s speech is transcribed by Deepgram, reasoned over by OpenAI GPT in the context of the active call flow, and answered in synthesized speech from ElevenLabs; the documented capabilities can be organized around this common flow across supported channels, though the exact internal implementation shared across channels is not documented (see Section 7).
How does multi-tenancy work? [Documented Capability] Each Business Customer operates in its own workspace, with its own AI agents, CRM records, campaigns, credits, phone numbers, and staff, administered platform-wide by a Super Administrator (Section 18).
How does billing work? [Documented Capability] Tenants purchase credit plans and consume credits based on calling and SMS usage; the exact metering unit and deduction mechanics are addressed as [Recommended Architecture] in Section 20.
How does AI hand off leads to human agents? [Documented Capability] An AI agent’s qualification outcome can generate a CRM task, which routes the lead to a Call Center/Sales Agent working from the same contact record (Section 11).
What other integrations are documented? [Documented Capability] A CRM (contacts, leads, tasks, notes), SMS (individual and bulk), a phone-number marketplace, KYC review, a support-ticket system, and Google/Facebook authentication.
Who are the platform’s user roles? [Documented Capability] Super Administrator (platform-wide operator), Business Customer (tenant owner), and Call Center/Sales Agent (works assigned leads inside a tenant).
4. Solution Overview
The platform’s architecture is built around a straightforward goal: whatever channel a customer interaction arrives on, phone, SIP line, or website widget, it should be able to reach the same documented AI capabilities and update the same CRM, so that call handling, lead data, and billing do not fork into channel-specific silos.
[Recommended Architecture] The diagram below presents this as a logical architecture — a recommended way to organize the documented capabilities into layers — rather than a claim about the platform’s internal service boundaries, which are not specified.
Figure 1. Recommended logical architecture. Every calling channel converges on the same AI processing pipeline and the same CRM/data layer, which is the structural property that keeps a lead’s history and a tenant’s credit balance consistent regardless of entry channel.
The presentation layer is covered in Section 6, the AI pipeline in Section 7, the individual calling channels in Sections 8 through 14, CRM/SMS/numbers/tenancy/roles in Sections 15 through 19, and billing/compliance/recording in Sections 20 through 22.
5. Architecture Principles
[Recommended Architecture] The following principles are not stated verbatim in the platform specification; they are the design logic this case study infers from the documented capabilities, offered as the reasoning a team building or extending a system like this would apply.
Separation of concerns. AI processing, communications, CRM, and billing are treated as distinct responsibilities. Deepgram, OpenAI GPT, and ElevenLabs each do one job in the voice pipeline (Section 7); telephony providers are interchangeable beneath the call-flow engine (Section 14); CRM and billing consume events rather than owning them.
Provider independence. Because the platform documents more than one telephony provider (Twilio, Plivo, SIP) and names distinct AI vendors for STT, reasoning, and TTS, no single external vendor relationship is structurally load-bearing for the whole platform. A vendor swap is a narrower change than a platform redesign.
Tenant-scoped data. Multi-tenancy (Section 18) implies that CRM records, calls, credits, and phone numbers belong to exactly one tenant. This case study treats tenant scoping as a cross-cutting requirement on every data-holding module, not a feature of any one module.
Event-aware processing. A call, an SMS send, a lead qualification, and a credit deduction are each treated conceptually as discrete, traceable events rather than only as row updates. This framing supports the usage-ledger and observability recommendations in Sections 20 and 33.
Human-AI continuity. The hybrid workflow (Section 11) only works if the AI agent’s output is structured enough for a human agent to use without re-listening to a call. This case study treats “AI output becomes human input” as a first-class design constraint on the CRM.
Usage traceability. Because billing is usage-based (Section 20), every usage event that can reduce a tenant’s credit balance should be attributable to a specific call, message, or action, so that billing accuracy can be reasoned about and reconciled.
These principles frame the recommendations that appear throughout this document; they are not, on their own, evidence of a specific internal implementation.
6. High-Level System Architecture
[Documented Capability] The platform organizes into five layers.
Presentation Layer. A user (Business Customer) dashboard for configuring AI agents, call flows, campaigns, contacts, numbers, billing, and KYC. An agent dashboard with an integrated web dialer for a Call Center/Sales Agent’s assigned tasks and calls. An embeddable website voice widget for visitor-facing AI conversations. An admin dashboard for the Super Administrator’s cross-tenant operations.
Application Layer. AI agent management, the call-flow engine, CRM, campaign management, SMS, phone-number management, billing and credits, KYC, and support tickets — each owning a distinct domain of platform functionality.
AI Processing Layer. Deepgram for speech-to-text, OpenAI GPT for reasoning, ElevenLabs for text-to-speech.
Communication Layer. Twilio Voice API and Plivo Voice API for standard telephony; SIP calling, including an ElevenLabs SIP Trunk, for SIP-based AI agents and bulk calling.
Data Layer. A relational (SQL) database is documented as the data store; the specific engine, schema, and ORM are not documented and are not assumed. [Recommended Architecture] A conceptual schema appears in Section 28.
| Layer | Components | Responsibility |
|---|---|---|
| Presentation | User dashboard, agent dashboard + web dialer, website voice widget, admin dashboard | Where each role interacts with the platform |
| Application | AI agent mgmt, call-flow engine, CRM, campaigns, SMS, phone numbers, billing/credits, KYC, support | Core tenant-facing business logic |
| AI Processing | Deepgram, OpenAI GPT, ElevenLabs | Converts audio to text, reasons over it, converts the response back to audio |
| Communication | Twilio, Plivo, SIP / ElevenLabs SIP Trunk | Places/receives calls; carries audio to and from the AI processing layer |
| Data | Relational (SQL) database | System of record for every module above it |
[Recommended Architecture] Grouping the application layer’s modules into an orchestration sub-layer (the call-flow engine and AI agent management directing traffic into the AI and communication layers) is a useful way to reason about the platform, but is presented here as an analytical lens, not a documented internal boundary.
7. AI Voice Agent Architecture
[Documented Capability] The platform’s core capability is an AI voice agent that can hold a spoken conversation: a caller’s speech is transcribed (Deepgram), reasoned over in context (OpenAI GPT), and answered in synthesized speech (ElevenLabs). This same capability is documented as usable across inbound calling, outbound calling, SIP calling, and the website voice widget.
Figure 2. One AI-handled conversational turn — the unit of work every calling channel in this document builds on.
Each stage has one job: Deepgram turns audio into an accurate transcript and nothing more; OpenAI GPT decides what to say next given the transcript, the agent’s configured persona, the active call-flow node, and prior turns; ElevenLabs turns that decision into natural speech in the agent’s configured voice. [Recommended Architecture] Keeping these three responsibilities separate — rather than folding transcription or synthesis into the reasoning step — is what allows any one stage to be reasoned about, monitored, or replaced independently (see Section 33 on observability).
[Conceptual Implementation] The specification does not document the exact mechanism for holding conversation state across turns within one call. A per-call session object carrying transcript history, the active call-flow node, and any data collected so far is one reasonable way this could work; it is not a documented implementation detail.
8. Inbound AI Calling
[Documented Capability] An inbound call arrives on one of the tenant’s platform-managed numbers. The platform must resolve which tenant and AI agent own that number and which call flow governs the call before AI processing begins.
Figure 3. Inbound call handling. Number and call-flow resolution happen before any AI processing, since they determine which agent and logic apply.
| Stage | What Happens | Owning Component |
|---|---|---|
| Number routing | Incoming call matched to the tenant and AI agent assigned to the dialed number | Phone-number management |
| Call-flow selection | The configured flow for that number/agent is loaded | Call-flow engine |
| Speech processing | Caller audio is transcribed | Deepgram |
| AI response | Transcript reasoned over, response generated | OpenAI GPT + call-flow context |
| Lead capture | Extracted details written toward a contact/lead | CRM |
| Call logging | Duration, outcome, timestamps recorded | Call logs |
| Recording | Audio stored where enabled | Call recording |
9. Outbound AI Calling
[Documented Capability] Outbound calling covers both a single call against one contact and a bulk campaign against many contacts. The platform documents bulk AI calling and SIP bulk calling as capabilities; it does not document specific concurrency limits, retry-attempt counts, or throughput figures, so none are stated here.
Figure 4. Outbound / bulk calling. A bulk campaign repeats this path once per contact in the assigned list; volume and scheduling are the documented difference from a single call, not the mechanism.
[Recommended Architecture] Retry policy, backoff timing, and concurrency control for a bulk campaign are not documented. A reasonable approach is a per-tenant call queue that throttles concurrent outbound attempts so one tenant’s campaign volume cannot degrade calling capacity for other tenants (expanded in Section 31).
10. Visual Call Flow Engine
[Documented Capability] The call-flow builder is a node-based workflow engine: a Business Customer constructs call-handling logic by connecting nodes on a canvas rather than writing code.
Figure 5. Example call-flow graph: a trigger starts the flow, and a condition node branches a qualified lead into a CRM action and a follow-up task.
| Node Type | Purpose | Example |
|---|---|---|
| Trigger | Starts the workflow | Incoming call received |
| Message | Plays a scripted prompt | Greeting |
| Input | Collects a caller response | Customer states their need |
| AI | Hands the turn to the AI agent | Open-ended qualification |
| Condition | Branches on an outcome | Is the lead qualified? |
| CRM | Writes/updates CRM data | Create lead, update contact |
| Action | Triggers a side effect | Create task, send SMS |
Because the flow is edited directly by a Business Customer rather than through code, changing call-handling logic, for example adding a qualifying question, does not require a software deployment.
11. Hybrid AI + Human Agents
[Documented Capability] The platform supports a hybrid model: an AI agent performs first-pass qualification and data capture, then hands qualified work to a human Call Center/Sales Agent through a CRM task, rather than requiring every interaction to be fully automated or fully staffed from the start.
Figure 6. Hybrid AI-plus-human workflow. The AI agent’s qualification output becomes the human agent’s starting context.
[Recommended Architecture] The technical value of this pattern depends on the AI’s qualification output being written as structured CRM data (not just a raw transcript), so a human agent can act on it without re-listening to the call. This is treated as a design requirement, not a documented data format.
12. Website AI Voice Agent
[Documented Capability] The platform documents an embeddable website voice widget, a website voice conversation capability, AI agent interaction through that widget, lead capture from it, and website voice logs.
Figure 7. Website voice component flow. The widget supplies the interaction channel; the platform routes the conversation into its documented AI capabilities.
[Conceptual Implementation] The specification does not state a specific browser transport or streaming protocol for the widget. A browser-based voice widget of this kind is commonly implemented with a real-time audio transport such as WebRTC or a WebSocket audio stream, but this is offered strictly as an illustrative example of how such a channel could be built, not as a documented implementation detail of this platform.
13. SIP Calling Architecture
[Documented Capability] SIP calling covers SIP-based calling, SIP AI agents, SIP bulk calling, SIP call logs, and an ElevenLabs SIP Trunk specifically.
Figure 8. SIP calling architecture. SIP is a transport for a single call or a bulk campaign into the same documented AI voice-processing capabilities used by other channels.
[Recommended SIP Architecture] SIP registration flow, codecs, RTP handling, port configuration, and session border controller topology are implementation considerations for a team building this kind of system — they are not documented platform internals, and none are asserted as fact here.
14. Telephony Provider Architecture
[Documented Capability] The platform supports more than one telephony path so calling logic is not permanently tied to one vendor’s API.
| Provider | Primary Role | Documented Use |
|---|---|---|
| Twilio Voice API | Placing/receiving calls over the public telephone network | General inbound/outbound calling, campaigns |
| Plivo Voice API | Placing/receiving calls over the public telephone network | Alternative provider for the same role |
| ElevenLabs SIP Trunk | SIP-based call transport | SIP AI agents, SIP bulk calling |
[Recommended Architecture] A provider-abstraction layer, where the call-flow engine and AI pipeline call a common interface rather than a specific vendor SDK, would explain how a tenant can be assigned a Twilio number, a Plivo number, or a SIP line with the rest of the system behaving consistently. This is a reasonable implementation pattern for what is documented, not a documented internal design.
15. CRM Architecture
[Documented Capability] The CRM is built around contacts and leads, tasks, call notes, and agent access to that data. A contact is a stored person or business the tenant may call, message, or has interacted with. A lead is a contact that has entered a qualification process. Tasks connect AI-driven qualification to human follow-up (Section 11). Call notes are what a human agent leaves after a conversation, sitting on the same contact timeline as the AI agent’s own interaction history.
Figure 9. CRM data flow. Regardless of which channel or handler produced an interaction, it resolves to one contact record that everything downstream builds on.
[Recommended Architecture] Routing AI-extracted data (name, intent, qualifying answers) into the contact/lead record as it is captured, rather than leaving it only in a transcript, is what would make later qualification, task creation, and human follow-up possible without re-deriving details from a recording. This describes a data-flow requirement the documented feature set implies, not a documented storage mechanism.
16. SMS Architecture
[Documented Capability] SMS sits alongside voice calling as a second outreach channel: individual SMS to one contact, and bulk SMS as part of a campaign, for promotions, follow-ups, or notifications.
Figure 10 (a). SMS campaign flow. A single campaign definition drives bulk delivery, with history written back to each contact the same way a call is logged.
[Recommended Architecture] Because the platform provides contact, campaign, calling, and SMS capabilities within the same SaaS environment, a single outreach effort can plausibly combine a calling step and an SMS step against the same contact list. This is a reasonable way to use the documented capabilities together, not a separately documented “multi-channel campaign” feature.
17. Phone Number Marketplace
[Documented Capability] The platform includes phone-number marketplace and management capabilities, with Twilio and Plivo numbers documented on the provider side. A Business Customer can search for available numbers, purchase a chosen number, and assign it to a specific AI agent or campaign.
[Conceptual Implementation] One reasonable way this could be implemented is that the marketplace queries provider-side number inventory (Twilio’s and Plivo’s own search APIs) at search time and completes a purchase against the selected provider, recording the resulting number in the tenant’s own inventory. This describes a plausible provisioning flow; it is not a documented statement of how number search and purchase are implemented internally.
Without an assigned number, an AI agent has no way to receive inbound calls (Section 8), and a campaign has no caller ID to place outbound calls from (Section 9).
18. Multi-Tenant SaaS Architecture
[Documented Capability] Multi-tenancy lets one platform deployment serve many Business Customers, each with tenant-scoped AI agents, CRM data, contacts, campaigns, credits, phone numbers, and staff.
Tenant-scoped; the isolation approach is a Recommended Architecture.
Figure 11. Multi-tenant architecture. The Super Administrator sits above every tenant; each tenant’s workspace (agents, CRM, calls, credits, numbers, staff) is scoped to that tenant alone.
[Recommended Architecture] The specification establishes that tenants are isolated from each other; it does not document the database-level isolation mechanism. A shared-schema model, where every tenant-owned record carries a tenant identifier enforced at the query layer, is a common and reasonable pattern for this class of system, but it is a recommendation, not a documented fact — as is any claim about exactly how isolation is enforced in this platform’s own database.
19. Roles & Permissions
[Documented Capability] Three roles run through the platform. The Super Administrator operates the SaaS business itself: provisioning and managing tenant accounts, reviewing and approving KYC submissions, overseeing support tickets platform-wide, and managing platform-level configuration. The Business Customer is the tenant account owner: configuring AI agents and call flows, managing their own CRM, launching campaigns, purchasing and assigning phone numbers, managing their own credit balance and billing, submitting KYC, and raising support tickets. The Call Center / Sales Agent works inside a Business Customer’s workspace on assigned tasks and contacts, using the agent dashboard and web dialer, and can log call notes and update the contact records for the leads they work.
| Capability | Super Administrator | Business Customer | Call/Sales Agent |
|---|---|---|---|
| Provision and manage tenant accounts | Yes | No | No |
| Review and approve KYC submissions | Yes | Submits own KYC only | No |
| Manage support tickets platform-wide | Yes | Raises own tickets only | No |
| View global call logs across tenants | Yes (platform administration) | Own tenant's logs only | Assigned contacts and agent-level calling activity |
| Manage platform-level settings | Yes | No | No |
| Configure AI agents and call flows | No | Yes | No |
| Manage CRM (contacts, leads) | No | Yes | Works assigned contacts/leads |
| Launch and manage campaigns | No | Yes | No |
| Purchase and assign phone numbers | No | Yes | No |
| Manage credits and billing | No | Yes | No |
| Use the agent dashboard and web dialer | No | Yes (as a workspace user) | Yes |
| Log call notes on assigned contacts | No | Yes | Yes |
[Recommended Architecture] The Super Administrator’s cross-tenant visibility, as documented, is scoped to platform administration functions — global call logs, KYC review, support management, and platform settings — rather than to unrestricted read/write access into every tenant’s day-to-day CRM or campaign data. This case study does not assert that the Super Administrator has unrestricted access to all tenant business data; that scope is not documented and should not be assumed by an implementer.
20. Credit & Usage-Based Billing
[Documented Capability] The platform supports credit-based, usage-based billing for calling and SMS activity: a tenant selects a credit plan and purchases credits through a payment gateway. The exact metering and deduction lifecycle — precisely when and how credits are consumed relative to a call or message, the billing unit, and the price per minute or per SMS — is not documented; none of those specifics are asserted here. The Usage Ledger recommendation below describes one reasonable way to implement this accurately.
Figure 10 (b). Credit and billing flow. Usage recording and credit deduction are shown as adjacent steps here because that is the design goal a usage-based billing system needs to satisfy — see the ledger recommendation below for how that goal could be met accurately.
[Recommended Architecture]: Usage Ledger. A durable usage ledger is a reasonable way to keep metering auditable. Each entry would conceptually record: the tenant, the usage type (e.g., a call-minute or an SMS send), a reference to the originating event (a call ID or message ID), the quantity consumed, the calculated cost, the resulting credit deduction, a timestamp, and a reconciliation state (pending, applied, or flagged for review). Representing usage this way, rather than only updating a running balance, is what would let a billing discrepancy be traced back to the event that caused it. This is a design recommendation, not a documented data structure.
21. KYC & Compliance Considerations
[Documented Capability] KYC verification is a structured workflow: a Business Customer submits identity information, a Super Administrator reviews it, and the outcome, along with platform-configured KYC settings, determines what the tenant is permitted to do. The workflow covers submission, verification/admin review, KYC settings, and KYC requests as a standing record.
On compliance, stated plainly: this document does not claim that the KYC workflow, or any other capability described here, automatically satisfies any country’s specific legal requirements. Regulatory obligations around identity verification, call recording, and outbound messaging vary by jurisdiction and use case. Call recording and outbound communication must be configured by the tenant in line with the laws applicable to their own operation; documented recording and calling capabilities are not, on their own, evidence of compliance with any particular legal regime.
22. Call Recording & Call Logs
[Documented Capability] The platform provides call logging and supports call recording where enabled.
[Conceptual Implementation]
Call Event → Call Record → Optional Recording (if enabled) → CRM / Call History → Usage Metering
A call log entry (duration, outcome, timestamps) can reasonably be expected for every completed call, since it is the basis for CRM history and usage metering; a recording is stored only where recording is enabled for that call and tenant configuration. This lifecycle is offered as one reasonable way the documented logging and recording capabilities relate to each other, not as a guarantee that every call always produces every listed artifact. As with Section 21, retention and legality of recordings is a tenant configuration responsibility.
23. End-to-End AI Call Workflow
Pulling the preceding sections together, here is the conceptual path a single contact takes through the platform’s documented capabilities:
Figure 14. End-to-end workflow, from an existing contact through AI conversation, CRM capture, logging, and human follow-up.
This figure references, rather than repeats, the mechanics already covered: the AI conversational turn is Figure 2 (Section 7), the CRM data flow is Figure 9 (Section 15), and the human handoff is Figure 6 (Section 11).
24. Human Agent Workflow
The human side of the hybrid model (Section 11), as a named sequence: a CRM contact is identified for follow-up because an AI qualification or a Business Customer’s own process flagged it; the task appears on the agent dashboard alongside the contact’s full history, including any prior AI interaction; the agent places or receives the call through the integrated web dialer; the human conversation happens; call notes are written to the same contact timeline the AI agent’s own interactions live on; a further follow-up task is created if the outcome warrants it. This workflow depends on the CRM architecture (Section 15) for context and the web dialer (Section 6) for execution.
25. Website Voice Workflow
Restating Section 12’s component flow as a named workflow: a website visitor engages the embedded voice widget; that starts an AI voice conversation using the same documented AI capabilities described for other channels in Figure 2; whatever the AI agent captures becomes a lead-capture event written into the CRM; a follow-up task is created where the interaction warrants it, closing the loop back into the hybrid AI-plus-human model (Section 11).
26. Technical Data Flow
The data flow underlying every workflow above can be organized consistently regardless of entry point:
Figure 15. Technical data flow. A voice interaction can produce a CRM outcome such as a contact or lead update, while supported usage can produce a billing/credit event, both traceable back to the same call record.
| Data Element | Produced By | Consumed By |
|---|---|---|
| Transcript segment | Deepgram | OpenAI GPT (reasoning), call record |
| AI response text | OpenAI GPT | ElevenLabs (synthesis), call record |
| Call record | Telephony/SIP layer + AI pipeline | Call logs, CRM, billing/usage |
| Lead/contact update | AI-extracted data or human agent notes | CRM, task assignment |
| Usage event | Every completed call or SMS send | Billing and credit ledger |
| Credit deduction | Billing service, keyed to a usage event | Tenant's credit balance |
27. API / Backend Examples
[Documented Capability] The specification states the platform’s backend is built on Node.js. Every example below is an illustrative Node.js implementation — not a documented production endpoint. They are written to show how the concepts in this document could plausibly be implemented with reasonable production practices (validation, tenant scoping, error handling, idempotency), not to document a real API contract, real route names, or real request/response shapes.
AI Agent API — illustrative example
// Illustrative Node.js implementation — not a documented production endpoint.
// Creates a new AI voice agent configuration scoped to the requesting tenant.
app.post('/api/agents', authenticateTenant, async (req, res) => {
const { name, persona, voiceProvider, telephonyChannel, callFlowId } = req.body;
if (!name || !persona || !callFlowId) {
return res.status(400).json({ error: 'name, persona, and callFlowId are required' });
}
const ALLOWED_CHANNELS = ['twilio', 'plivo', 'sip'];
if (telephonyChannel && !ALLOWED_CHANNELS.includes(telephonyChannel)) {
return res.status(400).json({ error: `telephonyChannel must be one of ${ALLOWED_CHANNELS.join(', ')}` });
}
try {
// req.tenantId is attached by authenticateTenant; every write below is
// scoped to it so one tenant's request can never create or touch
// another tenant's data (see Section 18).
const callFlow = await CallFlowService.findByIdForTenant(callFlowId, req.tenantId);
if (!callFlow) {
return res.status(404).json({ error: 'callFlowId not found for this tenant' });
}
const agent = await AgentService.create({
tenantId: req.tenantId,
name,
persona,
voiceProvider: voiceProvider || 'elevenlabs', // ElevenLabs is the documented TTS provider
telephonyChannel,
callFlowId,
});
return res.status(201).json({ agent });
} catch (err) {
req.logger?.error('agent_create_failed', { tenantId: req.tenantId, error: err.message });
return res.status(500).json({ error: 'Unable to create agent' });
}
});
// ── Start Call API — illustrative example ──
// Illustrative Node.js implementation — not a documented production endpoint.
// Initiates an outbound call. Uses an idempotency key so a client retry
// after a timeout cannot place the same call twice.
app.post('/api/calls', authenticateTenant, async (req, res) => {
const { agentId, contactId, channel, idempotencyKey } = req.body;
if (!agentId || !contactId || !idempotencyKey) {
return res.status(400).json({ error: 'agentId, contactId, and idempotencyKey are required' });
}
try {
const existing = await IdempotencyStore.find(req.tenantId, idempotencyKey);
if (existing) {
return res.status(200).json({ call: existing.result, replayed: true });
}
const agent = await AgentService.findByIdForTenant(agentId, req.tenantId);
const contact = await CrmService.findContactForTenant(contactId, req.tenantId);
if (!agent || !contact) {
return res.status(404).json({ error: 'Agent or contact not found for this tenant' });
}
// Recommended Architecture: check the usage ledger's projected balance
// before committing to a call that will consume credits (Section 20).
const hasCredits = await BillingService.hasSufficientCredits(req.tenantId, 'call');
if (!hasCredits) {
return res.status(402).json({ error: 'Insufficient credits to place this call' });
}
// Delegate to the provider adapter matching the requested channel
// (Twilio, Plivo, or SIP) so this route is not written against one
// vendor's SDK (see Section 14's provider-abstraction recommendation).
const call = await TelephonyService.placeCall({
tenantId: req.tenantId,
agentId: agent.id,
contactId: contact.id,
channel: channel || agent.telephonyChannel,
});
await IdempotencyStore.record(req.tenantId, idempotencyKey, call);
return res.status(202).json({ call });
} catch (err) {
req.logger?.error('call_initiate_failed', { tenantId: req.tenantId, error: err.message });
return res.status(500).json({ error: 'Unable to initiate call' });
}
});
// ── Lead Capture API — illustrative example ──
// Illustrative Node.js implementation — not a documented production endpoint.
// Called by the AI reasoning step when a call flow's AI node determines
// a caller should be captured as a lead.
app.post('/api/leads', authenticateTenant, async (req, res) => {
const { contactId, callId, qualification, extractedFields } = req.body;
const ALLOWED_QUALIFICATIONS = ['hot', 'warm', 'cold']; // illustrative values only
if (!contactId || !callId || !ALLOWED_QUALIFICATIONS.includes(qualification)) {
return res.status(400).json({ error: 'contactId, callId, and a valid qualification are required' });
}
try {
const lead = await CrmService.upsertLead({
tenantId: req.tenantId,
contactId,
callId,
qualification,
extractedFields, // e.g. { intent: 'pricing question' }
});
// Task creation is a separate, explicit step so call-flow logic controls
// whether a qualified lead actually needs human follow-up (Section 11).
if (qualification === 'hot') {
await TaskService.createFollowUp({ tenantId: req.tenantId, leadId: lead.id });
}
return res.status(201).json({ lead });
} catch (err) {
req.logger?.error('lead_capture_failed', { tenantId: req.tenantId, error: err.message });
return res.status(500).json({ error: 'Unable to capture lead' });
}
});
// ── Usage & Credit Deduction API — illustrative example ──
// Illustrative Node.js implementation — not a documented production endpoint.
// Writes a usage-ledger entry (Section 20) and deducts the corresponding
// credits. Usage is written before the deduction is confirmed so a partial
// failure is recoverable rather than silently lost (see Section 32).
app.post('/api/usage', authenticateTenant, async (req, res) => {
const { usageType, referenceId, quantity } = req.body; // e.g. usageType: 'call_minute' | 'sms'
if (!usageType || !referenceId || !quantity || quantity <= 0) {
return res.status(400).json({ error: 'usageType, referenceId, and a positive quantity are required' });
}
try {
const ledgerEntry = await UsageLedger.record({
tenantId: req.tenantId,
usageType,
referenceId,
quantity,
state: 'pending',
});
const cost = await BillingService.calculateCreditCost(usageType, quantity);
const updatedBalance = await BillingService.deductCredits({
tenantId: req.tenantId,
amount: cost,
ledgerEntryId: ledgerEntry.id,
});
await UsageLedger.markState(ledgerEntry.id, 'applied');
return res.status(200).json({ deducted: cost, balance: updatedBalance });
} catch (err) {
// A deduction failure is flagged for reconciliation, not silently dropped.
req.logger?.error('usage_deduction_failed', { tenantId: req.tenantId, referenceId, error: err.message });
return res.status(500).json({ error: 'Usage recorded; deduction failed and was flagged for reconciliation' });
}
});
// ── Webhook Handling — illustrative example ──
// Illustrative Node.js implementation — not a documented production endpoint.
// A generic handler for inbound call-event webhooks from a telephony
// provider (Twilio- or Plivo-style status callbacks).
//
// Recommended security concept: verify the webhook signature against the
// provider's documented signing scheme before trusting the payload. This
// is a general recommendation, not a description of an implemented check.
app.post('/webhooks/call-events', async (req, res) => {
const signatureValid = verifyProviderSignature(req); // recommended concept, not a documented mechanism
if (!signatureValid) {
return res.status(401).json({ error: 'Invalid webhook signature' });
}
const { providerCallId, status, tenantId } = req.body;
try {
switch (status) {
case 'completed':
await CallService.markCompleted(providerCallId);
// Recommended Architecture: process the documented usage event
// according to the platform's configured billing rules and timing.
await BillingService.recordUsageFromCall(providerCallId); // idempotent on providerCallId
break;
case 'no-answer':
case 'busy':
case 'failed':
await CallService.markUnsuccessful(providerCallId, status);
break;
default:
await CallService.logUnhandledStatus(providerCallId, status);
}
// Acknowledge receipt so the provider does not keep retrying a
// webhook the platform has already processed.
return res.status(200).json({ received: true });
} catch (err) {
await WebhookFailureLog.record({ providerCallId, tenantId, error: err.message });
return res.status(500).json({ received: false });
}
});
28. Conceptual Database Schema — Illustrative Only
This is a conceptual schema inferred from the documented feature set, not the platform’s actual production schema. Tables are ordered so that every table a foreign key references already exists above it, and foreign keys are added with ALTER TABLE after the base tables are created, avoiding forward-reference errors. The JSON column type used below is illustrative; the actual supported type (a native JSON/JSONB type or a validated TEXT column) depends on the SQL engine selected, which is not documented.
— Conceptual SQL Schema — Illustrative Only. Not a documented production schema.
-- Conceptual SQL Schema — Illustrative Only. Not a documented production schema.
-- 1. Base tables with no dependencies
CREATE TABLE tenants (
id BIGINT PRIMARY KEY,
business_name VARCHAR(255) NOT NULL,
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE credit_plans (
id BIGINT PRIMARY KEY,
name VARCHAR(255),
credit_amount INT,
price_amount DECIMAL(10,2)
);
-- 2. Tables that depend only on tenants
CREATE TABLE users (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
role VARCHAR(50) NOT NULL, -- 'super_admin' | 'business_customer' | 'agent'
email VARCHAR(255) NOT NULL, -- unique per tenant, not globally (see constraint below)
auth_provider VARCHAR(50), -- 'google' | 'facebook'
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE call_flows (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
name VARCHAR(255),
definition JSON NOT NULL -- node/edge graph (Section 10); JSON type is illustrative, adapt to the selected SQL engine
);
CREATE TABLE contacts (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
full_name VARCHAR(255),
phone_number VARCHAR(32),
email VARCHAR(255),
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE phone_numbers (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
number VARCHAR(32) NOT NULL,
provider VARCHAR(50) -- 'twilio' | 'plivo'
);
CREATE TABLE credits (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL UNIQUE,
balance DECIMAL(12,2) NOT NULL DEFAULT 0
);
CREATE TABLE kyc_requests (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
status VARCHAR(50) DEFAULT 'pending', -- 'pending' | 'approved' | 'rejected'
submitted_at TIMESTAMP,
reviewed_by BIGINT,
reviewed_at TIMESTAMP
);
CREATE TABLE support_tickets (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
raised_by BIGINT NOT NULL,
subject VARCHAR(255),
status VARCHAR(50) DEFAULT 'open',
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
-- 3. Tables that depend on users
CREATE TABLE agents_staff (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
user_id BIGINT NOT NULL,
display_name VARCHAR(255)
);
-- 4. Tables that depend on agents_staff / call_flows
CREATE TABLE ai_agents (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
name VARCHAR(255) NOT NULL,
persona TEXT,
voice_provider VARCHAR(50) DEFAULT 'elevenlabs',
telephony_channel VARCHAR(50), -- 'twilio' | 'plivo' | 'sip'
call_flow_id BIGINT NOT NULL
);
-- 5. Tables that depend on contacts / ai_agents
CREATE TABLE calls (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
ai_agent_id BIGINT,
contact_id BIGINT NOT NULL,
channel VARCHAR(50), -- 'twilio' | 'plivo' | 'sip' | 'website_widget'
direction VARCHAR(20), -- 'inbound' | 'outbound'
status VARCHAR(50),
started_at TIMESTAMP,
ended_at TIMESTAMP
);
CREATE TABLE sms_messages (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
contact_id BIGINT NOT NULL,
direction VARCHAR(20),
body TEXT,
sent_at TIMESTAMP
);
-- 6. Tables that depend on calls
CREATE TABLE call_logs (
id BIGINT PRIMARY KEY,
call_id BIGINT NOT NULL UNIQUE,
duration_seconds INT,
outcome VARCHAR(50)
);
CREATE TABLE call_recordings (
id BIGINT PRIMARY KEY,
call_id BIGINT NOT NULL,
storage_url VARCHAR(1024),
recorded BOOLEAN DEFAULT FALSE -- subject to tenant configuration, Section 21
);
CREATE TABLE call_notes (
id BIGINT PRIMARY KEY,
call_id BIGINT NOT NULL,
author_agent_id BIGINT,
note TEXT,
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE leads (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
contact_id BIGINT NOT NULL,
qualification VARCHAR(50), -- illustrative values, e.g. 'hot' | 'warm' | 'cold'
source_call_id BIGINT,
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE tasks (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
lead_id BIGINT NOT NULL,
assigned_agent_id BIGINT,
status VARCHAR(50) DEFAULT 'open',
due_at TIMESTAMP
);
-- 7. Usage / billing ledger tables (Section 20)
CREATE TABLE usage_ledger (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
usage_type VARCHAR(50), -- 'call_minute' | 'sms'
reference_id BIGINT,
quantity DECIMAL(10,2),
credits_deducted DECIMAL(10,2),
reconciliation_state VARCHAR(20) DEFAULT 'pending', -- 'pending' | 'applied' | 'flagged'
occurred_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE billing (
id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
credit_plan_id BIGINT NOT NULL,
amount_paid DECIMAL(10,2),
paid_at TIMESTAMP
);
-- 8. Foreign keys added after every referenced table exists
ALTER TABLE users ADD CONSTRAINT fk_users_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE users ADD CONSTRAINT uq_users_tenant_email UNIQUE (tenant_id, email);
ALTER TABLE call_flows ADD CONSTRAINT fk_callflows_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE contacts ADD CONSTRAINT fk_contacts_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE phone_numbers ADD CONSTRAINT fk_numbers_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE credits ADD CONSTRAINT fk_credits_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE kyc_requests ADD CONSTRAINT fk_kyc_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE kyc_requests ADD CONSTRAINT fk_kyc_reviewer FOREIGN KEY (reviewed_by) REFERENCES users(id);
ALTER TABLE support_tickets ADD CONSTRAINT fk_tickets_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE support_tickets ADD CONSTRAINT fk_tickets_user FOREIGN KEY (raised_by) REFERENCES users(id);
ALTER TABLE agents_staff ADD CONSTRAINT fk_staff_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE agents_staff ADD CONSTRAINT fk_staff_user FOREIGN KEY (user_id) REFERENCES users(id);
ALTER TABLE ai_agents ADD CONSTRAINT fk_aiagents_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE ai_agents ADD CONSTRAINT fk_aiagents_callflow FOREIGN KEY (call_flow_id) REFERENCES call_flows(id);
ALTER TABLE calls ADD CONSTRAINT fk_calls_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE calls ADD CONSTRAINT fk_calls_agent FOREIGN KEY (ai_agent_id) REFERENCES ai_agents(id);
ALTER TABLE calls ADD CONSTRAINT fk_calls_contact FOREIGN KEY (contact_id) REFERENCES contacts(id);
ALTER TABLE sms_messages ADD CONSTRAINT fk_sms_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE sms_messages ADD CONSTRAINT fk_sms_contact FOREIGN KEY (contact_id) REFERENCES contacts(id);
ALTER TABLE call_logs ADD CONSTRAINT fk_calllogs_call FOREIGN KEY (call_id) REFERENCES calls(id);
ALTER TABLE call_recordings ADD CONSTRAINT fk_recordings_call FOREIGN KEY (call_id) REFERENCES calls(id);
ALTER TABLE call_notes ADD CONSTRAINT fk_notes_call FOREIGN KEY (call_id) REFERENCES calls(id);
ALTER TABLE call_notes ADD CONSTRAINT fk_notes_agent FOREIGN KEY (author_agent_id) REFERENCES agents_staff(id);
ALTER TABLE leads ADD CONSTRAINT fk_leads_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE leads ADD CONSTRAINT fk_leads_contact FOREIGN KEY (contact_id) REFERENCES contacts(id);
ALTER TABLE leads ADD CONSTRAINT fk_leads_call FOREIGN KEY (source_call_id) REFERENCES calls(id);
ALTER TABLE tasks ADD CONSTRAINT fk_tasks_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE tasks ADD CONSTRAINT fk_tasks_lead FOREIGN KEY (lead_id) REFERENCES leads(id);
ALTER TABLE tasks ADD CONSTRAINT fk_tasks_agent FOREIGN KEY (assigned_agent_id) REFERENCES agents_staff(id);
ALTER TABLE usage_ledger ADD CONSTRAINT fk_usage_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE billing ADD CONSTRAINT fk_billing_tenant FOREIGN KEY (tenant_id) REFERENCES tenants(id);
ALTER TABLE billing ADD CONSTRAINT fk_billing_plan FOREIGN KEY (credit_plan_id) REFERENCES credit_plans(id);
29. Data Lifecycle
[Conceptual Implementation] Four lifecycles recur throughout this document:
Voice data. Call → AI/human processing → call log → optional recording (Section 22).
CRM data. Contact → interaction (AI or human) → qualification → task → follow-up (Sections 15, 11, 24).
Billing data. Usage event → pricing calculation → credit deduction → billing history, ideally traced through a usage ledger (Section 20).
KYC data. Submission → admin review → decision → retained record of the outcome (Section 21).
Each lifecycle is a conceptual synthesis of the documented capabilities that touch it, not a documented internal pipeline. They are restated here as lifecycles, rather than repeated as separate architecture descriptions, to avoid duplicating Sections 7–22.
30. Architecture Decisions & Trade-offs
[Recommended Architecture] Each row below states a design decision this class of platform embodies, why it exists, the benefit it buys, and the trade-off it accepts. These are architectural judgments this case study draws from the documented feature set, not statements the specification makes explicitly.
| Architecture Decision | Why It Exists | Technical Benefit | Trade-off |
|---|---|---|---|
| Separate STT / LLM / TTS layers (Deepgram / OpenAI GPT / ElevenLabs) | Each stage of a voice conversation has a different failure mode and a different vendor market | Provider independence per stage; a slowdown or outage in one stage is isolable | Three external dependencies in the critical path of every live call, each with its own latency and reliability profile |
| Multi-provider telephony (Twilio, Plivo, SIP) | Coverage, pricing, and SIP relationships differ by tenant and geography | Reduced single-vendor lock-in; a tenant is not stuck with one provider's rates or coverage | More integration surface area to build and test against three transport paths instead of one |
| Visual, node-based call-flow engine | Business Customers need to change conversation logic without engineering involvement | Faster iteration on call logic; no deployment needed for a workflow change | The workflow engine itself becomes a piece of infrastructure that must be built, versioned, and kept reliable |
| Built-in CRM (rather than a bolt-on integration) | AI and human agents need to work from the same contact history | Continuity between AI qualification and human follow-up (Section 11) | A larger application surface than a pure calling product; the CRM must be maintained as core, not optional |
| Multi-tenant SaaS model | One codebase needs to serve many independent businesses economically | Shared platform economics; one deployment amortized across tenants | Tenant isolation becomes a hard requirement at every layer, not just the database (Section 18) |
| Credit-based usage billing | Revenue should track actual consumption of calling/SMS resources rather than a flat fee | Billing that scales naturally with tenant usage | Requires accurate, auditable metering (Section 20) — a billing bug here is a trust problem, not just a bug |
| Website voice widget as a distinct channel | Lead capture should not require a phone number at all | An additional, low-friction lead channel embedded directly on a tenant's site | Adds browser/channel complexity (audio capture, session handling) beyond the telephony-based channels |
| Hybrid AI-plus-human workflow | Not every qualified lead should be closed by an AI agent alone | Combines automation's reach with a human's judgment on higher-value interactions | Requires the AI's output to be structured well enough for a human to act on without re-listening to the call |
31. Scalability Considerations
[Recommended Architecture] No specific performance numbers, throughput figures, or infrastructure products are documented for this platform; none are asserted here. The categories below describe where scaling pressure would concentrate, based on what each documented capability does, and a corresponding design direction — mentioning a technology (a queue, a cache, a specific cloud provider) only as an example of a category of tool a team might choose, never as a claim that this platform uses it.
| Scaling Pressure | Failure Risk if Unaddressed | Recommended Design Response |
|---|---|---|
| Real-time call concurrency | Growing latency or dropped turns as concurrent AI conversations increase | Treat each pipeline stage (STT/LLM/TTS) as independently scalable; monitor per-stage latency (Section 33) |
| AI provider latency | A slow reasoning or synthesis response reads to the caller as an unresponsive agent | Set per-stage timeout budgets and a defined fallback response (Section 32) |
| Telephony provider limits | Outbound campaigns throttled or rejected by a provider's own rate limits | Provider-abstraction layer (Section 14) so load can shift across Twilio/Plivo/SIP |
| Bulk campaign scheduling | One tenant's large campaign starves calling capacity for others | Per-tenant queueing/throttling on outbound call placement (Section 9) |
| Tenant fairness under shared infrastructure | Noisy-tenant effects degrade service for smaller tenants | Resource quotas or fair-share scheduling scoped by tenant identifier |
| Call recording storage growth | Binary storage grows unbounded with call volume | Lifecycle/retention policy tied to tenant configuration (Section 21), not indefinite retention by default |
| CRM query growth | Contact/lead/task queries slow down as tenant data accumulates | Indexing strategy aligned to tenant-scoped query patterns; pagination on list views |
| Usage-event throughput | Usage ledger writes lag behind actual call/SMS volume | Asynchronous, queued usage recording decoupled from the call path itself |
| Billing consistency under load | Concurrent usage events cause race conditions in credit deduction | Idempotent, ledger-based deduction (Section 20) rather than only updating a running balance in place |
| Webhook bursts from telephony providers | A burst of simultaneous call-completion webhooks overwhelms synchronous handling | Queue-and-process webhook ingestion, acknowledging receipt before processing completes |
| Single-provider outage | An AI or telephony vendor outage stops all affected calls platform-wide | Provider isolation per stage/channel (Sections 7, 14) limits blast radius, but does not eliminate it |
32. Reliability & Failure Handling
Each documented calling and billing capability has a corresponding failure mode worth designing around explicitly.
Figure 16. Failure-handling decision tree. Each failure category is handled at the stage where it occurs, so one failed stage cannot silently corrupt a downstream record such as a usage or credit deduction.
Telephony provider failure (a call fails to place) is addressed by the retry/fallback logic already described in Section 9’s outbound flow. Speech recognition failure (no usable transcript) means the AI agent has nothing to reason over; prompting the caller to repeat or escalating to a human agent are reasonable conceptual fallbacks. AI reasoning failure should not leave a caller in silence; a fallback response or a graceful call end are the alternatives. Speech synthesis failure means the AI has a response it cannot speak; retrying synthesis or falling back to a pre-recorded prompt are options. Call interruption should still leave the call marked accurately, with whatever partial transcript exists preserved. Webhook failures are logged and retried with backoff (illustrated in Section 27). Credit deduction failure is flagged for reconciliation rather than silently ignored, since silently dropping a usage record misstates a tenant’s actual consumption. SMS provider failure is handled by re-queuing for a later attempt.
33. Observability & Operations
[Recommended Architecture] A system with this many external dependencies and billing implications needs visibility into more than just “is the server up.” No specific monitoring tool or vendor is documented as part of this platform; the signals below are what a team operating a system like this would want visibility into, regardless of tooling choice.
| Signal | Example | Why It Matters |
|---|---|---|
| Call lifecycle state | Calls stuck in "ringing" or "in-progress" beyond an expected duration | Indicates a stuck pipeline stage or a lost provider callback |
| AI processing latency per stage | Time spent in Deepgram / OpenAI GPT / ElevenLabs per turn | Isolates which pipeline stage is degrading caller experience |
| Telephony failure rate | Spike in failed call placements from one provider | Signals a provider outage or a provider-side rate limit being hit |
| STT/TTS error rate | Empty transcripts or synthesis failures | Distinguishes an AI-quality issue from a telephony issue |
| Webhook processing lag | Growing delay between webhook receipt and processing completion | Signals a downstream bottleneck (CRM write, billing update) |
| Usage-event throughput vs. call volume | Usage events lagging behind completed calls | Early warning that billing may undercount consumption |
| Credit reconciliation backlog | Entries stuck in a "flagged" state in the usage ledger | Direct signal of billing accuracy risk (Section 20) |
| SMS delivery status | Rising bounce/failure rate from an SMS provider | Signals a provider-side delivery problem affecting campaigns |
| Tenant activity anomalies | A sudden spike in one tenant's call volume | Distinguishes legitimate growth from a misconfigured campaign or abuse |
| Support ticket volume by category | A cluster of tickets referencing the same capability | Early signal of a platform-wide issue before it is otherwise detected |
34. Security Architecture
[Recommended Architecture] Organizing the documented security-relevant capabilities into categories, without asserting any compliance certification:
Identity. [Documented Capability] Google Login and Facebook Login are the documented authentication paths.
Authorization. Role-based permissions enforce the boundaries in Section 19, so a Call Center/Sales Agent’s session cannot reach Business Customer administrative functions, and a Business Customer’s session cannot reach another tenant’s data or the Super Administrator’s platform-wide functions.
Tenant isolation. Discussed architecturally in Section 18 as a Recommended Architecture; enforcing it consistently would also serve as a security boundary, preventing one tenant’s session or API access from reaching another tenant’s contacts, calls, or billing records.
Secrets. Provider API credentials (Twilio, Plivo, OpenAI, Deepgram, ElevenLabs) need to be stored so that a compromise scoped to one tenant’s context cannot expose platform-wide credentials.
Webhooks. Signature or token-based verification, illustrated conceptually in Section 27, should confirm a webhook genuinely originates from the claimed provider before it is acted on.
Sensitive data. Contact details, call transcripts, and call recordings warrant access restricted to the owning tenant and the roles within it permitted to view them.
Billing. Credit balances and payment history warrant similarly restricted access, scoped to the owning tenant and platform administration.
Auditability. Usage events and administrative actions (KYC decisions, ticket handling) benefit from being individually traceable, which the usage-ledger recommendation in Section 20 is partly designed to support.
This document makes no compliance claims. It does not state or imply SOC 2, ISO 27001, GDPR, HIPAA, or PCI DSS compliance, or any other certification; none of that is documented in the specification.
35. Technical Complexity
What makes a platform like this genuinely difficult to build is not any single documented feature — it is the intersection of several hard problems at once, which is what separates this from a feature brochure and makes it an engineering case study.
Real-time audio orchestration. Three external AI services (Deepgram, OpenAI GPT, ElevenLabs) must be chained within a live call’s latency budget, per conversational turn, for every concurrent call. Multi-provider communications. Twilio, Plivo, and SIP each have distinct APIs, failure modes, and call-state semantics that a provider-abstraction layer must reconcile. AI latency variability. A reasoning step’s response time is not fully within the platform’s control, yet it directly shapes caller-perceived responsiveness. Call-state management. A call’s state (ringing, connected, in a specific call-flow node, ended) must stay consistent across telephony events, AI processing, and CRM writes happening concurrently. Workflow execution. The visual call-flow engine must execute a Business Customer’s own node graph reliably and predictably, including branches and CRM/action side effects. Tenant isolation. Every one of the above must be correctly scoped per tenant, under concurrent load from many tenants simultaneously. Usage metering. Every unit of consumption that could affect billing must be captured reliably, without unintended duplication or loss — not missed (lost revenue) and not double-counted (overcharging). Billing consistency. Credit deductions must remain accurate under concurrent usage events without double-deducting or under-deducting. Webhook reliability. Provider callbacks arrive asynchronously and must be processed reliably despite retries, using idempotency and reconciliation safeguards to avoid unintended duplication. Recording storage. Binary call recordings accumulate continuously and must be retained, and made unavailable, in line with each tenant’s own configuration. AI-to-human handoff. The AI’s qualification output must be structured well enough that a human agent picking up a task does not need to re-derive context from a transcript.
None of these problems is exotic on its own; the difficulty is that a platform like this must solve all of them simultaneously, and consistently, across every tenant and every channel at once.
36. Business Use Cases
[Documented Capability, applied] The platform can support workflows such as the following. These are potential applications of the documented capabilities, not claims of actual customer deployments or results.
| Use Case | How the Platform's Documented Capabilities Apply |
|---|---|
| AI call center | AI voice agents (Section 7) handle inbound and outbound conversations |
| Sales automation | Outbound campaigns (Section 9) combined with CRM lead qualification (Section 15) |
| Customer support | Inbound AI calling (Section 8) paired with support tickets for issues needing escalation |
| Lead qualification | AI-driven qualification written into the CRM (Sections 8, 15) |
| Telemarketing | Bulk outbound and SIP bulk calling campaigns (Sections 9, 13) |
| Appointment follow-ups | Hybrid AI-plus-human workflow with task creation (Section 11) |
| Customer surveys | Outbound calling with AI-driven data capture into the CRM (Sections 9, 15) |
| Marketing campaigns | Combined calling and SMS campaign management (Sections 9, 16) |
| SMS marketing | Bulk SMS campaigns (Section 16) |
| Website voice assistance | Embeddable website voice widget (Section 12) |
| AI receptionist | Inbound AI calling with call-flow-driven routing (Section 8) |
| Agency AI calling services | Multi-tenant architecture (Section 18), letting an agency operate calling services for multiple end clients as separate tenants |
| White-label call center SaaS | Multi-tenant, role-based architecture (Sections 18, 19) as the underlying structure for a reseller model |
37. SaaS Revenue Model
[Documented Capability] The specification documents a usage-oriented monetization model: revenue is generated through credit packages purchased by Business Customers (Section 20), usage-based billing metered against calling and SMS consumption, AI calling and bulk calling as the primary consumption drivers, SMS marketing as a second consumption channel, and phone numbers acquired through the marketplace (Section 17) as an additional line of paid provisioning. No specific pricing figures, margins, or revenue projections are documented, and none are estimated here.
38. Technology Stack
| Technology / Integration | Role |
|---|---|
| Node.js | Backend application runtime |
| JavaScript | Application development language |
| JSON | Data exchange and configuration format (including call-flow definitions) |
| HTML / CSS | Front-end structure and styling |
| SQL | Relational data layer |
| OpenAI GPT | AI conversation reasoning |
| Deepgram | Speech-to-text |
| ElevenLabs | Text-to-speech and SIP trunk |
| Twilio | Voice API / telephony provider |
| Plivo | Voice API / telephony provider |
| ElevenLabs SIP Trunk | SIP infrastructure |
| SMTP | Email delivery |
| Google Login | Authentication |
| Facebook Login | Authentication |
39. Final Architecture Diagram
[Recommended Logical Architecture] — this consolidates the documented capabilities into the layered model used throughout this case study; it is an analytical framing, not a claim about the platform’s internal service boundaries.
Figure 17. Final consolidated architecture: actors and the surfaces they use, the application modules those surfaces expose, the AI capabilities every calling channel can draw on, the communication layer that carries audio, and the shared data layer beneath everything.
40. Technical Outcomes
Describing the architecture’s outcomes in neutral technical terms, without invented metrics:
Unified communications. Voice calling across Twilio, Plivo, and SIP, plus SMS, share the same contact, campaign, and logging structures rather than requiring manual reconciliation across separate systems. AI-plus-human workflows. The hybrid model gives a structured way to combine automated first-contact handling with human follow-up through one CRM task system. Centralized CRM. Contacts, leads, tasks, and call notes live in one place regardless of whether an AI agent or a human produced them. Automated lead capture. Data an AI agent extracts during a conversation is written directly into the CRM as part of the call itself. Multi-channel entry points. Phone calls, SIP, bulk campaigns, SMS, and a website widget all feed into the same underlying AI and CRM capabilities. Configurable call flows. The visual call-flow builder lets call-handling logic change without a software deployment. Usage-aligned monetization. Credit-based billing ties revenue to platform consumption rather than a flat, usage-unrelated fee. Centralized administration. The Super Administrator role gives platform-wide oversight of tenants, KYC, and support, distinct from any tenant’s own operation. Extensible integrations. Because telephony, AI, and voice functions are each handled by a distinct, named provider rather than a single monolithic implementation, a given provider relationship can, in principle, be reconsidered without redesigning the surrounding call-flow, CRM, or billing logic.
41. Frequently Asked Questions
1. What is an AI Cloud Call Center SaaS platform?
A multi-tenant cloud platform that lets a business run AI-handled and human-handled inbound and outbound calling, alongside SMS, CRM, and billing, without assembling those systems separately.
2. How does an AI voice agent process a phone call?
Caller speech is transcribed by Deepgram, reasoned over by OpenAI GPT in the context of the active call flow and prior turns, and answered with ElevenLabs-synthesized speech — repeating each turn until the call ends (Section 7).
3. Which AI technologies are used in the platform?
[Documented Capability] OpenAI GPT for reasoning, Deepgram for speech-to-text, and ElevenLabs for text-to-speech and its SIP trunk.
4. Does the platform support inbound and outbound AI calling?
[Documented Capability] Yes — inbound calling resolves the number to a tenant, agent, and call flow before starting the AI conversation (Section 8); outbound calling and bulk campaigns place calls against a CRM contact list using the same documented AI capabilities (Section 9).
5. How does SIP integrate with AI voice agents?
[Documented Capability] SIP, including an ElevenLabs SIP Trunk, is a transport alongside Twilio and Plivo; an AI agent configured for SIP calling, or a SIP bulk campaign, hands its audio to the same documented STT/reasoning/TTS capabilities used by other channels (Section 13).
6. Can AI calls create CRM leads automatically?
[Documented Capability] Yes — an AI agent’s call-flow logic can capture qualifying details during a call and write them toward a CRM contact or lead record (Sections 10, 15).
7. How does multi-tenant architecture work?
[Documented Capability] Each Business Customer operates in an isolated workspace with its own AI agents, CRM, campaigns, credits, numbers, and staff, with a Super Administrator providing platform-wide administration (Section 18).
8. How does credit-based billing work?
[Documented Capability] A tenant purchases a credit plan; calling and SMS usage consume credits from that balance. The exact metering unit and deduction timing are not documented; Section 20 proposes a usage-ledger approach as a recommended architecture for implementing this accurately.
9. Can human agents work alongside AI agents?
[Documented Capability] Yes — the hybrid model lets an AI agent qualify a lead and create a CRM task that routes it to a Call Center/Sales Agent, who works from the AI’s captured context via the agent dashboard and web dialer (Section 11).
10. Can an AI voice agent be embedded into a website?
[Documented Capability] Yes — a website voice widget lets a visitor start a live AI voice conversation directly on a Business Customer’s site, feeding the same CRM and the same documented AI capabilities as a phone call (Section 12).
11. Is this document describing an existing, certified, or benchmarked product?
No. This case study documents a feature specification and labels every claim as a documented capability, a recommended architecture, or a conceptual implementation. It does not assert performance benchmarks, customer results, or compliance certifications.
42. Final Technical Summary
| Area | Documented Capability |
|---|---|
| AI reasoning | OpenAI GPT |
| Speech-to-text | Deepgram |
| Text-to-speech | ElevenLabs |
| Telephony | Twilio, Plivo |
| SIP | ElevenLabs SIP Trunk |
| CRM | Contacts, leads, tasks, notes |
| SMS | Individual + bulk |
| Call-flow logic | Visual, node-based builder |
| SaaS model | Multi-tenant |
| Billing | Credits + usage-based pricing |
| KYC | Submission + admin review |
| Authentication | Google Login, Facebook Login |
| Backend | Node.js |
| Data layer | SQL (relational) |
| Support | Ticketing system |
| Roles | Super Administrator, Business Customer, Call Center/Sales Agent |
What the architecture demonstrates. The platform brings AI voice, multi-provider telephony, a visual workflow engine, CRM, SIP, SMS, usage-based billing, and multi-tenant administration together into one SaaS operating model rather than a collection of separately integrated tools. The engineering significance is not in any individual capability — each exists elsewhere as a standalone product — but in making a call event, a CRM update, an SMS send, and a credit deduction behave as one connected, tenant-scoped data flow. That integration discipline, sustained across every documented channel and role, is what this case study has aimed to make explicit, distinguishing throughout between what the platform is documented to do and what remains a reasonable, clearly labeled recommendation for how it could be built.
Ready to Build an AI-Powered SaaS Platform?
The architecture patterns walked through in this case study — provider abstraction, tenant-scoped data isolation, a usage ledger for accurate billing, and a hybrid AI-plus-human workflow — aren’t unique to call center software. They’re the same engineering patterns Zipprr applies across its own white-label SaaS products, so a team evaluating this kind of build doesn’t have to start from a blank page.
- Explore a live example of these patterns → Zipprr AI Chat
- Browse the full product catalog → Zipprr Products
- See how this fits your own use case → Book a free demo
Every Zipprr product is a one-time purchase, owned outright with complete source code included, so there is no monthly platform fee or vendor lock-in on top of what this case study has already described.



