Best AI Voice Agents in 2026: 14 Platforms for Phone Support, Sales, and Booking

Turn this article into takeaways for your work.
Each assistant summarizes the article only for you and suggests best practices for your work.
Vapi and Retell AI lead for engineering teams building a custom phone agent from scratch, PolyAI and Parloa lead for enterprise contact centers that want a fully managed rollout, Sierra leads for teams that already run a chat agent and want the same brain answering the phone, and Regal and Thoughtly lead for sales and revenue teams doing outbound qualification and booking. This guide ranks 14 AI voice agent platforms on end-to-end response latency, the one number that decides whether a caller feels like they're talking to something useful or just hangs up, with every price checked against each vendor's own pricing page in August 2026.
A voice agent is not the same product as a voice generator, and mixing the two up is the fastest way to buy the wrong tool. Text-to-speech platforms turn a script into audio. A voice agent listens on a live call, decides what to say next, and can call a tool mid-conversation to look up an order, check a calendar, or transfer to a human. If narration or voice cloning is what you actually need, see our best AI voice generators guide instead, a different product class entirely. This article sits inside a wider AI agent platforms buying guide, and if your team would rather build a voice agent than buy one, what is an AI agent and the AI voice call agent blueprint cover the build side.
Updated August 2026: What Changed
- Air.ai, the company that popularized the "AI sales rep on the phone" pitch in 2023, is gone. The FTC sued it in August 2025 and settled for an $18 million judgment in March 2026, banning its operators from marketing business opportunities. Its voice service went dark, and the air.ai domain now belongs to an unrelated company. It's off this list.
- Aircall acquired Vogent in May 2026 to strengthen its own native voice agent, but Vogent's developer platform is still open for new signups as a standalone product as of this writing.
- ElevenLabs cut Conversational AI pricing roughly in half this year, down to 8 to 10 cents a minute depending on plan, and added pay-as-you-go billing on top of its credit system.
- Synthflow's live pricing page no longer shows the self-serve monthly plans some third-party trackers still list. As of August 2026 it's Enterprise-only, with contracts starting around $30,000 a year.
- Cartesia, known mainly as a low-latency voice model supplier, shipped Line, a full voice agent development platform, putting it in more direct competition with Vapi and Retell instead of just powering their text-to-speech.
Key Facts
- Conversational AI is projected to cut contact center agent labor costs by $80 billion in 2026, though Gartner's own research notes only about 1 in 10 agent interactions will be fully automated this year, per Gartner.
- The global AI voice agents market is estimated at $3.51 billion in 2026, projected to reach $35.24 billion by 2033, a 39% compound annual growth rate, per Grand View Research.
- 91% of customer service leaders say they're under executive pressure to implement AI in 2026, per a Gartner survey of 321 leaders.
- Humans hand off a conversational turn in roughly 200 milliseconds. Replies past 800ms start to feel delayed, and past 1,500ms a call feels broken, a bar most voice agent platforms are still chasing in production, not just in demos, per Telnyx's 2026 latency benchmark.
- A fully loaded US onshore call center agent costs about $26 an hour in 2026 (roughly $0.43 a minute), versus $6 to $14 an hour offshore and $12 to $18 an hour nearshore, the baseline every voice agent's per-minute rate is really competing against, per Call Force Global.
- The FCC confirmed in February 2024 that AI-generated voices count as an "artificial voice" under the TCPA, so outbound AI voice calls need the same prior express consent as a prerecorded robocall, per the FCC.
Quick Comparison Table
| Platform | Best For | Starting Price | Key Strength | Key Limitation |
|---|---|---|---|---|
| Vapi | Developers building custom agents | $0.05/min platform fee + provider costs | Deepest telephony and model flexibility | Real all-in cost is opaque until you configure it |
| Retell AI | Production-grade custom agents | $0.07-0.31/min pay-as-you-go | Itemized pricing, strong support reputation | Effective rate climbs fast with premium LLMs |
| Bland AI | Flat-rate agents, no surprise bills | $0/mo + $0.14/min (Start tier) | One rate covers LLM, STT, and TTS | Per-minute rate is higher than usage-based rivals |
| ElevenLabs Agents | Voice quality and language coverage | From $0.08/min | 70+ languages, top-rated voice realism | LLM costs currently absorbed, likely to become billable |
| PolyAI | Enterprise contact centers, managed | Custom (reported $150K+/yr) | Deep enterprise deployment experience | No self-serve tier, sales-led only |
| Parloa | High-volume contact centers | Custom (reported $300K+/yr) | Outcome-based pricing tied to resolution | Built for 2M+ calls/year, not smaller teams |
| Sierra (Voice) | One agent across chat and phone | Custom, sales-led | Unifies voice with an existing chat agent | Voice is new, thinner track record than the chat product |
| Regal | Outbound sales and retention calling | Custom, 25-seat/100K-minute minimum | Built for structured outbound campaigns | Not designed for small teams or low call volume |
| Synthflow | Managed build, now enterprise-only | Custom, from $30,000/yr | Fully managed STT, LLM, and TTS stack | Self-serve tiers no longer published |
| Thoughtly | Real estate, insurance, education | $500/mo unlimited minutes (10 concurrent) | Unlimited minutes at a flat rate | Concurrency-capped, not built for huge call volume |
| Phonely | Cheapest real self-serve entry point | Free (100 min/mo); $50/mo Starter | Lowest cost of entry with a free tier | Overage rates ($0.25-0.35/min) add up fast |
| Cartesia Line | Lowest-latency agent development | Free; Pro $5/mo + $0.06/min | Fastest voice model in the category | Newer agent platform, thinner track record |
| Deepgram Voice Agent API | Developers who bring their own LLM | $0.056-0.075/min pay-as-you-go | Cheapest per-minute rate, BYO discounts | Developer API only, no built-in agent builder UI |
| Vogent | Self-improving agents via API | $0.09/min, from $0.05/min at volume | Agents that retest themselves against past calls | Now owned by Aircall, standalone future less certain |
Latency and Barge-In: The Number That Actually Decides If a Call Feels Human
Two things matter more than any feature list on a pricing page: how long a caller waits for the agent to start talking, and what happens when they talk over it.

| Platform | Vendor-Claimed Latency | Independent Test Result | Notes |
|---|---|---|---|
| Bland AI | ~400ms (methodology undisclosed) | 850ms median / 1,180ms p95 (Tested Media) | Largest gap between claim and measured result |
| Retell AI | "As low as ~600ms" | 680ms median / 920ms p95 (Tested Media) | Vendor claim and independent test roughly agree |
| Vapi | p50 under 500ms, p95 under 800ms (target) | 720ms median / 1,050ms p95 (Tested Media) | A target, not a guarantee; varies with model choice |
| ElevenLabs Agents | ~75ms (TTS generation only) | 1.73s p50 for a full agent turn (Cekura) | The 75ms figure is the voice model alone, not a full turn |
| Cartesia Line | Sub-90ms TTS/STT models (Sonic 3.5 / Ink 2) | Not independently benchmarked at press time | Fastest underlying models in the category by spec |
But raw speed isn't the whole story. Humans hand off a turn in conversation in about 200 milliseconds without thinking about it; once a voice agent's reply crosses 800ms, callers start to notice, and past 1,500ms the call feels broken, not futuristic, per Telnyx's 2026 benchmark comparing vendor claims against independent tests from Tested Media and Cekura. The gap between what a vendor advertises and what gets measured on a fixed stack is often 300 to 900 milliseconds, which is the difference between a natural pause and a caller repeating themselves.
Barge-in, letting a caller interrupt the agent mid-sentence and having it actually stop and listen, is the other half of feeling human. Vogent and Cartesia both build their pitch around turn-taking and interruption handling specifically, not just raw response speed, because a fast agent that talks over the caller is worse than a slightly slower one that yields the floor cleanly. When you pilot any of these platforms, test barge-in with a real interruption, not a silent pause, before you trust the latency number on its marketing page.
1. Vapi: Developer-First Infrastructure for Custom Voice Agents
Vapi's bet is that voice agents belong in code, not a drag-and-drop canvas. It's infrastructure: you pick your own STT, LLM, and TTS providers (or use Vapi's defaults), wire in tools, and Vapi handles the real-time orchestration, telephony, and call state. That flexibility is why it's become the reference platform competitors measure themselves against.
The tradeoff is that Vapi's headline $0.05/min platform fee is not what you pay. STT, LLM, and TTS are billed at provider cost on top, and most teams land between $0.07 and $0.25 a minute once they've picked real models, climbing past $0.30 with premium voices and large context windows. Budget for the full stack, not the platform fee alone.
| What you get | What you don't |
|---|---|
| Full control over STT, LLM, and TTS provider choice | Headline $0.05/min price is a fraction of the real cost |
| Native and BYO Twilio telephony, including full SIP trunking | No flat all-in rate, every model swap changes your bill |
| Large community and reference integrations | Enterprise pricing and support are custom-quoted |
| 10 concurrent calls included, scalable by line | Steeper setup than a no-code builder like Thoughtly |
Pricing: Build plan usage-based at $0.05/min platform fee plus provider costs at cost (real-world total commonly $0.07-0.25/min); Scale (Enterprise) uses a fixed platform fee plus committed volume pricing, custom-quoted. HIPAA add-on $2,000/mo, zero data retention add-on $1,000/mo. See vapi.ai/pricing.
Best for: Engineering teams that want full control over their voice stack and are comfortable assembling STT, LLM, and TTS themselves
2. Retell AI: Vapi's Closest Rival, Tuned for Reliability
Retell AI competes on the same turf as Vapi, developer-first infrastructure for a custom agent, but leans harder on production reliability and support response time. The base voice engine rate covers speech-to-text and Retell's own infrastructure; you then choose an LLM and TTS provider on top, each billed separately and itemized in the console rather than buried in a pass-through line.
Where Retell pulls ahead for some teams is the breadth of its telephony options: buy a number directly, import one from an existing carrier, or connect a SIP trunk from Twilio, Telnyx, or Vonage, with local numbers available across the US, UK, Canada, and Australia. Independent latency tests put it roughly on par with Vapi, occasionally slightly ahead.
| What you get | What you don't |
|---|---|
| Transparent, itemized pricing by component | $0.07/min headline is still just the voice engine layer |
| Three telephony paths: buy, import, or SIP trunk | Premium LLM choices (GPT-5.5-class) push cost toward $0.31/min |
| $10 in free credits to test before committing | Enterprise tier starts around $8,000, a real budget line |
| Consistently rated highly for support responsiveness | Less telephony documentation depth than Vapi's SIP guides |
Pricing: Pay-as-you-go from $0.07/min (voice infrastructure and STT) up to roughly $0.31/min with a premium LLM and add-ons like knowledge base (+$0.005/min) or PII removal (+$0.01/min); Enterprise from about $8,000. See retellai.com/pricing.
Best for: Product teams shipping a production phone agent who want itemized, predictable component pricing over a bundled rate
3. Bland AI: Flat Per-Minute Pricing With Nothing Bolted On
Bland's whole pitch is one number. Where Vapi and Retell make you assemble and price a stack, Bland bundles the LLM, speech-to-text, and text-to-speech into a single per-minute rate with no token charges and no per-provider pass-through, so the invoice matches what sales quoted.
That simplicity costs more per minute than usage-based competitors past the free Start tier, but it buys predictability finance teams like: a platform fee plus a flat rate. Bland is also one of the few platforms here stating HIPAA eligibility with a signed BAA outright, not gated behind a custom enterprise conversation.
| What you get | What you don't |
|---|---|
| One flat rate covers LLM, STT, and TTS entirely | Per-minute rate is higher than assembling your own stack |
| HIPAA-eligible with a signed BAA | Bring-your-own-telephony still needed to avoid transfer fees |
| Platform fee drops your per-minute rate at higher tiers | $499/mo Scale tier needed before the best rate kicks in |
| No token charges or model-provider surprise bills | Custom rates only available at Enterprise volume |
Pricing: Start $0/mo + $0.14/min; Build $299/mo + $0.12/min; Scale $499/mo + $0.11/min; Enterprise custom. Call transfers billed separately at $0.03-0.05/min depending on tier. See bland.ai/pricing.
Best for: Teams that want one predictable rate and no surprise model-provider bill at the end of the month
4. ElevenLabs Agents: Voice Quality and 70+ Languages From the TTS Leader
ElevenLabs built its name on some of the most natural-sounding text-to-speech on the market, covered in our best AI voice generators guide, and ElevenAgents is its conversational layer on top: real-time voice and chat agents that call tools, deployed across telephony, web, WhatsApp, and SMS from the same account.
The pitch for teams already on ElevenLabs is straightforward: no new vendor for voice quality, plus native connectors into Twilio, Genesys, Zendesk, and Salesforce. The catch is that LLM costs are currently absorbed by ElevenLabs rather than billed, which the company has said won't last forever, so today's per-minute rate is likely to rise once that changes.
| What you get | What you don't |
|---|---|
| Widest language coverage on this list (70+ languages) | LLM costs absorbed now, expected to become billable |
| Native connectors: Twilio, Genesys, Zendesk, Salesforce | No published flat-rate plan, credits and PAYG mixed |
| Same voice quality as ElevenLabs' TTS product | Full per-minute breakdown isn't on the main pricing page |
| 15 minutes free to test before paying | Best rates need an annual Business or Enterprise plan |
Pricing: Calls from $0.10/min on Creator and Pro plans (down about 50% after a 2026 price cut), $0.08/min on annual Business plans, lower at Enterprise; 15 minutes free to start. See elevenlabs.io/agents.
Best for: Teams that already use ElevenLabs for voice quality and want the broadest language coverage without a second vendor
The next four platforms trade self-serve flexibility for a managed, enterprise-led rollout. If that's the lane you're actually in, our best enterprise AI agent platforms roundup covers the wider category beyond just voice.
5. PolyAI: Enterprise Contact Center Voice AI, Fully Managed
PolyAI sells outcomes to enterprises, not a self-serve console. There's no free tier and no public rate card. Deployments go through a sales-led process where PolyAI's own team handles conversation design, ongoing model tuning, and 24/7 support as part of the contract, closer to a managed service than software you configure yourself.
That's the right model for banks, hotel groups, and retailers that need a voice agent live across hundreds of locations without staffing in-house prompt engineers, which is why Marriott, Caesars Entertainment, and Foot Locker use it. It's the wrong model if you want to test an idea this week for a few hundred dollars.
| What you get | What you don't |
|---|---|
| Managed conversation design and ongoing tuning included | No self-serve signup, sales cycle before you can test it |
| Proven at enterprise scale across 100+ customers | Reported minimum contracts start around $150,000 a year |
| Deep experience in banking, hospitality, and retail | Not a fit for testing an idea before committing budget |
| 24/7 support included as part of the managed contract | Pricing opacity makes early-stage budgeting harder |
Pricing: Not published; sales-led enterprise contracts. Third-party estimates (reported, not vendor-confirmed) put minimums around $150,000/yr, with mid-tier deployments running $10,000-$20,000/month. PolyAI raised an $86 million Series D in December 2025 at a $750 million valuation. See PolyAI's funding announcement.
Best for: Enterprise contact centers in banking, hospitality, and retail that want a managed rollout, not a DIY build
6. Parloa: Outcome-Based Pricing for High-Volume Contact Centers
Parloa's pricing logic is the most unusual on this list: instead of charging per seat or per minute, it charges per successfully resolved conversation. If a call escalates to a human, you don't pay the AI resolution price for it, which aligns the vendor's incentive with actually solving the caller's problem instead of just keeping them on the line.
That model only makes sense at real scale. Parloa is explicit that it's built for organizations handling roughly 2 million calls a year, not a 30,000-order-a-month e-commerce brand, and the "build it with us" engagement means budgeting $300,000 or more annually once implementation and integration are included.
| What you get | What you don't |
|---|---|
| Outcome-based pricing tied to resolved conversations | Not viable below roughly 2 million calls a year |
| Full agent lifecycle: design, test, scale, optimize | Reported budgets start at $300,000+ annually |
| Financing options for multi-year contracts (Capchase) | Build-it-with-us model, not a self-serve product |
| Strong fit for regulated, high-stakes call volume | Long implementation runway versus a self-serve platform |
Pricing: Outcome-based, priced per resolved conversation rather than per seat or per minute; not publicly listed. Budget guidance in Parloa's own materials is $300,000+ annually for platform licensing plus implementation. See Parloa's pricing guide.
Best for: Large enterprises running millions of calls a year that want to pay for resolutions, not minutes
7. Sierra (Voice): One Agent Across Chat and Phone
Sierra, co-founded by former Salesforce co-CEO Bret Taylor, built its reputation on chat-first "agent OS" deployments for brands like Sonos, Casper, and SiriusXM before extending the same agent to phone calls in early 2026. The pitch is consolidation: one agent, one set of policies and integrations, answering a customer whether they text, chat on the website, or call in.
Sierra's voice agent slots in front of or behind an existing IVR rather than replacing it, with escalation to a human handled through AI-generated summaries. And because voice is newer than Sierra's chat product, ask any demo for phone-specific reference customers, not just chat metrics.
| What you get | What you don't |
|---|---|
| Same agent and knowledge base across chat and voice | Voice is newer than Sierra's established chat product |
| Works in front of or behind an existing IVR | No public pricing, sales-led enterprise motion only |
| AI-generated summaries for human handoff | Reported budgets run $150,000-$350,000+ a year |
| Backed by a $950M Series E at a $15.8B valuation | Built for large enterprises, not a self-serve trial |
Pricing: Not published; outcome-based, sales-led enterprise contracts, with reported estimates of $150,000-$350,000+ annually. Sierra raised a $950 million Series E in May 2026 at a $15.8 billion valuation. See TechCrunch's coverage.
Best for: Enterprises that already run Sierra, or a similar agent OS, for chat and want the same agent to answer the phone
8. Regal: Outbound Sales and Retention Calling at Scale
Regal grew up as an AI-assisted dialer for human sales and retention teams before layering in fully autonomous voice agents, and that lineage shows: it's built for structured outbound campaigns, appointment reminders, and win-back calls in insurance, healthcare, and financial services, where the phone still drives revenue.
Regal doesn't publish pricing, and its own pricing page states a real floor: a minimum of 25 seats and 100,000 minutes of calls a month before Regal will scope a contract. That rules it out for a small team testing an idea, but it's a reasonable bar for a call center already running that volume with human agents and looking to blend in AI.
| What you get | What you don't |
|---|---|
| Built specifically for structured outbound campaigns | Published minimum of 25 seats and 100K minutes/month |
| Native CRM integrations and real-time analytics | No pricing visible until you're through a sales call |
| Agents that work across calls, SMS, and chat | Not built for a small team's first pilot |
| Backed by $100M+ raised, led by Emergence Capital | Positioned for revenue teams, not general customer support |
Pricing: Not published; usage-based custom contracts with a stated minimum of roughly 25 seats and 100,000 call minutes per month. See regal.ai/pricing.
Best for: Sales, collections, and retention teams running structured outbound campaigns at real call volume
9. Synthflow: No-Code Builder Now Behind an Enterprise Wall
Synthflow started as a no-code voice agent builder aimed at agencies and SMBs, stacking GPT-4o reasoning on ElevenLabs voices inside a drag-and-drop flow editor, the same category our best no-code AI agent builders roundup covers more broadly. That positioning still shows up across the web in third-party reviews quoting monthly self-serve plans.
But it doesn't show up on Synthflow's own site anymore. As of August 2026, synthflow.ai/pricing lists a single Enterprise tier starting at $30,000 a year, scoped around call volume, concurrency, telephony setup, and integrations, with no visible self-serve signup. If you've seen a cheaper monthly plan quoted elsewhere, verify it directly with Synthflow before budgeting around it; our own check of the live page found no such tier.
| What you get | What you don't |
|---|---|
| Fully managed STT, LLM, and TTS, nothing to assemble | No visible self-serve tier as of August 2026 |
| Native telephony, SIP trunking, or approved carriers | $30,000/yr floor prices out small teams and solo builders |
| CRM, calendar, and knowledge-source integrations | Third-party sites still quote stale self-serve pricing |
| MSA/DPA support and a formal security review | Slower sales cycle than a true self-serve competitor |
Pricing: Enterprise-only as of August 2026, contracts starting at $30,000 annually, final price scoped to volume and integrations. See synthflow.ai/pricing.
Best for: Enterprises that want a fully managed voice AI build without assembling STT, LLM, and TTS themselves
10. Thoughtly: Unlimited Minutes for Real Estate, Insurance, and Education
Thoughtly targets high-consideration, lead-driven industries specifically: real estate brokerages, insurance carriers, mortgage lenders, and education enrollment teams, where a missed inbound call is a missed sale. The agent qualifies a lead by voice, follows up by SMS, email, or WhatsApp, books the meeting, and writes it back to the CRM without a human touching it. Our AI lead qualifier agent blueprint covers the same workflow if you'd rather build it yourself.
Its pricing stands out here: the Flex plan is a flat $500 a month for unlimited voice minutes, capped at 10 concurrent calls, rather than metering by the minute. For a brokerage with bursty call volume around listings, that flat rate can beat a per-minute competitor; for a call center needing 100 simultaneous conversations, the concurrency cap forces an upgrade to a custom Scale quote.
| What you get | What you don't |
|---|---|
| Flat $500/mo for unlimited voice minutes | Capped at 10 concurrent calls on the Flex tier |
| Full workflow: call, SMS, email, WhatsApp, CRM sync | Per-minute overage pricing isn't published |
| 34+ languages and 200+ integrations | Enterprise tier requires 1M+ minutes/mo or a regulated vertical |
| Built specifically for real estate, insurance, education | Less general-purpose than Vapi or Retell for other verticals |
Pricing: Flex $500/mo (unlimited minutes, up to 10 concurrent calls); Scale and Enterprise custom-quoted, with Enterprise requiring 1M+ minutes/month or a regulated industry. See thoughtly.com/pricing.
Best for: Real estate, insurance, mortgage, and education teams that want unlimited calling at a flat monthly rate
11. Phonely: The Cheapest Real Self-Serve Entry Point
Phonely is one of the few platforms on this list a small business can sign up for and use the same day without a sales call. The free plan includes 100 minutes a month, a real phone number, and unlimited concurrent calls, and the $50/mo Starter tier is priced for a business answering its own phones, not a contact center.
The tradeoff shows up in overage: past your plan's included minutes, Phonely charges $0.25 to $0.35 a minute, well above the $0.05-0.10/min self-serve rates elsewhere on this list, so it works best for predictable, moderate call volume rather than bursty or high-volume lines. Enterprise customers can negotiate down to about 5 cents a minute.
| What you get | What you don't |
|---|---|
| Free tier with a real phone number, no card required | Overage runs $0.25-0.35/min, high next to usage-based rivals |
| Self-serve signup, live the same day | Best per-minute rate (5 cents) needs an Enterprise deal |
| Unlimited concurrent calls even on the free plan | Advanced API and CRM integrations gated to the Pro tier |
| HIPAA BAA available at Enterprise | Thinner enterprise track record than PolyAI or Parloa |
Pricing: Free (100 min/mo, 1 number); Starter $50/mo or $33/mo annual (250 min, $0.25-0.35/min overage); Pro $150/mo or $100/mo annual (750 min, $0.30/min overage); Enterprise from roughly $0.05/min, custom. See phonely.ai/pricing.
Best for: Small businesses and solo operators who want a free-to-start agent answering their phones without a sales cycle
12. Cartesia Line: Built on the Fastest Voice Model in the Category
Cartesia spent its first few years as an infrastructure supplier, the low-latency speech models other platforms sometimes license, before shipping Line, a full code-first platform for building voice agents directly. The pitch: the company building the fastest underlying model should also build the best agent platform.
Line agents default to Cartesia's own Sonic 3.5 text-to-speech and Ink 2 speech-to-text models, both built on a state-space-model architecture the company says gets sub-90ms generation, faster than the transformer-based models most competitors license from third parties. Voximplant added native support for Line agents in February 2026, extending where a Cartesia-built agent can actually take calls.
| What you get | What you don't |
|---|---|
| Fastest underlying TTS/STT models in the category by spec | Newer full-agent platform, thinner production track record |
| Free tier to start, self-serve through Enterprise | Reasoning still needs a connected LLM, not bundled into Sonic |
| $0.06/min calls plus $0.014/min for a Cartesia number | Smaller voice library than ElevenLabs or Murf |
| 40+ language support with consistent low latency | Fewer telephony integrations than Vapi or Retell today |
Pricing: Free; Pro $5/mo; Startup $49/mo; Scale $299/mo; Enterprise custom. Line calls run $0.06/min plus $0.014/min for telephony on a Cartesia-provided number. See cartesia.ai/pricing.
Best for: Developers who want to build on the lowest-latency voice model available and are comfortable on a newer agent platform
13. Deepgram Voice Agent API: Bring Your Own LLM, Pay by the Hour
Deepgram's Voice Agent API skips the builder UI entirely. It's speech-to-text, LLM orchestration, and text-to-speech combined into one real-time WebSocket API, priced by connection time, for teams that want to write their own agent logic instead of configuring someone else's canvas.
The pitch is cost and flexibility at the infrastructure layer: bring your own LLM or TTS model and the published rate drops further, and Deepgram positions itself as roughly a quarter cheaper than ElevenLabs' Conversational AI and three-quarters cheaper than OpenAI's Realtime API for comparable usage. It's the right layer for a team with engineers who'd rather write orchestration code than learn a third-party agent builder.
| What you get | What you don't |
|---|---|
| STT, LLM orchestration, and TTS in a single real-time API | No drag-and-drop builder, this is a developer API only |
| Cheapest published per-minute rate on this list | You write the agent logic and call flow yourself |
| BYO LLM or BYO TTS discounts on top of base pricing | No native telephony, bring your own carrier |
| $200 in free credits for new accounts | Best suited to teams with in-house engineering capacity |
Pricing: Pay-as-you-go from $0.056-0.075/min (Standard); Growth plan (prepaid annual) from $0.051/min; BYO LLM and BYO TTS configurations priced lower; Advanced tier $0.122-0.163/min. $200 free credit for new accounts. See deepgram.com/pricing.
Best for: Engineering teams that want the cheapest real-time voice infrastructure and are building their own agent logic on top
14. Vogent: Self-Improving Agents, Now Backed by Aircall
Vogent's differentiator is a testing loop most competitors treat as an afterthought: run a new agent version against every past call automatically, flag where it would have handled something differently, and fix the prompt before it talks to a live caller. Paired with purpose-built models for IVR navigation, scheduling, and order processing, it's built for teams that want an agent that improves over time, not one that launches and stays static.
Aircall, the business phone platform, acquired Vogent in May 2026 to strengthen its own native voice agent, and as of this writing Vogent's standalone developer platform is still accepting new signups. Ask directly about product roadmap and continuity before committing to a multi-year build on it.
| What you get | What you don't |
|---|---|
| Automated evaluation suite testing new versions against past calls | Acquired by Aircall in May 2026, standalone future less certain |
| Purpose-built models for IVR nav, scheduling, order processing | Smaller ecosystem and community than Vapi or Retell |
| Self-serve from $0.09/min, down to $0.05/min at volume | Enterprise pricing details thinner than more established rivals |
| Full API and SDK access alongside no-code tools | Newer brand, less independent latency benchmark coverage |
Pricing: Self-serve from $0.09/min, dropping to about $0.05/min for enterprise clients; premium voices priced separately; Enterprise custom. See Vogent's own comparison page.
Best for: Teams that want an agent that automatically re-tests itself against real past calls as it evolves
Telephony: Bring Your Own Twilio or Take the Bundled Number
Every platform above solves the same plumbing problem differently: how does a phone call actually reach the agent? Most developer-first platforms support both a bundled number you buy in-platform and a bring-your-own carrier setup through Twilio, Telnyx, or a raw SIP trunk, which matters if you already have numbers, minutes, or compliance history tied to an existing carrier account.

| Platform | Numbers Provided | BYO Twilio / SIP | Notes |
|---|---|---|---|
| Vapi | Buy in-platform or import | Full SIP trunk support, inbound and outbound | Most flexible telephony story of the developer-first tools |
| Retell AI | Local numbers in US, UK, Canada, Australia+ | Twilio, Telnyx, Vonage via Elastic SIP | Three paths: buy, import, or connect a trunk |
| Bland AI | Bundled numbers | BYOT available, waives transfer fees | Flat per-minute rate covers everything but transfers |
| ElevenLabs Agents | Native plus Twilio, Genesys, Amazon Connect | Yes | Multi-channel beyond voice: WhatsApp, SMS, web |
| Cartesia Line | Cartesia-provided numbers at $0.014/min | Via Voximplant and other telephony partners | Newest telephony layer, fewer direct integrations |
| Deepgram Voice Agent API | None native | Fully BYO, it's an API | Telephony is entirely the developer's responsibility |
| Phonely | 1-5 free numbers depending on tier | SIP trunking at Enterprise | Numbers bundled at every self-serve tier |
| Thoughtly | Bundled or BYO carrier | Twilio, Telnyx | No-code builder sitting over a bundled or BYO number |
| PolyAI, Parloa, Regal, Sierra, Synthflow | Enterprise telephony or existing contact center stack | Scoped per contract | Built to sit in front of or behind an existing dialer or IVR |
If your team already runs Twilio, a platform with full SIP trunk support (Vapi, Retell AI, Thoughtly) keeps your numbers and call history in one place. Starting from zero, a bundled number gets you dialing in minutes, which is where Bland, Phonely, and Cartesia's own numbers are the faster path.
The Real Per-Minute Math: AI Voice Agent vs. a Human Agent
Every voice agent vendor's per-minute rate is really competing against one number: what a human costs to do the same job. A fully loaded US onshore call center agent runs about $26 an hour in 2026, roughly $0.43 a minute, versus $6 to $14 an hour offshore ($0.10-0.23/min) and $12 to $18 an hour nearshore ($0.20-0.30/min), per Call Force Global's 2026 cost data.

| Comparison | Per-Minute Cost | What It Includes |
|---|---|---|
| US onshore human agent (fully loaded) | ~$0.43/min | Wages, benefits, software, overhead, recruiting, turnover |
| Nearshore human agent (Caribbean, Latin America) | $0.20-0.30/min | Same as above, lower-cost labor market |
| Offshore human agent (Philippines, India) | $0.10-0.23/min | Same as above, lowest-cost labor market |
| Deepgram Voice Agent API (self-serve) | $0.051-0.075/min | STT, LLM orchestration, TTS; BYO telephony |
| Cartesia Line (self-serve) | $0.06/min + $0.014/min telephony | TTS, STT, orchestration, Cartesia-provided number |
| Vapi / Retell AI (self-serve, all-in) | $0.07-0.31/min | STT, LLM, TTS, varies heavily by model choice |
| Bland AI (flat rate) | $0.11-0.14/min + platform fee | LLM, STT, TTS bundled, no token charges |
| Phonely (self-serve overage) | $0.25-0.35/min | Full stack plus a phone number, small-business tier |
On paper, most self-serve platforms undercut even offshore human labor once volume justifies the setup, and that's the comparison every vendor pitch leads with. It's a fair one for a narrowly scoped job, qualify a lead, confirm an appointment, answer a FAQ, that a script can handle end to end. It's a less honest one for anything needing judgment or de-escalation, which is why Gartner's $80 billion labor-savings forecast for 2026 pairs with a note that only about 1 in 10 agent interactions will be fully automated this year. Price the agent against the specific call type, not your total headcount budget.
Compliance and Consent: What Changes When the Caller Is Recorded
An AI voice agent isn't just a cheaper call. It inherits a specific set of legal obligations the moment it dials out or records what it hears.

Outbound consent. The FCC confirmed in February 2024 that an AI-generated or cloned voice counts as an "artificial voice" under the TCPA, so an outbound call from a voice agent needs the same prior express consent as a prerecorded robocall, plus clear caller identification and a working opt-out. An existing business relationship doesn't exempt an AI voice call the way it can exempt a manual one, and penalties run $500 to $1,500 per call with no aggregate cap.
Call recording and PCI. The moment a caller reads a card number to an agent that records the call, that recording is in scope for PCI DSS, and storing a spoken CVV after authorization is a straightforward violation. Platforms built for regulated industries (Bland, Phonely, Synthflow, Thoughtly, Sierra) offer pause-and-resume recording or a separate secure capture flow to keep card data out of the transcript. Verify this explicitly rather than assuming it's handled by default.
HIPAA. If the call touches protected health information, you need a signed Business Associate Agreement, not just a vendor's security page. Bland AI states HIPAA eligibility with a BAA outright; Vapi, Thoughtly, and Phonely gate it behind their top enterprise tiers; PolyAI, Parloa, and Sierra scope it per contract.
None of this is optional or unique to one vendor. Confirm consent workflows, recording controls, and BAA availability in writing before a pilot goes near real customer calls, not after.
How to Choose: Decision Framework
Start with the operating constraint that matters most: build control, contact-center support, outbound structure, pricing predictability, or voice latency.

| If you need... | Pick... | Why |
|---|---|---|
| Full control over your own STT, LLM, and TTS stack | Vapi | Most flexible telephony and model choice on this list |
| A production agent with itemized, predictable pricing | Retell AI | Transparent per-component billing, strong support reputation |
| One flat rate with zero model-provider surprises | Bland AI | LLM, STT, and TTS bundled into a single per-minute price |
| The widest language coverage and best voice realism | ElevenLabs Agents | 70+ languages, same engine as the market's top TTS product |
| A fully managed enterprise rollout across many locations | PolyAI | Managed conversation design and tuning included in the contract |
| Outcome-based pricing at very high call volume | Parloa | Pay per resolved conversation, built for 2M+ calls a year |
| One agent answering chat and phone the same way | Sierra (Voice) | Unifies an existing chat agent with a new voice channel |
| Structured outbound sales or retention calling | Regal | Purpose-built for outbound campaigns at real volume |
| A managed build with no in-house AI engineering | Synthflow | Fully managed stack, now scoped as an enterprise contract |
| Unlimited minutes for real estate, insurance, or education leads | Thoughtly | Flat $500/mo covers unlimited voice minutes |
| A free-to-start agent for a small business's own phone line | Phonely | Real self-serve signup, free tier, no sales call required |
| The lowest-latency voice model as your foundation | Cartesia Line | Sub-90ms TTS/STT models, purpose-built Line platform |
| The cheapest infrastructure if you're building agent logic yourself | Deepgram Voice Agent API | Lowest published per-minute rate, full BYO LLM support |
| An agent that automatically re-tests itself against past calls | Vogent | Built-in evaluation suite most competitors lack |
What to Do Next
Pick your lane before you pick a vendor: building your own agent (Vapi, Retell AI, Cartesia Line, Deepgram) or buying a managed one (PolyAI, Parloa, Synthflow, Sierra). Then run a real pilot, not a demo script, with a genuine interruption mid-sentence and a caller who mumbles, and time the response yourself instead of trusting a vendor's latency page. Confirm consent workflows and BAA availability in writing before a single real customer call touches the system.
If you'd rather design the agent's workflow yourself than shop for a platform, our AI voice call agent blueprint and how to build an AI agent guide cover the build path, and the choosing an AI agent platform framework applies just as well to voice as it does to chat. For phone-adjacent support workflows, see our AI support triage agent blueprint. And for the wider set of AI agent products, start with our best AI agent platforms guide.

Principal Product Marketing Strategist
On this page
- Updated August 2026: What Changed
- Key Facts
- Quick Comparison Table
- Latency and Barge-In: The Number That Actually Decides If a Call Feels Human
- 1. Vapi: Developer-First Infrastructure for Custom Voice Agents
- 2. Retell AI: Vapi's Closest Rival, Tuned for Reliability
- 3. Bland AI: Flat Per-Minute Pricing With Nothing Bolted On
- 4. ElevenLabs Agents: Voice Quality and 70+ Languages From the TTS Leader
- 5. PolyAI: Enterprise Contact Center Voice AI, Fully Managed
- 6. Parloa: Outcome-Based Pricing for High-Volume Contact Centers
- 7. Sierra (Voice): One Agent Across Chat and Phone
- 8. Regal: Outbound Sales and Retention Calling at Scale
- 9. Synthflow: No-Code Builder Now Behind an Enterprise Wall
- 10. Thoughtly: Unlimited Minutes for Real Estate, Insurance, and Education
- 11. Phonely: The Cheapest Real Self-Serve Entry Point
- 12. Cartesia Line: Built on the Fastest Voice Model in the Category
- 13. Deepgram Voice Agent API: Bring Your Own LLM, Pay by the Hour
- 14. Vogent: Self-Improving Agents, Now Backed by Aircall
- Telephony: Bring Your Own Twilio or Take the Bundled Number
- The Real Per-Minute Math: AI Voice Agent vs. a Human Agent
- Compliance and Consent: What Changes When the Caller Is Recorded
- How to Choose: Decision Framework
- What to Do Next