Guide de l'acheteur

Best AI Phone Agents in 2026

In-depth 2026 buyer guide to AI phone agents. Inbound vs outbound playbooks, Retell, Vapi, Bland, Synthflow, Goodcall, and Smith.ai—with transfer tests, cost stacks, and TCPA-safe rollout.

Best AI Phone Agents in 2026 — buyer guide visual

TL;DR

Pick the call lane first, then the control surface. Inbound coverage (answer, book, route) and outbound automation (dial, qualify, suppress) share voice tech but fail in different ways. The agent worth buying is the one that finishes a narrow call type, hands off cleanly when confidence drops, and writes a CRM record a human will keep.

If you need…Commencez iciDo not start here
After-hours answer + bookingGoodcall, Smith.ai, or AI receptionist softwarePrimary revenue line on day one
Configurable inbound with tools/CRMRaconter l'IA, SynthflowUnowned knowledge that changes weekly
Full STT/TTS/LLM + SIP controlVapiShipping without transfer QA
High-volume outbound follow-upIA fade or Retell with counselCold lists without stored consent
AI plus live human backupSmith.aiTreating hybrid coverage as set-and-forget

How to use this page: decide whether the job is mostly inbound or mostly outbound, shortlist two vendors in the same category, run the pressure tests on your own numbers, then pilot after-hours or a small consented batch. Adjacent guides: AI receptionist software, Agents IA SDR, sales AI agents, customer support AI agents, AI voice agents, best AI agent tools by category.

How this page was built: product packaging shapes and public pricing pages were reviewed as of June 2026. We did not run paid multi-vendor load tests on every platform in this shortlist. Where numbers appear, treat them as planning ranges and re-check vendor pages before you model finance. Legal notes cite federal sources and are not legal advice.

AI phone agents category map for buyers
Inbound routing and outbound follow-up are different products—map the lane before vendors
Platform comparison: watch transfer and interruption handling, not the happy path Watch on YouTube

What this page covers and what it leaves out

This guide is for phone agents: systems that place or receive telephone calls, hold a natural conversation, call tools (calendar, CRM, ticketing), and hand off to humans.

It is not:

If the job is never miss a call and book it with human backup, open the receptionist guide. If the job is programmable voice workflows, SIP, and tool graphs, stay here.


What an AI phone agent actually is

An AI phone agent is a stack, not a single model:

  1. Telephony — numbers, SIP/BYOC, concurrency, caller ID reputation
  2. Realtime loop — STT → policy/LLM → tools → TTS, with barge-in
  3. Outils — calendar, CRM, tickets, knowledge, status lookups
  4. Control plane — prompts, versions, evals, transcript QA
  5. Transfert — warm/cold transfer, queues, fail-soft recovery

Sales demos sell the middle of the loop. Production pain lives at the edges: transfers, hold traps, CRM field quality, recording consent, and outbound suppression.


Inbound vs outbound

Mixing an inbound coverage RFP with an outbound dialer RFP is how teams buy the wrong platform. Treat them as two products that share voice tech.

Inbound

Caller starts. Your number rings (or SIP delivers) and the agent answers.

Jobs that fit: after-hours capture, overflow when humans are busy, appointment book/reschedule, simple FAQs, lead intake for one service type, routing with a summary.

Primary risks: latency, wrong bookings, failed transfers, knowledge rot—not cold-call TCPA exposure.

Go deeper: AI receptionist software for front-desk coverage · customer support AI agents for post-call deflection · omnichannel AI support platforms when phone is one channel among many.

Inbound AI phone agent routing desk
Inbound agents win on answer rate, booking accuracy, and clean human handoff

Metrics that matter: answer rate, booking/task completion, transfer success, time-to-first-word, summary edit rate, repeat-myself complaints.

Pilot pattern: test number → after-hours → overflow → primary line only with proof.

Sortant

Your system starts. The agent dials and speaks first.

Jobs that can fit with governance: no-show recovery, appointment reminders with reschedule, renewal nudges, consented lead follow-up, post-purchase check-ins. Collections only under counsel-approved policy.

Primary risks: TCPA and artificial-voice rules, DNC and suppression, spam-labeled caller ID, brand damage from aggressive or invented claims.

Go deeper: Agents IA SDR for multi-channel sequences · sales AI agents et AI for sales for CRM-native motions · AI workflow automation agents for post-call tasks.

Outbound AI phone agent compliance desk
Outbound only after consent storage, suppression, and brand-safe scripts are real

Metrics that matter: connect rate, right-party contact, opt-out rate, complaint rate, number reputation, cost per successful outcome.

Pilot pattern: consented lists only → low volume → counsel review of transcripts → scale. If consent cannot be explained in plain English, outbound stays blocked.

Side-by-side

DimensionInboundSortant
Who startsCallerYour system
First winAfter-hours / overflowConsented follow-ups / no-shows
Hardest QATransfer + booking accuracyConsent + brand claims
Legal heatRecording + data handlingTCPA, DNC, artificial voice
Adjacent guideAI receptionistAgents IA SDR
Kill criteriaTransfer fails, wrong booksComplaints, spam labels, bad consent

Decision rule

This page ranks call completion under failure:

  1. Transfer quality first
  2. Latency that feels present
  3. Structured outcomes over pretty transcripts
  4. Full stack cost, including failed handoffs
  5. Compliance designed into the agent graph, not attached as a PDF

Comparison matrix

Same axes for every finalist. Symbols are editorial judgments from public packaging and category fit as of June 2026—not lab scores. Re-validate in your own environment.

ProductBest lanePublic starting price (USD, checked July 17, 2026)Pricing shapeSkip if…
Raconter l'IAInbound + careful outbound$0.07–$0.31/minUsage varies with selected model, voice, telephony, and add-onsYou need live humans on every hard call
VapiCustom inbound or outbound infra$0.05/min platform feeSTT, TTS, LLM, and telephony are additional at-cost layersYou have no eng/ops for evals
IA fadeOutbound-heavy volume$0.14/min Start$0.12/min + $299/mo Build; $0.11/min + $499/mo ScaleYou cannot model min charges and spam risk
SynthflowAgency / client inboundEnterprise from $30,000/yearSales-led contract around volume, security, and launch scopeMulti-brand enterprise queues are the core job
GoodcallSMB inbound intake$79/month/agent100 unique customers included; $0.50 above allowanceYou need deep SIP and custom tool chains
Smith.aiInbound coverage + hybridFree AI tier for 25 callsPaid usage is per call; live reception starts at $300/moYou only want raw programmable voice

Product fit notes

Each note: who it is for, what proof to demand, what disqualifies it, verified public pricing as of July 17, 2026, and a narrow first pilot.

Raconter l'IA

For: teams that want a production builder for inbound booking/routing and carefully governed outbound without assembling every provider from zero.

Retell walkthrough: skip the intro—watch calendar tools and live transfer behavior Watch on YouTube

Points forts : approachable agent configuration, function/tool patterns, warm-transfer style flows, knowledge sync patterns many ops teams can run.

Limits: production QA is still yours. Outbound still needs your compliance program. Low-code is not no-ops.

Demand this proof: one live warm transfer with a one-sentence brief to a human, then open the transcript and CRM write from that same call.

Disqualify if: they cannot show fail-soft behavior when the human line does not answer.

Pricing (checked July 17, 2026): Retell lists pay-as-you-go AI Voice Agents at $0.07–$0.31/minute and enterprise pricing by quote. The range depends on the selected model, voice, telephony, and add-ons; it is not one all-in voice rate. Plan on connected minutes × AHT × transfer factor, not the lowest headline number.

First pilot: one inbound reschedule + capture flow. Review the first 30 real transcripts before expanding.

Vapi

For: engineering teams that need maximum control—BYO STT/TTS/LLM, SIP/BYOC, deep internal tools—especially when voice is product infrastructure.

Vapi setup: watch provider wiring and test-call latency, not only dashboard polish Watch on YouTube

Points forts : flexible architecture, explicit provider control for security reviews, solid base for custom IVR replacement and specialized outbound if you build governance.

Limits: ownership cost is real. Without evals and transcript review, flexibility becomes brittle production. Effective cost is platform + providers + eng time.

Demand this proof: p50/p95 time-to-first-word on a mobile call, plus a written path for STT/TTS/LLM provider outage.

Disqualify if: the team cannot staff weekly transcript review and prompt change control.

Pricing (checked July 17, 2026): Vapi lists $0.05/minute for platform hosting. STT, TTS, LLM, and telephony are at-cost provider layers (or $0 to Vapi when you bring the relevant key), so $0.05 is not the finished call cost. Ten concurrent call lines are included; extra lines are listed at $10/line/month.

First pilot: rebuild one existing IVR outcome with tools and a hard transfer. Measure transfer success and eng hours spent on edge cases.

IA fade

For: operators with high-volume voice workflows—especially outbound-heavy programs—who already think in campaign volume and detailed runtime control.

Bland live test: watch interruption, escalation, and tool behavior under stress Watch on YouTube

Points forts : scale-oriented product design, detailed logging/control narratives in market use, useful when dial volume is the point.

Limits: pricing stacks (tiers, per-minute, transfer legs, short-call minimums). Outbound brand risk is on you. Finance must model your AHT and transfer rate.

Demand this proof: a full bill-of-materials quote for 10,000 connected minutes at your AHT with a stated transfer rate, including short-call minimums.

Disqualify if: opt-out does not write a searchable suppression record within minutes.

Pricing (checked July 17, 2026): Bland lists Start at $0.14/connected minute with no platform fee, Build at $299/month + $0.12/min, and Scale at $499/month + $0.11/min. Transfers on Bland-provided numbers are also metered; carrier path and short outbound-call minimums still belong in the model.

First pilot: consented follow-up or no-show recovery only—not cold lists. Review unit economics after roughly 500 connected minutes.

Synthflow

For: agencies and operators shipping client inbound agents quickly with visual builders and packaged telephony.

Synthflow walkthrough: watch knowledge updates and transfer configuration Watch on YouTube

Points forts : speed-to-live for booking/qualification templates; agency-friendly packaging; concurrent-call oriented positioning.

Limits: complex enterprise queue logic and multi-brand policy may need Retell/Vapi-class control. Always test the client transfer graph, not a demo clinic script.

Demand this proof: change a knowledge fact mid-session and show the next call reflects it without a full redeploy mystery.

Disqualify if: agency margin models ignore QA and revision time.

Pricing (checked July 17, 2026): Synthflow says enterprise contracts start at $30,000/year. The final package is scoped around volume, concurrency, telephony, integrations, security, and launch support—so it is not comparable to a self-serve per-minute plan without a proposal.

First pilot: one vertical template end-to-end for a single client. Score setup time, transfer reliability, and knowledge-update pain.

Goodcall

For: SMB service businesses that need inbound intake/booking without hiring a voice engineer. Close cousin to receptionist use cases—compare AI receptionist software if hybrid humans matter more.

Points forts : configurable workflows aimed at real service businesses; lower operational load than raw developer platforms.

Limits: multi-entity enterprises, heavy regulated workflows, and deep custom telephony may need a platform or managed hybrid.

Demand this proof: after-hours booking accuracy across a week of real calls, plus human-request rate.

Disqualify if: you already know you need multi-brand SIP and custom tool chains in quarter one.

Pricing (checked July 17, 2026): Goodcall lists Starter at $79/month per agent with 100 unique customers, Growth at $129 with 250, and Scale at $249 with 500. It says minutes and tokens are unlimited; extra unique customers cost $0.50 each. That makes it more predictable for repeat callers than per-minute services, but overage needs a real caller-volume forecast.

First pilot: after-hours inbound only for two weeks. Track booking rate, callback accuracy, human-request rate.

Smith.ai

For: teams that want inbound coverage with live human backup—when a missed call costs more than pure automation savings.

Smith.ai handoff: watch the human takeover path, not only AI greeting quality Watch on YouTube

Points forts : hybrid model reduces empty-line failure modes; useful bridge while you learn which call types automate safely. Deeper receptionist comparison: AI receptionist software.

Limits: hybrid is not free. Programmable multi-system voice still needs a platform under or beside it.

Demand this proof: abandoned-after-transfer rate and complaint tickets on overflow weeks—not only AI containment.

Disqualify if: leadership only wants raw programmable voice and will not pay for coverage.

Pricing (checked July 17, 2026): Smith.ai advertises an AI Receptionist tier with 25 calls free, then bills paid usage per call. Its live Virtual Receptionist service starts at $300/month for 30 calls. Compare it by cost per booked or qualified outcome, not against a raw platform minute rate.

First pilot: overflow + after-hours with a written escalation matrix.

YourGPT as a control layer

For: teams that run phone next to chat/web and need governed knowledge, structured post-call actions, and approval gates. Pair with a voice runtime (Retell, Vapi, and peers). See AI workflow automation agents for post-call orchestration patterns.

Points forts : structured summary schemas, approved-answer control, multi-channel policy consistency.

Limits: not a substitute for telephony reliability or STT/TTS selection.

Demand this proof: a schema-validated post-call object (intent, fields, next step, owner) that blocks bad CRM writes.

Disqualify if: you expect one product to replace SIP, STT, TTS, and the contact-center fabric.

Pricing posture: confirm current packaging with YourGPT; model as control cost on top of voice minutes.


Three call types to design before you buy

Abstract scorecards fail when nobody has written the actual call. Draft these three before vendor calls.

1) Dental or home-services reschedule (inbound)

Success: correct appointment moved, confirmation sent, CRM updated, no double-book. Kill criteria: wrong slot booked, transfer fails when caller is upset, summary missing date/time. Tools: calendar write, SMS/email confirm optional, human warm transfer. Who owns knowledge: front-desk lead updates hours and blackout dates weekly.

2) SaaS billing overflow (inbound)

Success: identity gate, balance/status lookup or clean ticket, escalate billing disputes to humans fast. Kill criteria: inventing balances, collecting card numbers on open voice, trapping callers in FAQ loops. Tools: CRM/billing read, ticket create, cold transfer to billing queue. Compliance note: payment capture belongs in a controlled path, not free conversation.

3) Consented no-show recovery (outbound)

Success: right party, offer to reschedule, log outcome, honor stop requests immediately. Kill criteria: dialing without consent evidence, spam labels, scripts inventing discounts. Tools: dialer, calendar, suppression list write. Who signs off: counsel on script + ops on suppression latency.

If a vendor cannot run your version of one of these live, stop the evaluation.


Inbound playbook

Call types that usually work first

  1. After-hours message + callback scheduling
  2. Appointment reschedule / cancel with calendar write
  3. Hours and location questions with hard knowledge bounds
  4. New lead intake for one service SKU
  5. Overflow only when queue wait exceeds a threshold

Call types to delay

  • Medical or legal advice, emergencies, crisis language
  • Complex pricing negotiation
  • Multi-policy insurance changes
  • Any flow that needs identity proof you have not designed

Architecture checklist

  • Business hours and holiday calendar as data, not buried only in a prompt
  • Warm transfer brief template (who, why, urgency)
  • Hold failure path: apology + callback capture + ticket
  • Recording announcement that matches jurisdiction
  • Structured CRM fields
  • Named owner for knowledge changes

When the same customer continues in chat or email after the call, plan handoff into customer support AI agents.


Outbound playbook

Programs that can be safe enough to pilot

  • Reminder + reschedule for existing appointments
  • No-show recovery inside a defined window
  • Renewal outreach to customers with a documented relationship
  • Lead follow-up only where consent and purpose are recorded

Programs that need counsel before automation

  • Cold prospecting to purchased lists
  • Artificial voice at scale without TCPA analysis
  • Collections and debt language
  • Health or financial product claims

Architecture checklist

  • Consent store with timestamp, source, purpose, retention
  • Suppression updates within minutes of opt-out
  • Instant stop path that logs and ends the call
  • Brand-safe claim library
  • Number reputation monitoring
  • Human takeover that does not trap the callee
  • Cap on daily attempts per number

If phone is one step in a multi-channel sequence, evaluate Agents IA SDR so you do not force an SDR problem into a pure phone platform.


Buyer scorecard

Score each row 1–5. Fifty points possible. Below 35 usually means you are not ready for a primary line or scaled outbound.

CriterionWhat strong looks likeScore
Transfer reliabilityWarm and cold paths work; failure returns options
Turn-taking / latencyFeels present on mobile
Tool accuracyBookings and CRM match the call
Knowledge boundsStays inside approved policy
Summary qualityStructured fields humans keep
ObservabilitySearchable transcripts and replay
Compliance controlsRecording, opt-out, retention, access logs
Telephony fitNumbers, SIP, concurrency, reputation
Ops ownershipFlows update without a fire drill
Unit economicsClear cost per successful outcome

What to pressure-test before you trust a phone agent

Vendor walkthroughs are optimized for clean audio and cooperative callers. Your business is not.

Run the checks below on your numbers, your calendar or CRM, and a normal mobile phone. If the vendor will not run them live—or cannot show the transcript and event log from the same call—you are buying a black box.

Conversation quality under real speech

Interruption handling. Talk over the agent mid-sentence. A usable agent stops, listens, and continues without restarting the whole script.

Ambiguous openings. Start with something vague about an account, order, or appointment. Strong systems ask one clarifying question at a time.

Background noise. Run a call from a car or open office. The question is graceful degradation, not perfection.

Getting a human when the caller needs one

Multiple ways to ask for a person. Agent, human, representative, operator, or I need someone—route all of them.

Warm transfer with context. The human should hear a one-sentence brief: who is calling, what they need, what already failed.

Immediate bridge under urgency. Billing fights, clinical anxiety, safety language, and furious customers need a fast cold transfer. Measure time from request to ring.

What happens when nobody answers. Let the human line ring long enough to fail. The agent should return with an apology and a concrete next step—not dead air.

Data the business depends on

Callback number confirmation. Capture a number and read it back. Wrong digits create pure waste.

What lands after the call. Open the CRM or ticket. You want outcome, next step, owner, and tags—not a paragraph nobody will read.

Boundaries that protect brand and ledger

Out-of-scope questions. Pricing exceptions, legal advice, clinical interpretation: refuse and escalate.

Sensitive data refusal. Card numbers and Social Security numbers should not be collected on an open conversational path.

Opt-out and do-not-contact. On outbound or marketing-adjacent flows, stop, log, and suppress in a way ops can audit.

How to score the session

Shared sheet: pass, partial, fail, one-line note. Weight transfer and post-call artifacts higher than voice charm. Use the same sheet for every finalist.


How phone agent pricing actually works

Headline per-minute rates almost never equal your bill. Voice stacks charge in layers, and transfers can bill more than one leg.

ComponentTypical unitWhat drives spend
TelephonyPer minute / per numberDuration, transfer legs, countries
Voice platformPer minute or monthly planPackaging, concurrency, features
Speech-to-text / text-to-speechPer minute or per characterLanguage, voice tier, length
Language modelTokensPrompt size, tools, summary length
Workflow writesPer actionCRM enrichment, tickets, SMS

Cost per successful outcome = (platform + telephony + speech + model + tools + human time on failed handoffs) ÷ successful outcomes

Define success narrowly: booked appointment, complete qualified lead, resolved FAQ without transfer, or a logged opt-out that actually suppressed the number.

Published price points buyers can use

These are vendor-published USD prices checked on July 17, 2026. They are planning inputs, not a promise of your invoice: taxes, carrier traffic, selected models, add-ons, contracts, and call patterns can change the result.

ProviderStarting public priceWhat the starting price coversDo not forget
Raconter l'IA$0.07–$0.31/minVoice-agent runtime at a configuration-dependent rateModel, voice, telephony, optional add-ons, and capacity
Vapi$0.05/minVapi hostingSTT, TTS, LLM, transport/telephony, and extra concurrency
IA fade$0.14/min on free StartLLM, STT, and TTS in Bland’s stated connected-minute ratePlan fee on Build/Scale, transfer time, carrier choice, short outbound calls
Synthflow$30,000/year enterprise starting pointThe contracted enterprise platform scopeVolume, integrations, telephony, implementation, and support terms
Goodcall$79/month/agentUnlimited minutes/tokens and 100 unique monthly customers$0.50 per unique customer above allowance and added agents
Smith.aiFree for 25 AI-handled callsEntry AI Receptionist coveragePaid per-call usage, human escalation, or live coverage plans

For a quick budget sanity check, do not compare $0.05/min to $79/month as if they buy the same thing. First estimate connected minutes, unique callers, expected human handoffs, and the share of calls that need a calendar or CRM action. Then model the vendor in the unit it actually sells.

Always re-quote with your AHT, concurrent peaks, transfer rate, and retention needs. Official starting points: Vapi pricing, Retell AI pricing, Bland AI pricing, Synthflow pricing, Goodcall pricing, et Smith.ai pricing.


Compliance for AI phone agents in the US

This is a buyer and operator briefing, not legal advice. Laws and enforcement posture change. Involve counsel before outbound automation, artificial-voice programs, healthcare, or financial use cases. State rules can be stricter than federal baselines.

Why compliance belongs in product design

Phone agents fail when consent is missing, opt-outs are slow, recordings lack a lawful basis, or numbers get labeled spam. Those failures create regulatory exposure and brand damage. Build compliance into prompts, tools, logging, and human paths.

TCPA and artificial or prerecorded voice

The core federal reference many teams start from is 47 CFR § 64.1200 (FCC delivery restrictions).

What operators should internalize:

  • Artificial or prerecorded voice is not a free marketing channel. Depending on line type, purpose, and content, rules can require prior express consent or prior express written consent for telemarketing and certain advertising uses. Exemptions exist but are narrow and fact-specific—read the current text with counsel.
  • Opt-out must work in practice. FCC rules address how consent can be revoked, including reasonable methods and time limits for honoring requests. If your agent cannot stop, log, and suppress quickly, you do not have a compliant program.
  • Identification matters for artificial or prerecorded messages.
  • Do-not-call obligations for telephone solicitations include honoring the national registry under the conditions in § 64.1200, plus company-specific lists and process requirements described in the regulation.

Outbound checklist:

  1. Map every campaign to purpose: informational, transactional, telemarketing, or healthcare-related.
  2. Store consent with timestamp, source, purpose, and authorized number.
  3. Implement in-call opt-out that writes durable suppression.
  4. Control retention and access on recordings and transcripts.
  5. Re-check packaging when script, audience, or dialer logic changes.

Telemarketing Sales Rule and DNC expectations

If calls are telemarketing, the FTC Telemarketing Sales Rule guidance sits beside the TCPA/FCC layer. Ask who classifies campaigns, how entity-specific DNC lists are maintained, and what abandonment rules apply if you blend AI with progressive dialing.

Recording rules vary by jurisdiction (one-party vs all-party regimes). Many teams default to a clear announcement at the start and re-announce when a human joins, because multi-party recordings can change the analysis. Build:

  • An announcement the agent cannot skip on recorded lines
  • A path for callers who refuse recording
  • Strict access control and retention windows

Use jurisdiction-specific counsel; do not rely on a single national assumption.

Caller ID authentication and reputation

Outbound programs die when carriers label numbers as spam. Ask about STIR/SHAKEN support, reputation monitoring, and auto-pause on complaint spikes. Separate AI outbound numbers from human sales lines. Warm volume instead of blast dialing.

Sensitive data, healthcare, and payments

Default: do not collect payment cards, government IDs, or detailed clinical information on an open voice path unless you designed a controlled workflow and matching contracts.

  • Payments: card collection can trigger PCI obligations—prefer secure links or human-assisted capture.
  • Soins de santé : HIPAA applies for covered entities and business associates handling PHI. BAAs, minimum necessary access, and audit logs are procurement items. Definitions: 45 CFR § 160.103.
  • General PII: transcripts often contain addresses and account numbers—redaction and role-based access matter.

Inbound is not risk-free

Inbound avoids many cold-outreach TCPA issues, but you still own recording consent, retention, accurate disclosures, and safe escalation. Emergency or clinical language needs human routing—never model-generated medical advice.

What to demand from vendors in writing

DemanderPourquoi c'est important
Data flow diagram for audio, transcripts, toolsYou cannot secure what you cannot see
Retention defaults and deletion SLAsTranscripts are long-lived liability
Subprocessors for STT, TTS, LLM, storageRisk travels with the chain
Opt-out and suppression exportsAuditors and counsel will ask for proof
Regional hosting and access controlsCross-border audio is board-level
Incident response for wrong disclosureVoice mistakes scale fast

Ship the smallest automated call type that creates value, with logging and human escape hatches on day one.


Metrics that prove the pilot worked

MétriqueInbound signalOutbound signal
Connect / answerCoverage realityDialer health
Containment by intentAutomation valueRarely the primary goal
Transfer successTrustQualité de l'escalade
Task completionBookings, captureRight-party outcomes
Summary edit rateCRM trustCRM trust
Opt-out / complaintsBrand riskBrand + legal risk
Cost / successRetour sur investissementRetour sur investissement

High containment plus high summary edit rate means you automated noise.


Rollout playbook

Week zero: scope

Write one call type, an escalation matrix, success metrics, and kill criteria. Be explicit whether the workstream is primarily inbound or primarily outbound so legal and ops review the right risks.

Days 1–7: shadow and QA

Test numbers only. Build a failure taxonomy. Review nearly every transcript at first.

Week 2: limited live

Inbound: after-hours. Outbound: a tiny consented batch.

Weeks 3–4: expand carefully

Inbound: overflow. Outbound: modest volume with reputation checks.

Later: primary line or scale

Only with proof. Keep a visible human path. Weekly samples. Change control on knowledge and scripts.

Kill criteria examples: transfer success below bar, booking error spike, complaint spike, spam labels, unexplained cost spike, counsel red flag.


Architecture choices that matter more than brand names

ChoicePrefer when…
Managed hybridCoverage reliability is the product (receptionist guide)
Self-serve platformCustom tools, multi-system writes, productized voice
Vendor telephonySpeed to first pilot
SIP / BYOCEnterprise control and reputation tooling
Single skill graphEarly pilots
Split skillsPrompts are long and error rates climb

Knowledge governance: versioned FAQs, do-not-answer lists, owner per change, no silent production prompt edits.


Where phone agents fail in production

  1. Transfer theater that becomes silence
  2. Policy hallucination
  3. CRM pollution
  4. Latency spiral on slow tool calls
  5. Outbound reputation death
  6. Knowledge rot on hours and pricing
  7. Unowned QA after week one
  8. Inbound and outbound requirements mashed into one RFP

Design against these before the pilot, not after the postmortem.


Questions worth asking every vendor

  1. Show a failed warm transfer and the recovery path.
  2. What is p50/p95 time-to-first-word on mobile?
  3. How are barge-in and endpointing configured?
  4. Which CRM fields can write, with what approvals?
  5. How do you search, redact, and retain transcripts?
  6. How fast do opt-outs hit suppression?
  7. Full bill of materials for 10,000 connected minutes at our AHT and 25% transfer rate?
  8. How are provider outages handled in the SLA?
  9. Can we export prompts, flows, and logs if we leave?
  10. What is out of scope by default?
  11. For outbound: how do you support consent audit evidence?
  12. For inbound: how do holiday hours update without a full redeploy?

FAQ

Are AI phone agents the same as IVR?

No. IVR is menu routing. A phone agent should clarify intent, complete tasks, and escalate with context.

Inbound or outbound first?

Almost always inbound after-hours or overflow first. Outbound only after consent, suppression, and brand-safe scripts are operational—and often as a separate workstream from Agents IA SDR.

What is the number one failure mode?

Transfers. If callers cannot reach a human quickly, trust collapses.

Retell vs Vapi vs Bland?

Vapi when you need max provider and SIP control. Retell when ops needs a production builder many teams can run. Bland when high-volume outbound-style control is central and full cost math is validated. Re-run the pressure tests on your call types either way.

How is this different from an AI receptionist?

Receptionist products optimize for front-desk coverage. Phone platforms optimize for programmable voice workflows. Start with AI receptionist software when hybrid humans and coverage matter more than custom tool graphs.

When should we use an AI SDR product instead?

When the job is multi-channel prospecting and sequence design—not only dialing. See Agents IA SDR et sales AI agents.

What metrics unlock the main line?

Stable transfer success, acceptable task and error rates, low complaints, trusted CRM summaries, and unit economics finance will sign—not merely that callers stayed on the line for a few seconds.

What to do next

  1. Write the one call type you actually care about—use one of the three patterns above if you need a template.
  2. Decide whether this workstream is primarily inbound or primarily outbound so legal and ops review the right risks.
  3. Shortlist two vendors in the same category (do not bake off hybrid coverage against raw developer infrastructure without adjusting criteria).
  4. Run the pressure tests and fill the scorecard on the same calls.
  5. Pilot after-hours (inbound) or a small consented batch (outbound) with heavy transcript review.
  6. Expand only when transfer success and cost per successful outcome clear your kill criteria.

Related lanes when the job spills outside pure phone: AI receptionist, Agents IA SDR, customer support AI agents, AI workflow automation agents, best AI agent tools by category.

Need a broader procurement frame? Use the AI agent buying checklist alongside the phone-specific pressure tests on this page.