Key takeaways
- Choose the responsibility you want to buy: a speech model, speech components, a managed agent platform, or an open framework.
- OpenAI Live separates conversation from backend work, Gemini 3.8 Live changes tool behavior, and hosted Cartesia Line agents face a December migration deadline.
- Session minutes, audio tokens, characters, provider fees, and reserved compute are different billing units; no universal all-in price multiplier applies.
- Evaluate verified actions, interruption recovery, and deployment boundaries alongside voice quality and response time.
FAQ
What is the best voice API for an AI agent?
It depends on which layer you need. Compare hosted voice models for conversation, managed platforms for phone workflows, and LiveKit or Pipecat for programmable control; test each against your actual tasks.
Is streaming speech the same as full-duplex conversation?
No. A service may stream output from an uploaded audio file without continuously listening while speaking. Full duplex also does not automatically cancel an already-running business action.
Which options can I self-host?
LiveKit Agents and Pipecat provide open frameworks; PersonaPlex and Ultravox offer model deployment paths under their applicable licenses. Self-hosting an agent framework does not make its external model providers local.
How should I compare voice-agent costs?
Use the same call recordings and tasks, then include model or session usage, telephony, provider charges, idle capacity, and optional services. Track cost per correctly completed task as well as cost per minute.
Executive Summary
This comparison covers 19 voice models, speech APIs, managed agent platforms, and open frameworks. They solve different parts of a spoken application: generating audio, maintaining a conversation, executing business actions, carrying a phone call, and operating the service. Buying one part does not necessarily buy the others.
Three September changes matter immediately. OpenAI's new GPT-Live separates full-duplex conversation from a delegated backend agent. Gemini 3.8 Live is now a stable model with different turn and tool behavior from the earlier native-audio models. Cartesia is moving hosted Line SDK agents to Managed Agents, with a December 1 migration deadline.[1][2][3]
The useful question is which responsibilities your team wants to own. A managed phone platform can supply workflows and call operations. A framework gives more control over providers and deployment. Open weights allow inference control but leave more application and serving work. This report recommends evaluation paths, not a universal winner: no comparative calls or latency tests were run for this refresh.
Market Definition and Inclusion
An agentic voice system connects spoken interaction to reasoning or actions. This comparison includes publicly usable live voice services, configurable voice-agent platforms, speech components used to build them, and frameworks or weights that developers can operate themselves. Funding announcements and recent publicity are not prerequisites for inclusion.
The four useful layers overlap:
| Layer | What it supplies | What still needs attention |
|---|---|---|
| Conversational models and realtime APIs | Audio understanding, response generation, turn behavior, and sometimes delegated reasoning | Business authorization, durable action state, media integration, and recovery |
| Speech components | Streaming transcription or speech synthesis | Reasoning, tool execution, turn management, and the rest of the call |
| Managed agent platforms | Configured agents, provider coordination, tools, and call operations | Business rules, integration correctness, evaluation, and contractual requirements |
| Open frameworks | Programmable media and agent coordination | Deployment, capacity, provider accounts or model serving, and application support |
These are functional distinctions, not exclusive company classifications. Cartesia and Deepgram offer both speech components and managed agents. LiveKit and Pipecat can connect to native speech-to-speech services as well as a transcription–LLM–speech pipeline.[4][5][6][7]
Comparison Matrix
The deployment column distinguishes running application code from running the actual models. Commercial private deployments and downloadable weights are also different from open-source framework licenses.
| Platform | Main role in this comparison | Deployment boundary | Decision-relevant constraint |
|---|---|---|---|
| OpenAI Live / Realtime | Hosted conversational models | Managed voice models; client-controlled backend possible | Live delegation and Realtime are different APIs and execution models[8] |
| Gemini Live | Hosted audiovisual conversation | Managed Gemini API | 3.8 changes tool defaults and removes earlier affective-dialogue configuration[2] |
| AWS Nova 2 Sonic | Unified speech model | Amazon Bedrock bidirectional stream | Verify region and application integration; not a complete phone platform[9] |
| Grok Voice | Hosted speech-to-speech | Managed API | Separate input/output audio metering; session and concurrency limits[10] |
| Hume EVI | Expressive conversation | Managed EVI service | EVI 3 and EVI 4 mini have different model and control requirements[11] |
| Inworld | Managed realtime speech pipeline | Hosted service, labeled research preview | Compatible-looking events do not ensure identical tool continuation behavior[12][13] |
| NVIDIA PersonaPlex | Full-duplex open-weight model | Self-hosted weights and live reference server | Reference serving and application orchestration remain your responsibility[14][15] |
| Ultravox | Audio-understanding model plus hosted voice service | Downloadable model variants or hosted API | The model generates text; spoken output requires TTS[16] |
| ElevenLabs Agents | Managed agent platform and speech stack | Hosted agents | Conversation minutes and LLM usage are separate charges[17] |
| Cartesia | Speech APIs and Managed Agents | Managed agents; self-hosted agent code path | Hosted Python Line SDK migration due December 1, 2026[3] |
| Gradium | Streaming STT/TTS components | Hosted API; contracted private deployments | Regional pinning and retention depend on enrollment and data type[18] |
| Rime | TTS component | Managed API or commercial private deployment | Coda and Mist v3 expose different controls[19][20] |
| Deepgram Aura / Voice Agent | Speech components and hosted agents | Managed speech and agent APIs | Aura pricing is not the price of a complete Voice Agent session[21] |
| Vapi | Managed provider orchestration | Hosted platform; chosen providers and custom backend | Hosting fee excludes provider usage[22][23] |
| Retell AI | Managed phone-agent operations | Hosted platform with carrier options | Model, voice, telephony, and optional QA affect the bill[24] |
| Bland AI | Managed phone agents and workflows | Hosted platform; enterprise infrastructure options | New agent dashboard and existing Pathways need separate configuration review[25][26] |
| Twilio Conversation Relay | Managed speech and phone connection to your backend | Twilio media; your WebSocket application and LLM | Application owns reasoning, actions, and reconnect recovery[27][28] |
| LiveKit Agents | Open agent and media framework | Self-hosted or LiveKit Cloud | Cloud session compute excludes inference and other services[6][29] |
| Pipecat | Open voice pipeline framework | Self-hosted or Pipecat Cloud | Active and warm compute, providers, and transport have separate economics[7][30] |
Cost Reality Check
Prices below were checked September 16, 2026 and are in USD. They describe published billing units, not measured all-in costs. There is no defensible universal “multiply the headline by two or three” rule: speech activity, context growth, carrier geography, model choice, call duration, reserved capacity, and optional services change the result.
| Product or layer | Published rate or unit | What the figure leaves out |
|---|---|---|
| OpenAI GPT-Live-1 | $0.05 per session minute, metered per second | Backend models and tools[31] |
| OpenAI GPT-Realtime-2.1 | $32 input / $64 output per million audio tokens; $0.40 cached audio input | Text/image usage; actual token mix and retained context determine the bill[32] |
| Gemini 3.8 Live | $3 input / $12 output per million audio tokens | Text, video/image input, transcription, and repeated active context[33][34] |
| Grok Voice | $0.08 per input-audio minute and $0.08 per output-audio minute, plus billable text items | Telephone carrier and application costs; these are two audio directions, not one call-minute price[10] |
| ElevenLabs Agents | Monthly plans with included minutes; published additional call minutes at $0.08 | LLM and applicable telephony charges[35] |
| Cartesia Managed Agents | $0.06/minute; Cartesia phone usage adds $0.014/minute | LLM passthrough after the published promotion ends October 1[36][3] |
| Gradium | TTS credits per character; STT credits per audio second | Your LLM, orchestration, transport, and account's credit price[37] |
| Rime | Mist $0.03 / Coda $0.05 per 1,000 characters | STT, LLM, media, and agent operations[20] |
| Deepgram | Aura-2 $0.030 per 1,000 characters PAYG; standard Voice Agent $0.075/minute PAYG | Alternative plans/configurations, telephony, and application services[21] |
| Vapi | $0.05/minute hosting plus provider passthrough | Success packages, extra concurrency, carrier, and optional services[23] |
| Retell | $0.055/minute infrastructure plus voice, LLM, and telephony | Optional QA and other add-ons[24] |
| Twilio Conversation Relay | $0.07 per active session minute | Twilio Voice and your LLM/backend[38] |
| Ultravox hosted | $0.05 per minute, billed in six-second increments | Telephony and applicable additional services[39][40] |
| LiveKit Cloud / Pipecat Cloud | Selected agent compute rates start at $0.01 per active minute | Inference, applicable transport, warm capacity, and other services[29][30] |
Hume's public pricing labels its EVI table EVI 3; do not silently apply those rates to every EVI version. Inworld meters component consumption through credits. Bland separates standard plans from its Agent Phone offering. PersonaPlex self-hosting requires a hardware and utilization calculation rather than a fictional zero-dollar call rate.[41][42][43][44]
Worked example: a configured phone agent
Consider an illustrative 1,000 billable minutes on Retell with $0.055 infrastructure, $0.015 standard voice, and $0.016 standard-tier GPT-4.1 mini per minute. The sum is $86 before telephony, optional services, and fixed charges. Adding a listed $0.015/minute US telephony route produces $101. This is arithmetic for one configuration, not a quote or a measured production average. The pricing page says connected silence remains billable and AI-agent fees stop after transfer while telephony continues.[24]
For GPT-Live, 1,000 session minutes similarly produce a $50 voice-model subtotal. The cost of the delegated task can vary substantially even when the spoken conversation lasts the same amount of time. Keep the voice session and backend usage as separate lines in your ledger.[31]
Context, silence, and idle capacity
Google documents that Live turns rebill the retained session context, including earlier raw audio. Its displayed audio-minute equivalents therefore do not define a flat call-minute price. Context compression changes which history remains; optional transcripts add text output. Proactive listening also keeps ingesting audio when the model chooses not to reply.[34]
Cloud framework pricing introduces another distinction: a warm worker can cost money before a call arrives. Pipecat Cloud prices reserved compute separately from active sessions, and its scaling guide says each bot instance handles one session. A larger machine tier is not automatically more concurrent calls per instance.[30][45]
Calculate cost per completed, correct task as well as cost per minute. Retries, failed transfers, and long clarification loops can make a nominally cheaper model more expensive for the workload. That is an evaluation method, not a claim that one provider has a higher failure rate.
Product Analysis
OpenAI: choose Live or Realtime deliberately
OpenAI now offers two distinct voice architectures. GPT-Live handles the spoken interaction while a separate backend performs delegated work. Responses delegation supplies a managed path; client delegation lets the application run its own model, agent harness, or service. Realtime instead combines speech, reasoning, and tool selection in one model. The Live guide documents browser WebRTC, server WebSockets, and phone integration paths.[8]
GPT-Realtime-2.1 accepts audio, text, and images and exposes configurable reasoning effort; increasing effort can also increase latency and output usage. Its model card identifies the Realtime endpoint, not the Live endpoint.[32] Prefer an explicit architecture decision over substituting one model identifier into the other's example. Neither product's availability establishes a universal accuracy lead on your business tasks.
Gemini: a migration with behavioral consequences
The current Gemini API documentation marks Gemini 3.8 Live stable. It supports audio/video/image input and audio output with transcription. Compared with earlier models, tool calls default to nonblocking, proactive audio is always on, and affective-dialogue configuration has been removed. The standard model does not accept a thinking_level setting.[2]
The separate Extended Thinking model adds another lifecycle distinction: a completed spoken turn need not mean the underlying interaction is finished. Its guide requires checking interaction_status and allows only nonblocking tool execution. Test old assumptions about “turn complete,” interruption, and when a task is safe to declare finished before migrating an existing agent.[46] This is a stronger decision criterion than assuming every product named Gemini Live has identical behavior or contractual terms.
AWS Nova 2 Sonic: Bedrock-native speech
AWS exposes Nova 2 Sonic as amazon.nova-2-sonic-v1:0 through Bedrock's InvokeModelWithBidirectionalStream. The model card lists four supported regions: Northern Virginia, Oregon, Stockholm, and Tokyo.[9] Its service card describes unified speech understanding and generation with asynchronous tools: the conversation can continue while an external operation finishes, but the application still implements those operations and business logic.[47]
That makes it a natural evaluation candidate for a team already operating its agent backend in Bedrock. It does not justify the old fixed “$0.015 per call minute” estimate. Measure actual speech and text usage for the chosen region and workflow before comparing costs with a service billed by session duration.
Grok Voice: inspect both audio directions
Grok's current speech-to-speech documentation supplies a hosted voice model with explicit session and concurrency limits. The material pricing distinction is that incoming and generated audio are separate meters, alongside billable text items.[10] Its SIP guide also makes the carrier boundary visible: phone numbers and carrier routing remain part of the application's setup.[48]
Evaluate it when a direct realtime model and a programmable phone connection fit the application. Check how a transfer failure returns control and how long a session may run; those details matter more than treating a single displayed audio rate as the price of a complete support line.
Hume EVI: version-specific expressive interaction
Hume belongs in the comparison because EVI is a usable conversational API, not because it had a new funding announcement. EVI 3 and EVI 4 mini differ: the latter requires a supplemental language model and does not support the same quick-response setting. Pin the version when comparing behavior or implementing turn controls.[11]
Its tool-use documentation includes tool failure and correction flows, which are useful for action-oriented conversations.[49] Expressive speech can help an interaction feel appropriate, but perceived emotion is not evidence of a person's internal state. Test whether the agent completes the task and handles mistakes without making psychological claims about the caller.
Inworld: modular realtime service, preview caveats
Inworld's Realtime API combines speech recognition, a configurable language-model route, and speech synthesis behind a conversational interface. Its product FAQ still labels availability a research preview, so that status should accompany production planning.[50][12]
Two integration details deserve tests. The migration guide says the service can continue automatically after receiving a tool result; blindly copying another API's explicit response-creation sequence can produce duplicate responses. Its memory documentation also describes expiration after 15 minutes of inactivity, so a session-memory feature should not replace durable business records.[13][51] Compatibility should be established with event traces and failure cases, not inferred from similar JSON names.
PersonaPlex: inspect the shipped checkpoint and serving code
NVIDIA's downloadable PersonaPlex checkpoint is an English full-duplex model. The current model card reports 170ms smooth-turn latency and 240ms interruption latency on its benchmark. The often-repeated 70ms result comes from a different experimental setup in the paper; it is not the released checkpoint's number.[44][52] These are author-reported measurements, not comparable end-to-end measurements against today's hosted services.
NVIDIA supplies a live browser server as well as an offline path. The reference server's shared state and conversation lock still require a separate concurrency design.[14][15] Code is MIT; weights use NVIDIA's separate Open Model License.[53][54] fal's hosted audio-file request interface is another access route, but streaming a result from a file is not evidence of a continuous bidirectional phone session.[55]
Ultravox: distinguish the model from the service
Ultravox's model consumes speech directly and generates text; a TTS stage supplies audible responses. The hosted service adds the surrounding conversation facilities, while downloadable model variants offer a different deployment path.[16] The repository identifies variant-specific language-model backbones, so its MIT code license should not be substituted for every weight license.[56]
For application design, its staged calls and background threads are worth evaluating when a conversation needs specialist behavior or concurrent work.[57][58] Self-hosting a model variant does not automatically reproduce all hosted service features or remove the need for transport, TTS, tools, and operational monitoring.
ElevenLabs: evaluate the agent platform as well as the voice
ElevenAgents combines transcription, a selected LLM, speech synthesis, and a turn model. The platform includes workflows, retrieval, tools, deployment channels, and evaluation features; it should be compared with managed agent products as well as with standalone TTS APIs.[59]
The published billing explanation charges connection duration, applies a large discount to qualifying extended silence, and bills the LLM separately. That is different from paying only for the characters the agent speaks.[17] Voice quality is a workload-specific evaluation: include your names, numbers, accents, and interruptions rather than treating a general listening preference as a guarantee of task success.
Cartesia: speech stack plus a hosting migration
Cartesia's Managed Agents documentation now describes Ink 2, a selected LLM, and Sonic 3.6, with configuration through the UI or API.[4] Existing custom Python agents hosted through the Line SDK must migrate by December 1. The same notice says self-hosted agent code continues working, so the deadline should not be described as the entire SDK disappearing.[3]
Check feature parity before moving a production workflow. The migration page still lists knowledge bases, multiple languages per agent, and call-event webhooks as forthcoming. Its temporary free-LLM promotion is a dated offer, not a durable cost assumption.[3] This is a concrete operational change that matters more to an existing customer than a new TTS leaderboard position.
Gradium: components and explicit data boundaries
Gradium's API supplies streaming transcription and synthesis in its documented language set. Credits map to text characters for TTS and audio duration for STT; they do not include a general reasoning agent.[37] It belongs on the shortlist for teams composing their own speech stack.
Its September data-residency announcement distinguishes ordinary regional endpoints from enrolled, pinned residency. Before enrollment, routing is best-effort; enrolled routing fails closed instead of silently crossing the boundary. Zero-retention commitments for inference payloads also do not automatically cover stored voice-cloning or voice-design assets. Private deployment options require their own agreement.[18] These distinctions are more useful than a blanket “EU endpoint means every byte stays in Europe” claim.
Rime: check controls model by model
Rime's current models are Coda and Mist v3, not simply the older Arcana label carried forward. The model guide recommends Coda generally but documents different capabilities: inline phoneme controls remain on older Mist models, while pauses and timestamps vary across current models. A voice migration can therefore change more than sound quality.[19]
Rime offers character-priced hosted TTS and commercial private deployment options.[20] It is a component choice for a phone or app pipeline; add speech recognition, reasoning, turn handling, and transport to the evaluation. Verify the specific model's language and pronunciation behavior instead of using an aggregate marketing language count as a compatibility promise.
Deepgram: two different purchasing decisions
Deepgram offers speech components, including Aura-2 and the newer Flux TTS listing, alongside its managed Voice Agent API. The latter coordinates listening, reasoning, and speaking, with configurable models and tools.[21][5]
That creates two distinct comparisons: replacing a TTS or transcription component in your existing application, or adopting an integrated agent session. PAYG, Growth, and bring-your-own-provider configurations have different rates. Do not combine the lowest rate from one configuration with the included features of another.[21] For an existing Deepgram customer, a controlled comparison can hold the speech component constant and vary who operates the agent loop.
Vapi: provider choice with managed coordination
Vapi's configuration separates transcriber, model, and voice, and can connect a custom language-model backend.[22] Squads coordinate specialist assistants with handoffs and shared context, which is useful when distinct call stages need different prompts or tools.[60]
The economic tradeoff is explicit: its hosting charge sits beside provider passthrough and plan-dependent operating features. Current packages also change included concurrency and retention.[23] Evaluate provider substitution and a failed specialist handoff as part of the pilot. A modular configuration gives flexibility, but the application still needs to decide which assistant owns the next action.
Retell: workflow and call operations
Retell documents both conversation flows and prompt-based agents, with function calls, warm transfers, DTMF/IVR behavior, versioning, testing, and live-call monitoring.[61] These features make it a useful candidate for teams seeking a managed operating surface for phone workflows.
Its component pricing makes a configuration reproducible, provided the comparison includes the selected LLM, voice, carrier route, and optional QA. Continuous QA has its own per-minute charge.[24] A practical pilot should include transfer-to-human failures and post-call evaluation, not only a successful inbound greeting. Those checks help determine which operational features actually reduce work for the team.
Bland: test the workflow, including promotion gates
Bland's new dashboard guide and existing Pathways documentation describe two relevant configuration surfaces. Existing pathway logic should be inspected before assuming a newly created agent is a drop-in replacement.[25][26] Its scenario tools support assertions and promotion gates, letting a team check required conversational outcomes before advancing a configuration.[62]
Bland belongs alongside managed phone-agent platforms. Compare the actual standard-plan and carrier terms for the intended deployment; the separate Agent Phone product is not a universal enterprise quote.[43] Test assertions about completed actions as well as assertions about spoken phrases, because saying a booking succeeded is weaker evidence than a confirmed booking record.
Twilio: keep your backend, delegate the speech connection
Conversation Relay manages the speech and call connection while an application exchanges text and events over WebSocket and runs its chosen reasoning backend.[27] It is particularly relevant when a team already owns business logic or telephony routing and wants to avoid assembling raw audio processing.
The recovery boundary is concrete. The protocol does not automatically reconnect a broken application WebSocket, and a handoff event is not proof a human answered. The application must use the resulting call-control path and implement recovery deliberately.[28] This makes Relay a different purchase from a fully configured agent platform, even though both may appear as a phone number to the end user.
LiveKit Agents: framework and media control
LiveKit Agents supports Python and Node applications, pipeline or native speech models, tools, and agent handoffs. Agents participate in rooms, with WebRTC handling frontend media and provider connections handled behind the agent. Teams can deploy through LiveKit Cloud or their own infrastructure.[6] The framework is Apache 2.0 licensed.[63]
This is a good evaluation path when a team needs control over the media application and agent lifecycle. Cloud plans reduce operational work, but the agent-session price is one part of a bill that can also include inference, telephony, and observability.[29] Self-hosting changes who operates that infrastructure; it does not make external model calls local.
Pipecat: programmable pipelines and capacity planning
Pipecat is a BSD-2-Clause Python framework for conversational applications, with multiple transports, provider integrations, and multi-agent coordination. It is not restricted to Daily transport or to a transcription–LLM–TTS chain.[7][64]
Pipecat Cloud supplies a managed deployment path, while self-hosting remains an option. Its scaling documentation describes warm pools, maximum instances, and overload behavior; the default maximum needs to be distinguished from broad marketing claims about scale.[45] For a pilot, test cold starts and a burst above configured capacity. Those observations tell you whether to reserve workers, raise limits, or add explicit call admission behavior.
Architecture Patterns That Affect Correctness
Native speech, cascades, and delegated backends
A conventional cascade passes audio through transcription, a language model, and speech synthesis. This makes provider replacement and inspection of intermediate text straightforward, but the application must coordinate partial output and interruption across components. Native speech models handle more of the conversation within one model. A delegated design can separate conversational responsiveness from longer reasoning or tool work, as GPT-Live documents.[8]
These choices are not a universal latency ranking. A useful measurement starts at the end of the user's actual utterance and ends when the user hears a meaningful answer. Separately measure interruption stop time and time to a verified tool result. A fast first audio chunk may contain silence or filler, and a fluent response may arrive before an action completes.
Gradium's September account of a Coval benchmark change illustrates the problem: adding leading silence to the measured time changed what the metric represented. The vendor explicitly says pre-change and post-change numbers are not directly comparable.[65] Do not place TTS first-byte numbers, model turn-taking scores, and full telephone round-trip measurements in the same “latency” column.
Full duplex does not settle action state
Full duplex means the system can listen while speaking. It is already available in a hosted product such as GPT-Live and in an open-weight model such as PersonaPlex; describing it as a feature that may only arrive in 2028 is outdated.[1][44] Whether an application hears an interruption, stops playback, cancels inference, and cancels a business operation are four different questions.
OpenAI's delegation guide explicitly keeps permissions, private tool execution, and durable task state in the application. An interruption does not automatically cancel backend work.[66] The same design problem should be tested in any voice application rather than assumed solved by an “interruptible” label.
Worked workflow: reschedule an appointment
Use a fictional booking service with explicit permissions and a reliable test database. This is a proposed evaluation, not a test performed for this report.
Caller requests a new time
→ backend verifies identity and reads availability
→ agent presents the available option
→ caller confirms
→ backend commits one booking change and returns its identifier
→ agent reports the confirmed result
Now interrupt while the change is running: “Wait, keep the old time.” The test should establish whether the operation was never started, canceled, completed, or compensated. A second call or retried tool request must not create a duplicate appointment. Keep a durable operation identifier so the next agent or human can resolve the state.
Repeat with a slow database, a failed tool, a disconnected socket, and an unavailable transfer destination. Record the actual booking state beside the transcript. This exposes the difference between conversational naturalness and reliable task execution without assuming a particular vendor will pass or fail.
Practitioner Evidence and Its Limits
A Gemini developer-forum thread provides a useful billing caution. In April 2026, DamianFilipek described unexpectedly high live-audio costs; later discussion addressed rebilled context and compression. Another participant, AlisaFortin, described updated guidance, while the original developer reported different results for text tests and actual audio. These are dated reports about the earlier model and test setup, not proof of a Gemini 3.8 defect.[67] They justify replaying real audio and inspecting usage events instead of validating cost with text-only tests.
For most of the newest releases, public primary documentation is stronger evidence of supported interfaces than public discussion is of comparative reliability. This review does not convert individual complaints, vendor demos, or leaderboard scores into representative accuracy rankings. The linked profiles contain additional implementation reports and limitations where substantive evidence was available.
Deployment, Privacy, and Tool Boundaries
Start with a data-flow inventory: microphone or carrier, speech provider, language model, tool server, transcript store, recordings, logs, and analytics. A regional model endpoint does not locate every downstream system. Gradium's enrollment and voice-asset distinctions are one documented example of why the scope of a promise matters.[18]
Similarly, hosting a framework yourself gives control over that process; it does not relocate a managed inference API. Downloadable weights provide a different degree of control, while the applicable model license still needs review. PersonaPlex's code/weight split and Ultravox's backbone-specific variants make those distinctions concrete.[53][54][56]
A function call or MCP connection is an integration surface, not authorization by itself. For example, Retell documents MCP and external functions, while Bland provides tool connections and scenario checks.[61][62] Keep the permission check where the business action is performed, and evaluate how the spoken response changes when that action is rejected.
Where Tembo Fits
For a voice interface that hands off software work, the backend execution platform becomes a separate decision. Tembo runs coding agents in cloud environments with repository and integration context, produces reviewable pull requests or artifacts, and supports team visibility and self-hosted deployment. Its current positioning addresses that execution and review lifecycle.[68] Disclosure: I am Tembo's founder and CEO.
That makes Tembo adjacent to this market rather than a twentieth voice API. The reviewed product material does not establish a realtime speech model, media transport, or a built-in integration with the voice services above. A team exploring a spoken software-work assistant would still need a custom bridge to authenticate the user, submit the authorized task, track its identifier, and communicate progress or a review link. Do not equate a spoken acknowledgment with a completed code change.
GPT-Live's client-delegation design demonstrates why the distinction matters: the application owns the backend and durable task state.[66] A short conversation can hand off a longer task, while approval and review remain in the system that actually performs the work. That is an architectural evaluation path, not a claim of an existing Tembo voice connector.
Recommendations by Requirement
These are shortlists based on documented fit. Test the actual workload before choosing a winner.
| Requirement | Start by evaluating | What should decide the trial |
|---|---|---|
| Conversational voice with long-running backend tasks | OpenAI Live; Inworld or Ultravox for their different execution models | Task lifecycle, interruption, result validation, and recovery |
| Audio/video context in a Google application | Gemini Live | Current model behavior, context accounting, and tool completion |
| Bedrock-centered deployment | Nova 2 Sonic | Region, authentication, backend integration, and measured usage |
| Direct hosted speech-to-speech | OpenAI Realtime, Gemini Live, Grok Voice, Hume EVI | Names/numbers, accents, version-specific controls, and action accuracy |
| Managed business phone workflows | Retell, Vapi, Bland, ElevenLabs Agents, Cartesia Managed Agents | Transfer failures, workflow changes, QA, and full configured cost |
| Existing Twilio application and owned reasoning backend | Conversation Relay | WebSocket lifecycle, call-control fallback, and chosen speech settings |
| Replacing a speech component | Gradium, Rime, Cartesia, Deepgram | Your audio, pronunciation controls, language coverage, and integration effort |
| Programmable media and provider control | LiveKit Agents, Pipecat | Transport needs, deployment skills, capacity, and lifecycle customization |
| Control over model inference | PersonaPlex; a suitable Ultravox model variant | Weight license, hardware, actual concurrency, and missing application features |
| Regional or private processing requirements | Contracted deployment paths and the full data flow | Written scope for inference, assets, logs, subprocessors, and failure routing |
For a small team, begin with a narrowly scoped managed workflow and a test set that includes failure paths. For a team with media or infrastructure expertise, an open framework can make provider and deployment choices easier to control. For model research, downloadable weights expose behavior that a managed platform may not. These recommendations reflect ownership tradeoffs, not evidence that one team size requires a particular vendor.
Exclusions and What to Watch
Hume, Bland, Grok, Inworld, Twilio Conversation Relay, and Ultravox are now included because they meet the functional criteria. A lack of recent funding news is not grounds to exclude an available developer product.
Cekura is adjacent testing and monitoring infrastructure, rather than one of the runtime voice choices counted here.[69] Sesame remains outside this developer-API comparison: its reviewed public site describes its consumer voice work, and this review did not verify a generally available developer API from that material.[70] That is a bounded availability finding, not a claim that no private program or future API exists.
The developments worth watching are concrete: whether preview services change their production terms, whether migration guides reach feature parity, how context and idle capacity appear in real bills, and whether evaluation tools catch incorrect actions as reliably as awkward speech. The September releases already invalidate a simple forecast that full duplex is years away. They do not establish that one architecture will replace every other one.
Research by Ry Walker Research. Sources checked September 16, 2026. See the research methodology.
Sources
- [1] OpenAI — GPT-Live-1 API launch (September 10, 2026)
- [2] Google — Gemini 3.8 Live stable model (September 15, 2026)
- [3] Cartesia — Line SDK migration and December 1 deadline
- [4] Cartesia — Managed Agents architecture
- [5] Deepgram — Voice Agent architecture and configuration
- [6] LiveKit — Agent framework architecture and deployment
- [7] Pipecat — Framework, transports and provider integrations
- [8] OpenAI — GPT-Live architecture and connection guide
- [9] AWS — Nova 2 Sonic model card and regions
- [10] xAI — Speech-to-speech model, pricing and limits
- [11] Hume — EVI version differences
- [12] Inworld — Voice Agents availability FAQ
- [13] Inworld — Migration and automatic tool continuation
- [14] NVIDIA — PersonaPlex live and offline reference implementation
- [15] NVIDIA — PersonaPlex WebSocket server source
- [16] Ultravox — Audio-to-text model and speech synthesis
- [17] ElevenLabs — Conversation duration, silence and LLM billing
- [18] Gradium — Residency, retention and private deployment (September 2, 2026)
- [19] Rime — Coda and Mist model capabilities
- [20] Rime — Hosted and private-deployment pricing
- [21] Deepgram — Speech and Voice Agent pricing
- [22] Vapi — Provider and custom-backend configuration
- [23] Vapi — Hosting, passthrough and operating packages
- [24] Retell — Component, carrier and QA pricing
- [25] Bland — First agent in the new dashboard
- [26] Bland — Conversational Pathways
- [27] Twilio — Conversation Relay architecture
- [28] Twilio — WebSocket events, handoff and recovery
- [29] LiveKit — Cloud agent and inference pricing
- [30] Daily — Pipecat Cloud active and reserved compute pricing
- [31] OpenAI — GPT-Live-1 model and pricing
- [32] OpenAI — GPT-Realtime-2.1 model card
- [33] Google — Gemini API pricing
- [34] Google — Live API best practices and context billing
- [35] ElevenLabs — Agents plans and additional minutes
- [36] Cartesia — Agent and speech billing units
- [37] Gradium — Speech API capabilities and credit units
- [38] Twilio — Conversation Relay pricing
- [39] Ultravox — Hosted pricing
- [40] Ultravox — Billing FAQ
- [41] Hume — Published EVI 3 pricing table
- [42] Inworld — Component billing and limits
- [43] Bland — Plans, usage units and extra charges
- [44] NVIDIA — Released PersonaPlex checkpoint model card
- [45] Pipecat Cloud — Session instances, warm pools and limits
- [46] Google — Gemini 3.8 Live Extended Thinking
- [47] AWS — Nova 2 Sonic service card
- [48] xAI — SIP setup and transfer recovery
- [49] Hume — EVI tool failure and correction flows
- [50] Inworld — Realtime architecture
- [51] Inworld — Memory lifecycle
- [52] PersonaPlex paper — Experimental and released-checkpoint results
- [53] NVIDIA — PersonaPlex MIT code license
- [54] NVIDIA — Open Model License (October 24, 2025)
- [55] fal — PersonaPlex audio request schema
- [56] Fixie AI — Ultravox repository and model variants
- [57] Ultravox — Call stages
- [58] Ultravox — Background threads
- [59] ElevenLabs — ElevenAgents platform overview
- [60] Vapi — Specialist assistant handoffs
- [61] Retell — Agent building and call operations
- [62] Bland — Scenario assertions and promotion gates
- [63] LiveKit — Apache 2.0 framework license
- [64] Pipecat — BSD-2-Clause license
- [65] Gradium — Perceived first-audio benchmark methodology (September 9, 2026)
- [66] OpenAI — Live delegation and application responsibilities
- [67] Google developer forum — Live speech pricing discussion (April–June 2026)
- [68] Tembo — Agent execution, review and deployment platform
- [69] Cekura — Voice and chat agent testing and monitoring
- [70] Sesame — Public product site