← Back to research
•·11 min read·company

NVIDIA PersonaPlex

NVIDIA's open-weight full-duplex voice model combines role prompts and voice conditioning. Released-checkpoint benchmarks, live serving, hosted options, licensing, and deployment tradeoffs.

Key takeaways

  • PersonaPlex combines simultaneous listening and speaking with text-based role prompts and audio-based voice conditioning.
  • The released checkpoint reports 170ms smooth-turn latency and 95% interruption takeover rate; the older 70ms and 100% figures describe a different experimental setup.
  • NVIDIA supplies a live browser server and offline evaluation code, while fal offers a separately priced hosted audio API.
  • Commercially usable weights and MIT code provide control, but application tools, scaling, transport, and operational checks remain separate work.

FAQ

What is NVIDIA PersonaPlex?

PersonaPlex is a Moshi-based 7B-class speech-to-speech model that listens and speaks concurrently. Text and voice prompts set its conversational role and vocal style.

Can PersonaPlex be used commercially?

The weights use the NVIDIA Open Model License, which permits commercial use subject to its terms; the code uses MIT. Downloading the gated Hugging Face files requires accepting the model's conditions, and operating the model still costs money.

Does PersonaPlex require self-hosting?

No. Self-hosting is supported, and fal offers a hosted PersonaPlex audio API. That API's documented audio-file input and streaming results should not be assumed to provide a complete bidirectional phone-call service.

Does the reference implementation support live conversations?

Yes. NVIDIA documents a live browser server as well as an offline WAV evaluation path. Community ports have their own capabilities and performance limits.

Executive Summary

PersonaPlex is NVIDIA's open-weight speech-to-speech model for conversations in which listening and speaking overlap. Released in January 2026, it combines a text description of the agent's role with an audio prompt that conditions its voice. Its model card still identifies the public model as v1.0, with English speech input and output.[1][2]

The important September corrections concern deployment and evidence. NVIDIA already supplies a live browser server, and fal now offers a hosted PersonaPlex endpoint. Meanwhile, the paper distinguishes its experimental model from the released checkpoint; their latency and naturalness numbers should not be mixed.[3][4][5]

See agentic voice APIs for the wider landscape. PersonaPlex is a model and reference implementation, so evaluating it requires a different checklist from evaluating a complete contact-center platform.

Model and Architecture

Listening and speaking at the same time

PersonaPlex builds on Kyutai's Moshi architecture and Helium language-model foundation. Mimi encodes incoming audio into tokens and decodes generated tokens back into speech; temporal and depth transformers model the conversation. Text and audio are generated while user audio continues to arrive, at a documented 24kHz audio sample rate.[1]

Role text + voice conditioning
              ↓
User audio → Mimi encoder → conversational model → Mimi decoder → agent audio
                                      ↓
                                  agent text

This model can learn when to pause, speak over another speaker, or provide a short acknowledgment. It does not mean every interruption will be handled correctly, nor does it eliminate microphone, network, buffering, and playback delay. Full duplex describes the interaction architecture; measured responsiveness describes a particular implementation and workload.

Roles and voices

The reference package includes 18 voice presets grouped into natural and varied voices. Its prompting guide distinguishes a question-answering assistant role, customer-service roles supplied with relevant facts, and open-ended conversation.[3]

NVIDIA's examples include banking support, medical-office reception, casual conversation, and fictional emergency roleplay.[1] These demonstrate conditioned speech behavior. They do not establish identity verification, appointment persistence, medical competence, or access to a real account system. Those actions require application logic beyond a convincing spoken response.

Training and the released model

The research overview describes blending real Fisher conversations with synthesized assistant and customer-service dialogues. That combination targets both conversational timing and task adherence.[1] The paper's appendix explains that the released checkpoint adds real conversational data and changes synthetic voice generation relative to the experimental setup.[5]

The practical consequence is easy to miss: a benchmark row labeled PersonaPlex in the main paper is not automatically the score of the downloadable checkpoint. Use the released model's results when planning a deployment.

Benchmarks: Identify the Checkpoint and Metric

The current model card reports these FullDuplexBench results for the released checkpoint:[2]

MetricReleased checkpointInterpretation
Smooth turn-taking latency0.170 secondsTime between the user stopping and the agent starting
Smooth turn-taking takeover rate0.908Benchmark measure of taking the speaking turn
User-interruption latency0.240 secondsTime for the agent to stop after an interruption
User-interruption takeover rate0.950Benchmark interruption behavior, not a guarantee on every call

The main paper's experimental table reports 70ms smooth-turn latency and 1.000 interruption takeover rate. Those are distinct from the released-checkpoint values above. Its released-model naturalness study reports 2.95 ± 0.25 for PersonaPlex and 2.80 ± 0.24 for Gemini; the authors say this separate annotator pool is not directly comparable with the earlier 3.90/3.72 study.[5]

A small difference between mean scores should not become a universal ranking. The confidence intervals overlap, the baselines are those evaluated in the paper, and today's deployed services may differ. These are author-reported benchmark results, not measurements performed for this profile.

For an application, test both conversational behavior and factual adherence. A model that sounds natural while inventing a shipping promise can be less useful than a slower model that follows the actual policy. A benchmark emphasizing turn-taking cannot answer that product question by itself.

Running the Reference Implementation

Live browser use and offline evaluation

NVIDIA's GitHub README documents installation of the Opus development library, installation of the repository's moshi package, acceptance of the Hugging Face model conditions, and a server launched with temporary TLS certificates. The browser UI uses port 8998. An offline path consumes WAV audio and writes response audio and text.[3]

The two paths support different tests. An offline fixture helps compare repeated inputs, seeds, prompts, and model versions. The live UI reveals how microphone pickup, interruption timing, and playback affect an actual conversation. A deployment should pass both kinds of test before a favorable recording is treated as evidence of reliable interaction.

Hardware and concurrency

The model card lists Linux/PyTorch, A100 and H100 compatibility, and an A100 80GB as test hardware.[2] The README additionally documents CPU offload when GPU memory is insufficient, a Blackwell-specific PyTorch installation note, and CPU-only offline evaluation.[3] An offload option establishes a possible execution path, not a promised latency on a smaller GPU.

Reading the reference server shows WebSocket audio and text streaming, shared model state, and a lock around active conversation processing. It is not a ready-made multi-tenant serving system.[6] Production concurrency needs an explicit plan for session isolation, workers, routing, capacity, and cancellation. Increasing the number of connected browsers does not prove the model can sustain that many simultaneous conversations.

The same source logs role prompts and voice-prompt details.[6] If those prompts contain customer information, log handling belongs in the deployment design. The fact that inference runs locally does not automatically settle where application logs, recordings, or backups go.

Hosted Access and Cost

fal offers a hosted fal-ai/personaplex endpoint. Its documented schema accepts an audio-file URL, role prompt, and preset voice or voice-sample URL, with response audio and text. It supports streaming results, but the reviewed schema does not demonstrate continuous bidirectional microphone transport or a complete telephony service.[4]

RoutePublished cost or operating responsibility
Self-hosted modelNo-charge model license subject to terms; hardware, serving, and operations remain your costs[7]
fal hosted endpoint$0.001 per audio second as displayed September 15, 2026[8]
fal custom voice sampleAPI documentation says voice-sample conditioning is billed at twice the rate[4]

At the displayed standard fal rate, 60 billed audio seconds cost $0.06; that is arithmetic, not an estimate of a complete phone call's cost. Confirm the endpoint's billing units and workload behavior before extrapolating.

For self-hosting, this report does not assign a generic GPU hourly price. Instance type, utilization, simultaneous-session capacity, and latency requirements affect cost. A useful measure is infrastructure spend per successful conversation at the target concurrency, including idle capacity and failed sessions.

Licensing and Operational Responsibility

PersonaPlex's code is MIT-licensed.[9] The weights use the NVIDIA Open Model License, which permits commercial use and derivative models subject to its terms and says NVIDIA does not claim ownership of outputs. The agreement also includes conditions concerning guardrails, redistribution, and acceptable use.[7] The weights and code therefore do not share the same license.

The model card requires accepting conditions to access the gated Hugging Face files and identifies additional Moshi attribution information.[2] “Open weights” is a more precise deployment description than assuming the entire training corpus and every associated artifact are available under MIT.

A model license also does not grant permission to impersonate a person whose voice appears in an audio sample. For a practical evaluation, use supplied presets or a voice the application is authorized to use, and keep the voice choice separate from claims about identity or trustworthiness.

Community Ports and Practitioner Experience

Speech Swift provides a community Swift/MLX implementation for Apple Silicon, including PersonaPlex and its 18 presets.[10] Its author's current article describes a roughly 5.3GB four-bit model and streaming output; the reported faster-than-real-time result is tied to an M2 Max with 64GB, not every Mac.[11] This is a separate implementation, not NVIDIA's hardware support commitment.

The March Hacker News discussion contains both enthusiasm and concrete problems. armcat valued the model but emphasized that composable speech pipelines can also feel responsive. d4rkp4ttern reported slow, irrelevant responses on an M1 Max. Tepix questioned an example's unsupported shipping assurance. These are individual reports with different environments, rather than a controlled comparison.[12]

An often-repeated WAV-only complaint in that thread concerned the Apple Silicon demonstration. It should not be used to claim NVIDIA's reference implementation lacks live conversation: the official README and server source explicitly provide it.[3][6] Check which repository and revision a report describes before applying it to the model family.

Nemotron 3 VoiceChat is a separate model

NVIDIA's current VoiceChat card describes a 12B model using a Nemotron Nano V2 9B backbone, a different speech stack, and persona control based on PersonaPlex. It still labels the release early-access evaluation only, under API trial or model-evaluation terms.[13] It is evidence of continued work on persona-controlled speech, not a drop-in PersonaPlex checkpoint upgrade or a general production entitlement.

Choose the layer you need

OptionDecision-relevant distinction
PersonaPlexModel weights and reference code for teams willing to build and operate the surrounding application
LiveKit AgentsFramework and media infrastructure for model pipelines, tools, handoffs, and telephony; supports managed or custom deployment[14]
ElevenLabsAgent platform with workflow configuration, tools, knowledge retrieval, web/mobile/phone deployment, and evaluation features[15]
fal PersonaPlex endpointHosted model access for its documented audio request interface[4]

LiveKit's documented components include turn detection and interruption handling; ElevenLabs also documents configurable conversation flow.[14][15] A cascaded or managed system should not be dismissed as incapable of responsive interaction simply because PersonaPlex models duplex behavior directly.

If the requirement is a working support line with authenticated actions, compare the full application path. If the requirement is studying overlapping speech with control over inference, the model-level approach is more directly relevant. Neither architectural choice establishes the best instruction-following quality without testing the actual task.

A Practical Evaluation Workflow

Use a small fictional service with explicit facts: an item is unavailable, delivery takes three days, and the agent may explain the policy but cannot authorize a refund. Keep the prompt, voice, audio fixtures, and model revision fixed initially.

  1. Check factual adherence. Ask the same question several ways, including a request for an impossible delivery date. Record unsupported promises separately from fluent answers.
  2. Check timing. Include a mid-sentence pause, a brief acknowledgment, and a deliberate interruption. Measure when playback starts and stops, not just server computation time.
  3. Check the application boundary. Ask for a real action. Verify that only a successful authorized tool result can cause the interface to report completion; spoken confidence is insufficient.
  4. Check deployment behavior. Repeat under the target concurrency and network conditions, then interrupt a session or restart a worker. Observe recovery, retained state, logs, and cost.

This is an evaluation design, not a test run performed for this report. It preserves the distinction between a model producing plausible dialogue and a service completing a user's request correctly.

Strengths, Cautions, and Fit

PersonaPlex is a useful choice to investigate when control over weights, role conditioning, and overlapping speech matters. The live reference implementation, offline path, and community ports give developers several ways to inspect those behaviors.

The main cautions are integration and evidence. English-language model documentation does not establish multilingual reliability; experimental benchmark wins do not establish the released model's rank today; and a conversational voice is not a tool-execution system. NVIDIA's research and community momentum are reasons to evaluate the model, not guarantees about support or future releases.

Teams wanting a managed voice application should compare platform-level alternatives and the hosted endpoint's actual transport limits. Teams building their own system should budget for serving, media handling, authorization, evaluation, and human escalation in addition to inference. The right outcome is a conversation that remains accurate and useful under the application's real conditions.


Research by Ry Walker Research • methodology