AI Potluck
Back to Gap Map Model components / Speech & audio

PersonaPlex

NVIDIA
open weights / Overall score: 4.2(strong)

NVIDIA's full-duplex speech-to-speech conversation model, fine-tuned from Kyutai's Moshi: it listens and speaks at the same time, handling interruptions and fast turn-taking, and takes a text role prompt and a voice sample to set its persona. English only.

Openness

3 high confidence
3.0
weights
open(7B safetensors checkpoint on the Hub behind an automatically approved license click-through)
data
closed(Fisher English (an LDC corpus) plus synthetic conversations generated with open LLMs and TTS
code
partial(inference and server code only)
license
NVIDIA-Open-Model-License(the weights

The weights are published under the NVIDIA Open Model License, which permits commercial use and derivatives, and the code is MIT. NVIDIA trained it on the licensed Fisher corpus plus synthetic conversations it has not released, and publishes inference code only.

Adoption

3 high confidence
3.0

Hugging Face downloads of the single released checkpoint.

Capability

5 medium confidence
5.0

Holds a live two-way spoken conversation like Moshi, on which it is built, and adds control over who the assistant is and how it sounds.

Verified 2026-09-27