NVIDIA Releases Magpie: Open-Weights Multilingual TTS Model Targeting Low-Latency Voice Agents
Image via huggingface.co
NVIDIA has released Magpie TTS Multilingual, a 364-million-parameter open-weights text-to-speech model supporting 12 languages — including newly added Modern Standard Arabic, Korean, and Brazilian Portuguese. Designed for cascaded voice AI pipelines, the model is deployable on customer-owned infrastructure via NVIDIA NIM and achieves 32–79ms time-to-first-audio latency, making it viable for real-time conversational applications such as customer support agents, healthcare assistants, and enterprise copilots.
The release is aimed at teams who need fine-grained control over each layer of a voice pipeline — including data residency compliance, custom pronunciation dictionaries, and per-component latency tuning — rather than relying on a single managed speech API. Expanded code-switching support for Hindi and Japanese, backed by IPA grapheme-to-phoneme processing, allows more accurate handling of mixed-language content and technical terminology across a single shared model.