ARTFEED — Contemporary Art Intelligence

NVIDIA Magpie TTS: Open-Weight Multilingual Voice AI with Low Latency

ai-technology · 2026-08-10

NVIDIA has introduced Magpie Multilingual TTS, a text-to-speech model featuring 364M parameters and open weights, capable of supporting 12 languages: English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean, and Brazilian Portuguese. This model offers both male and female voice options, enhances code-switching specifically for Hindi and Japanese, and is designed for real-time speech generation. Key techniques include frame stacking and a local transformer, as outlined in the paper 'Frame-Stacked Local Transformers for Efficient Multi-Codebook Speech Generation' (ICASSP 2026). It boasts a Time to First Audio (TTFA) of 32ms on the B200 GPU and is available on Hugging Face and as an NVIDIA NIM container, improving synthesis quality and lowering character error rates for new languages. Magpie is part of the NVIDIA Nemotron Voice Agent Developer Example.

Key facts

  • NVIDIA Magpie Multilingual TTS is an open-weights model with 364M parameters.
  • Supports 12 languages: English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean, Brazilian Portuguese.
  • New languages added: Modern Standard Arabic, Korean, Brazilian Portuguese.
  • Achieves TTFA of 32ms on B200 for single stream; 239ms at 64 concurrent streams with 320x real-time throughput.
  • Architecture includes frame stacking and local transformer for low latency and quality.
  • Available on Hugging Face under NVIDIA Open Model License and as NVIDIA NIM container.
  • Part of NVIDIA Nemotron Voice Agent Developer Example for building complete voice agents.
  • Improved quality for French and Spanish with reduced CER and higher SSIM.

Entities

Institutions

  • NVIDIA
  • Hugging Face

Sources