Mistral AI launched Voxtral TTS, a text-to-speech model with 4 billion parameters.
The system generates speech in nine languages including English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic.
Voxtral TTS supports emotional expressiveness and contextual understanding for natural-sounding dialogue.
It adapts to custom voices using only a three-second audio sample as a reference prompt.
The model achieves a latency of 70 milliseconds for typical inputs while maintaining quality.



