OmniVoice Generates 24 kHz Speech in 600-Plus Languages From 3-Second Voice Samples
Summary
OmniVoice now generates 24 kHz speech across 600-plus languages from 3–10-second voice samples, though cross-lingual clones retain the accent of the reference speaker’s language.
Key Points
- OmniVoice supports zero-shot speech generation in more than 600 languages, with voice cloning, attribute-based voice design, and automatic voice selection.
- The model produces 24 kHz audio and reports real-time factors as low as 0.025, while FlashInfer delivers lossless inference acceleration of roughly 2x to 2.9x on NVIDIA GPUs.
- Voice-cloning references work best at 3–10 seconds; cross-lingual output retains an accent from the reference audio’s language.