Skip to content

OmniVoice Generates 24 kHz Speech in 600-Plus Languages From 3-Second Voice Samples

Sep 11, 2026
GitHub
Article image for OmniVoice Generates 24 kHz Speech in 600-Plus Languages From 3-Second Voice Samples

Summary

OmniVoice now generates 24 kHz speech across 600-plus languages from 3–10-second voice samples, though cross-lingual clones retain the accent of the reference speaker’s language.

Key Points

  • OmniVoice supports zero-shot speech generation in more than 600 languages, with voice cloning, attribute-based voice design, and automatic voice selection.
  • The model produces 24 kHz audio and reports real-time factors as low as 0.025, while FlashInfer delivers lossless inference acceleration of roughly 2x to 2.9x on NVIDIA GPUs.
  • Voice-cloning references work best at 3–10 seconds; cross-lingual output retains an accent from the reference audio’s language.

Tags

Read Original Article