Product Information
What is Voicebox?
Voicebox is a state-of-the-art speech generation model built on Meta's non-autoregressive flow-matching model. By learning to tackle text-guided speech inpainting tasks through large-scale data, Voicebox surpasses single-purpose AI models in speech tasks via in-context learning. It can synthesize speech across six languages, remove transient noise, edit content, transfer audio styles within and across languages, and produce diverse speech samples. Even more impressively, it generates speech 20 times faster than the leading autoregressive models.
How to use Voicebox?
Voicebox is an advanced voice generation model that surpasses single-purpose AI models in various voice tasks through contextual learning. It synthesizes speech in six languages, removes transient noise, edits content, converts audio styles, and generates diverse voice samples.
Core Functions of Voicebox
Text-to-Speech
Noise Elimination
AI-Driven
Multilingual
Voice Synthesis
Text to speech
Multiple Languages
Usage Scenarios of Voicebox
- Generate multilingual voice content.
- Noise removal and content editing for existing audio.
- Apply the style of specific audio to new voice recordings.
- Quickly create diverse voice samples.
Common Questions about Voicebox
What does Voicebox do?
How do I use Voicebox?
What are the core features of Voicebox?
What are the application scenarios for Voicebox?



















