Product Information
What is Voxtral?
The Voxtral models are state-of-the-art speech understanding models, available in two sizes—a 24B variant for production-scale applications and a 3B variant for local and edge deployments. Both versions are released under the Apache 2.0 license. We also offer these two models via API, along with a highly optimized transcription-only endpoint that delivers unmatched cost efficiency. Voxtral Small is an enhanced version of Mistral Small 3, incorporating cutting-edge audio input capabilities while maintaining top-tier text performance. It excels in speech transcription, translation, and audio comprehension. Voxtral mini is an upgrade to the sector's 3B model, integrating advanced audio input features while preserving first-class text performance. It performs exceptionally well in speech transcription, translation, and audio understanding.
How to use Voxtral?
Voxtral offers advanced voice understanding models with transcription, translation, and audio comprehension capabilities, along with cost-effective transcription services.
Core Functions of Voxtral
Ad-free, AI-driven, voice transcription, speech recognition
Usage Scenarios of Voxtral
- Production-Scale Applications
- On-Premises and Edge Deployments
- Voice Transcription
- Voice Translation
- Audio Understanding
Common Questions about Voxtral
What does Voxtral do?
How do I use Voxtral?
What are the core features of Voxtral?
What are the application scenarios for Voxtral?



















