Product Information
What is Gazelle speech language model?
Gazelle is Tincans' joint speech-language model—for more details and timely insights, check out our V0.2 announcement. This is an early research preview—please temper your expectations! Gazelle can take both text and audio as inputs (interchangeably) and generate text as output. You can further synthesize the text output into audio via a TTS provider (not implemented here). Example tasks include transcribing audio, answering questions, or understanding spoken audio. This approach will be superior for business use cases where latency and conversational quality matter, such as customer support, outbound sales, and more.
Known limitations exist! The model is trained only on English audio and isn’t expected to perform well with other languages. Similarly, it doesn’t yet handle accents well. The Gradio demo might have issues with audio sampling rates. We also only accept single audio inputs (microphone or upload).
Inference is powered by serverless GPUs. As a result, you may experience cold-start delays on first use (around 30 seconds), but subsequent responses will be faster. This demo isn’t optimized for inference speed but rather to showcase Gazelle’s capabilities. We do not store any responses.
How to use Gazelle speech language model?
Gazelle is a federated voice-language model developed by Tincans, capable of processing text and audio inputs to generate text outputs, enhancing interaction quality and efficiency in business scenarios.
Core Functions of Gazelle speech language model
Voice-to-Text
Voice Transcription
Voice Recognition
AI-Driven
Usage Scenarios of Gazelle speech language model
- Transcribe audio
- Answer questions
- Understand spoken audio
- Customer support
- Outbound sales
Common Questions about Gazelle speech language model
What does Gazelle do?
How do I use Gazelle?
What are the core features of Gazelle?
What are the application scenarios of Gazelle?



















