In the present day, we’re introducing Gemini 3.5 Transcribe, our most exact speech-to-text mannequin but, designed for clever voice interactions. Not like typical speech recognition fashions that battle with background noise, complicated jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts uncooked audio straight into correct, polished, formatted textual content.
Throughout our merchandise just like the Gemini app and on Android, we’ve seen customers already benefiting from this transcription mannequin with new voice capabilities like Rambler on Android and within the Gemini app on macOS. Now, builders can construct comparable capabilities with Gemini 3.5 Transcribe within the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
We have constructed 3.5 Transcribe to plug seamlessly into your developer workflows, whether or not you’re constructing voice brokers, real-time captioning instruments, or post-call analytics pipelines. The mannequin is obtainable throughout two separate APIs:
- Actual-time streaming: Delivers steady, bidirectional streaming with sub-second latency for interactive voice apps through the Reside API utilizing
gemini-3.5-transcribe-live. - Pre-recorded audio processing: Transcribes recorded audio, conferences, name logs, and extra with speaker attribution and word-level timestamps through the Interactions API utilizing
gemini-3.5-transcribe.
Get extra exact and clever transcription
Gemini 3.5 Transcribe is designed to seize your pure talking fashion to raised perceive your intent and acknowledge customized vocabulary, so you’ll be able to execute duties together with your voice.
- Good transcription: Seamlessly handles self-corrections (like “let’s meet Tuesday—no, Wednesday”), removes filler phrases (“ums” and ‘“ahs”), auto-formats your textual content.
- Perform calling: The mannequin can delegate complicated duties (similar to picture era and file evaluation) to different Gemini fashions through perform calls. At present obtainable within the Gemini macOS app.
- Extra exact transcription: As measured by Synthetic Evaluation, achieves a mean Phrase Error Price (WER) of 4.0% for streaming and a pair of.6% for non-streaming use-cases. It reveals robust efficiency throughout noisy, real-world environments, precisely capturing alphanumeric entities like postal codes and order IDs.
- Customized vocabulary: Acknowledges specialised jargon and distinctive spellings by seamlessly adapting transcriptions to your supplied customized vocabulary.
- World language assist: Mechanically detects and transcribes over 85 languages, seamlessly dealing with regional accents and various dialects.
- Multi-speaker identification: Precisely attributes speech in pre-recorded audio with timestamps for as much as three audio system (assist for 3+ audio system is experimental).







