Google Launches Gemini 3.5 Transcribe: Its Most Precise Speech-to-Text Model Yet
Google launched Gemini 3.5 Transcribe on August 26, 2026, introducing its most precise speech-to-text model designed to convert raw audio into polished, formatted text. This new model handles self-corrections, strips filler words, recognizes custom vocabulary, and automatically detects over 85 languages, aiming to enhance intelligent voice interactions. For broader context, explore our AI Tools by Platform.
Enhanced Accuracy and Speed
Gemini 3.5 Transcribe demonstrates notable improvements in accuracy and processing speed compared to previous models. For streaming audio, it achieves a word error rate of 4.0 percent, while for recorded audio, this rate drops to 2.6 percent. These figures indicate a focus on precision across different audio input scenarios.
Furthermore, the model offers a 70 percent reduction in time-to-final-transcription when compared to its predecessor, Chirp 3. This speed enhancement is designed to support high-volume, latency-sensitive tasks, making it suitable for real-time applications and rapid processing of recorded content.
Key Features and Capabilities
Gemini 3.5 Transcribe is designed with several advanced features to deliver high-quality transcriptions:
- Intelligent Text Formatting: It processes raw audio into polished text, automatically handling self-corrections and removing filler words.
- Custom Vocabulary Recognition: The model can recognize and adapt to custom vocabulary, which is crucial for specialized fields and industry-specific terminology.
- Multilingual Support: It automatically detects and transcribes in over 85 languages, broadening its applicability for global users.
- Function Calling Integration: A significant capability is its integration with function calling, allowing it to delegate tasks to other Gemini models. This effectively transforms dictation into an agent command channel, enabling more complex interactions and automation through voice commands.
Availability and Integrations
Google is making Gemini 3.5 Transcribe accessible through multiple channels and integrations:
- Developer APIs: Developers can access the model via two dedicated APIs: the Live API (
gemini-3.5-transcribe-live) for streaming audio and the Interactions API (gemini-3.5-transcribe) for recorded audio. These developer tools facilitate integration into custom applications. - Product Integrations: The model is integrated into Gboard's "Rambler" voice input feature on Android devices and is also available within the Gemini macOS app. Google has also announced that talk-to-type functionality powered by Gemini 3.5 Transcribe will soon be available in Chrome web fields, expanding its reach to web-based interactions.
- Platforms: Gemini 3.5 Transcribe is currently available in Google AI Studio and the Gemini Enterprise Agent Platform, providing immediate access for developers and enterprise users.
Performance Benchmarks
The FLEURS benchmark for Gemini 3.5 Transcribe is reported at 5.50% / 5.04%. These metrics provide an objective measure of the model's performance in speech recognition tasks.
Conclusion
The launch of Gemini 3.5 Transcribe represents Google's continued investment in advanced speech-to-text technology. With its focus on precision, speed, and intelligent features like function calling, the model aims to provide a robust solution for converting audio into high-quality text across a wide range of applications and user interfaces. Its availability through developer APIs and integration into key Google products positions it as a significant update for both developers and end-users.
Sources
Recommended AI tools
Freepik AI Image Generator
Image Generation
Generate on-brand AI images from text, sketches, or photos—fast, realistic, and ready for commercial use.
Google Antigravity
Productivity & Collaboration
Google Antigravity - Build the new way
Notebook LLM
Productivity & Collaboration
Turn complexity into clarity with your AI-powered research and thinking partner
Google Cloud Vertex AI
Data Analytics
Gemini, Vertex AI, and AI infrastructure—everything you need to build and scale enterprise AI on Google Cloud.
Google AI Studio
Productivity & Collaboration
The fastest way to build AI-first applications with Google Gemini.
Grammarly
Writing & Translation
Your AI writing partner for work
Was this article helpful?
Found outdated info or have suggestions? Send us a note.