Google introduces Gemini 3.5 Transcribe for more accurate voice-to-text

0
64
Google expands voice capabilities with Gemini 3.5 Transcribe Credit: 9to5google
Google expands voice capabilities with Gemini 3.5 Transcribe Credit: 9to5google

Google has introduced Gemini 3.5 Transcribe, a new speech-to-text model designed to deliver more accurate and polished transcriptions. The model is already powering products including Gboard Rambler and the Gemini app for macOS, with support also coming to Chrome.

Unlike conventional speech recognition systems, Gemini 3.5 Transcribe is designed to handle background noise, complex jargon, disfluencies and natural speech patterns. It can turn raw audio into formatted text while recognising custom vocabulary and understanding user intent.

The model can handle self-corrections such as “let’s meet Tuesday—no, Wednesday” and remove filler words such as “ums” and “ahs” from the final text. It also supports automatic formatting and natural voice editing.

According to Google, Artificial Analysis benchmarks show an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use cases. The model is also designed to accurately capture alphanumeric information such as postal codes and order IDs, even in noisy environments.

Gemini 3.5 Transcribe also supports custom vocabulary, allowing it to recognise specialised terminology and unique spellings. It can automatically detect and transcribe more than 85 languages, while handling regional accents and different dialects.

For pre-recorded audio, the model can identify and attribute speech to up to 3 speakers with timestamps, while support for more than 3 speakers remains experimental.

Google says the model delivers significantly lower latency than its Chirp 3 transcription model. Time to final transcription has improved by 70%, while on the FLEURS benchmark, it recorded a 5.50% WER in streaming and 5.04% in non-streaming use cases across selected languages and locales.

Gemini 3.5 Transcribe also supports voice-based task execution through function calling. This allows it to delegate complex tasks, including image generation and file analysis, to other Gemini models.

The technology is already available in Gboard Rambler on Android, the Gemini macOS app, and the microphone within Google Antigravity’s prompt box. It uses screen context and chat history, with user permission, to improve transcription across file names, agent thoughts and active documents.

The model is also coming to Chrome, where users will be able to dictate replies, draft posts and enter prompts into web fields using their voice.

For developers, Gemini 3.5 Transcribe is available in public preview through the Gemini API via Google AI Studio and Google Antigravity. For enterprises, it is available in public preview through the Gemini Enterprise Agent Platform and is coming soon to Gemini Enterprise for Customer Experience.

Also read: Viksit Workforce for a Viksit Bharat

Do Follow: The Mainstream LinkedIn | The Mainstream Facebook | The Mainstream Youtube | The Mainstream Twitter

About us:

The Mainstream is a premier platform delivering the latest updates and informed perspectives across the technology business and cyber landscape. Built on research-driven, thought leadership and original intellectual property, The Mainstream also curates summits & conferences that convene decision makers to explore how technology reshapes industries and leadership. With a growing presence in India and globally across the Middle East, Africa, ASEAN, the USA, the UK and Australia, The Mainstream carries a vision to bring the latest happenings and insights to 8.2 billion people and to place technology at the centre of conversation for leaders navigating the future.