
Gemini 3.5 Transcribe Cuts Word Error Rate to 2.6%
Google's Gemini 3.5 Transcribe replaces Chirp 3 with a 2.6% word error rate, automatic detection across 85+ languages and 70% faster final transcripts.
Speech recognition has spent two decades being *almost* good enough, and the gap has always shown up in the same place: the transcript you have to clean up before anyone else can read it. On August 26, 2026, Google DeepMind introduced Gemini 3.5 Transcribe, a dedicated speech-to-text model that succeeds Chirp 3 across Google's products and developer APIs. The interesting part is not the raw accuracy number, though that improved. It is that the model now does the editing pass for you.
- Gemini 3.5 Transcribe reports a 2.6% average word error rate on pre-recorded audio and 4.0% on streaming audio
- The model automatically detects and transcribes more than 85 languages, including regional accents and dialects
- Time to a final transcript is roughly 70% faster than Chirp 3, the model it replaces
- It is rolling out through Gboard Rambler on Android, the Gemini macOS app, Google Antigravity, and the Gemini API in public preview
What Makes Gemini 3.5 Transcribe Different From Chirp 3
Traditional speech-to-text systems are trained to produce a faithful record of the acoustic signal. That sounds like the right goal until you read the output. Real speech is full of filler words, false starts, and mid-sentence corrections, and a faithful transcript preserves every one of them. The cleanup work then falls to a human or to a second model.
Gemini 3.5 Transcribe folds that step into the model itself. Google says it removes disfluencies, handles self-corrections so the intended phrasing survives rather than the abandoned one, and applies formatting to unstructured speech. It also recognises custom vocabulary, which is the difference between a usable medical or engineering transcript and one that mangles every term that matters.
On the benchmark side, Google reports a 2.6% average word error rate for pre-recorded audio and 4.0% for streaming, measured across conditions that include background noise and conversational AI interactions. For context, the same comparison places Chirp 3 at 5.04% and 5.50% on FLEURS. Roughly halving the error rate is a meaningful step, though as we noted in our look at how speech recognition benchmarks are audited, aggregate WER figures always deserve a closer read than the headline invites.
How Fast Is Gemini 3.5 Transcribe in Practice?
Latency is the quieter half of this release. Google claims time to final transcription is about 70% faster than Chirp 3, and it exposes that through two distinct developer surfaces rather than one general-purpose endpoint.
The Live API handles continuous bidirectional streaming with sub-second latency, which is the profile you want for dictation, live captions, and voice agents that need to respond while someone is still talking. The Interactions API targets recorded material such as meetings, call logs, and interviews, and adds speaker attribution and timestamps. Google says it identifies up to three speakers reliably, with support for more currently experimental.
Splitting the model this way is a sensible acknowledgement that the two workloads have genuinely different requirements. A live captioning system that pauses to think has failed; a meeting transcript that takes four extra seconds and gets the speaker labels right has not.
Where the Model Shows Up First
The rollout is unusually broad for a launch-day announcement. Gemini 3.5 Transcribe already powers Gboard Rambler on Android, Google's long-form voice input mode, and it is live in the Gemini macOS app and in Google Antigravity. Google says Chrome will get it for voice typing in any web field, which would put high-quality dictation into every text box on the web without a per-site integration.
Developers can reach it in public preview through the Gemini API via Google AI Studio and Antigravity, with an enterprise preview arriving through the Gemini Enterprise Agent Platform. That is the same distribution pattern Google has used for recent model drops, including Gemini 3.7 Flash and its developer benchmarks.
Why 85+ Languages and Auto-Detection Matter
The accessibility case here is stronger than the productivity one. Automatic language detection across 85+ languages means a user does not have to declare, in advance, which language they are about to speak — a requirement that quietly excludes multilingual households, code-switching speakers, and anyone whose accent sits outside a model's training centre of gravity.
Pair that with disfluency handling and the practical result is that people who stammer, who think out loud, or who speak a regional variety of a widely-spoken language get a transcript closer to what they meant. Speech interfaces have been steadily getting cheaper and better all year, from Wispr's Canto speech model to open Arabic ASR, and the competitive pressure is showing up as capability rather than just pricing. More of our coverage on this front lives on the artificial intelligence page.
The honest caveat: every number above is Google's own, measured on Google's own evaluation mix. The rollout is broad enough that independent comparisons should arrive quickly, and those are the ones worth waiting for before retiring an existing pipeline.
Sources: Google DeepMind — August 26, 2026; 9to5Google — August 26, 2026; Android Authority — August 27, 2026.
More AI Stories

Meta MTIA 400 Puts 9.4TB/s HBM3e Behind FP4 Inference
Meta detailed MTIA 400 at Hot Chips 2026: eight HBM3e stacks, 9.4TB/s of bandwidth, hardware FP4, and scale-up domains reaching 72 accelerators.

Claude Memory Now Carries Between Chat and Cowork Tasks
Anthropic unified Claude's memory across chat and Cowork, with topic-by-topic editing and sensitive categories excluded by default on Free, Pro and Max.

Gemini Enterprise for Legal Brings AI Agents to Law Firms
Google Cloud launched Gemini Enterprise for Legal with four launch firms, iManage and Thomson Reuters connectors, plus audit logging and ethical walls.
