Announced on August 26, 2026, Google introduced Gemini 3.5 Transcribe, its most precise speech-to-text model yet. Designed to convert raw audio directly into accurate, polished, and formatted text, this model marks a significant leap forward for both live conversations and pre-recorded audio files like meetings and calls. At Kashtbhanjan Digital, an AI agent development and SEO agency serving global and DACH markets since 2004, we track releases like this closely, since they directly shape what's possible when we build voice-enabled AI agents for clients.
What Makes Google's New Transcription Model Different?
Unlike legacy transcription tools that output literal, messy transcripts filled with stutters and disfluencies, Google designed this model to understand intent and context. It handles self-corrections seamlessly, for example, when a speaker says, "let's meet Tuesday, no, Wednesday," the system automatically cleans up the phrasing in the final output. It also excels with alphanumeric entities like postal codes and order IDs, even when captured in noisy real-world audio environments.
Key Features of Gemini 3.5 Transcribe
The model comes with a set of capabilities aimed at developers, enterprises, and consumer applications. Its performance benchmarks show an average Word Error Rate (WER) of 4.0% in streaming mode and 2.6% in non-streaming mode. On the FLEURS multilingual benchmark, it achieves 5.50% WER streaming and 5.04% WER non-streaming, delivering a 70% faster time-to-final-transcription compared to Google's previous model, Chirp 3.
Automatic Filler Word Removal and Polishing
Raw human speech is naturally disorganized. The model processes audio input to automatically remove filler words such as "ums" and "ahs" while auto-formatting the resulting text with proper punctuation and structure. This cuts down the manual editing time typically needed for meeting notes, podcasts, and customer service call logs.
Support for 85+ Languages and Code-Switching
The model automatically detects and transcribes over 85 languages, including regional accents and dialects, and can handle live language switches mid-conversation, making it well suited for global enterprises and multilingual customer support.
It also supports multi-speaker identification (diarization), accurately attributing speech to up to 3 speakers with word-level timestamps in pre-recorded audio, with support for more than 3 speakers currently experimental. Developers can add custom vocabulary so domain-specific jargon and unique brand spellings get transcribed correctly.
How to Build and Integrate with the Gemini API
Developers can access Gemini 3.5 Transcribe through multiple public preview channels: the Gemini API in Google AI Studio, and Google Antigravity. Enterprise users can access it via the Gemini Enterprise Agent Platform, with availability coming soon to Gemini Enterprise for Customer Experience. Consumer implementations include the Gemini app on macOS (English), Rambler on Gboard for Android (select countries and languages), with Chrome talk-to-type support coming soon.
The model is also integrated with third-party developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. No public pricing has been disclosed yet.
Two Ways to Use It: Live API vs Interactions API
- Real-time streaming via the Live API (model name
gemini-3.5-transcribe-live): continuous, bidirectional streaming with sub-second latency, built for interactive voice apps and real-time assistants. - Pre-recorded audio processing via the Interactions API (model name
gemini-3.5-transcribe): built for batch processing meetings, call logs, and recorded audio, with speaker attribution and word-level timestamps.
For tasks that need to hand off to other models, such as triggering image generation or file analysis, the model supports function calling, currently available in the Gemini macOS app.
Why This Matters If You're Building AI Agents
This is where the announcement gets directly relevant to businesses, not just developers. Voice is the natural interface for a lot of the AI agents companies are building right now: customer support agents that pick up a call, meeting-transcription agents that turn a sales call into structured notes, and call-analysis agents that flag what actually got said versus what got promised. A model that removes filler words, corrects itself, attributes speakers, and holds up in noisy real-world audio is a meaningfully better foundation for all three.
If you're weighing whether a customer service AI agent is worth building for your business, better transcription infrastructure is exactly the kind of underlying improvement that makes those agents more reliable in production, not just in a demo.
Frequently Asked Questions
Can I use Google Gemini to transcribe audio?
Yes. Gemini 3.5 Transcribe converts raw audio files and speech into accurate, polished text while automatically detecting multiple languages.
Is there a free way to try Gemini 3.5 Transcribe?
It's in public preview via the Gemini API in Google AI Studio, and also shows up in consumer features like Rambler on Gboard and the Gemini app on macOS. No official pricing for API usage at scale has been disclosed yet.
Does it support word-level timestamps?
Yes, in the pre-recorded audio mode via the Interactions API. This is useful for subtitling, interactive transcripts, and detailed audio analysis.
Can it tell different speakers apart?
Yes, it supports speaker diarization for up to 3 speakers with word-level timestamps in pre-recorded audio. Support for more than 3 speakers is currently experimental.
What's the actual accuracy like?
Google reports a 4.0% Word Error Rate in streaming mode and 2.6% in non-streaming mode, with a 70% faster time-to-final-transcription than its previous model, Chirp 3.
Partner with Kashtbhanjan Digital
Models like Gemini 3.5 Transcribe are exactly what we build on when developing custom AI agents, whether that's a voice-enabled support agent, a meeting-notes automation, or a call-analysis tool for your sales team. If you want help figuring out what a model release like this actually means for your business, get in touch.
Select your location below for country-specific AI agent development services.
India
Custom AI agent development for Indian businesses and startups.
View India services →United Kingdom
AI agents built for UK businesses, UK GDPR compliant.
View UK services →Germany
AI agent development for German companies, GDPR and EU AI Act aligned.
View Germany services →Australia
Custom AI agents for Australian enterprises, Privacy Act 1988 compliant.
View Australia services →Canada
AI agent development for Canadian businesses, PIPEDA compliant.
View Canada services →New York
AI agents for New York enterprises and fast-growing startups.
View New York services →New Jersey
AI agent development for New Jersey businesses looking to automate operations.
View New Jersey services →Want to Build a Voice-Enabled AI Agent?
We build custom AI agents on top of the latest models, including voice-enabled customer support, meeting transcription, and call analysis agents.
Book a Free AI Agent Development Consultation