KBD
Back to Blog

29 August 2026

Gemini 3.5 Transcribe: Google's New AI Speech-to-Text Model Explained

Google announced Gemini 3.5 Transcribe on August 26, 2026, its most precise speech-to-text model yet. Here's what it actually does, the real benchmark numbers, the two API modes, and why it matters if you're building voice-enabled AI agents.

Gemini 3.5 Transcribe: Google's New AI Speech-to-Text Model Explained

Announced on August 26, 2026, Google introduced Gemini 3.5 Transcribe, its most precise speech-to-text model yet. Designed to convert raw audio directly into accurate, polished, and formatted text, this model marks a significant leap forward for both live conversations and pre-recorded audio files like meetings and calls. At Kashtbhanjan Digital, an AI agent development and SEO agency serving global and DACH markets since 2004, we track releases like this closely, since they directly shape what's possible when we build voice-enabled AI agents for clients.

What Makes Google's New Transcription Model Different?

Unlike legacy transcription tools that output literal, messy transcripts filled with stutters and disfluencies, Google designed this model to understand intent and context. It handles self-corrections seamlessly, for example, when a speaker says, "let's meet Tuesday, no, Wednesday," the system automatically cleans up the phrasing in the final output. It also excels with alphanumeric entities like postal codes and order IDs, even when captured in noisy real-world audio environments.

Key Features of Gemini 3.5 Transcribe

The model comes with a set of capabilities aimed at developers, enterprises, and consumer applications. Its performance benchmarks show an average Word Error Rate (WER) of 4.0% in streaming mode and 2.6% in non-streaming mode. On the FLEURS multilingual benchmark, it achieves 5.50% WER streaming and 5.04% WER non-streaming, delivering a 70% faster time-to-final-transcription compared to Google's previous model, Chirp 3.

Automatic Filler Word Removal and Polishing

Raw human speech is naturally disorganized. The model processes audio input to automatically remove filler words such as "ums" and "ahs" while auto-formatting the resulting text with proper punctuation and structure. This cuts down the manual editing time typically needed for meeting notes, podcasts, and customer service call logs.

Support for 85+ Languages and Code-Switching

The model automatically detects and transcribes over 85 languages, including regional accents and dialects, and can handle live language switches mid-conversation, making it well suited for global enterprises and multilingual customer support.

It also supports multi-speaker identification (diarization), accurately attributing speech to up to 3 speakers with word-level timestamps in pre-recorded audio, with support for more than 3 speakers currently experimental. Developers can add custom vocabulary so domain-specific jargon and unique brand spellings get transcribed correctly.

How to Build and Integrate with the Gemini API

Developers can access Gemini 3.5 Transcribe through multiple public preview channels: the Gemini API in Google AI Studio, and Google Antigravity. Enterprise users can access it via the Gemini Enterprise Agent Platform, with availability coming soon to Gemini Enterprise for Customer Experience. Consumer implementations include the Gemini app on macOS (English), Rambler on Gboard for Android (select countries and languages), with Chrome talk-to-type support coming soon.

The model is also integrated with third-party developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. No public pricing has been disclosed yet.

Two Ways to Use It: Live API vs Interactions API

  • Real-time streaming via the Live API (model name gemini-3.5-transcribe-live): continuous, bidirectional streaming with sub-second latency, built for interactive voice apps and real-time assistants.
  • Pre-recorded audio processing via the Interactions API (model name gemini-3.5-transcribe): built for batch processing meetings, call logs, and recorded audio, with speaker attribution and word-level timestamps.

For tasks that need to hand off to other models, such as triggering image generation or file analysis, the model supports function calling, currently available in the Gemini macOS app.

Why This Matters If You're Building AI Agents

This is where the announcement gets directly relevant to businesses, not just developers. Voice is the natural interface for a lot of the AI agents companies are building right now: customer support agents that pick up a call, meeting-transcription agents that turn a sales call into structured notes, and call-analysis agents that flag what actually got said versus what got promised. A model that removes filler words, corrects itself, attributes speakers, and holds up in noisy real-world audio is a meaningfully better foundation for all three.

If you're weighing whether a customer service AI agent is worth building for your business, better transcription infrastructure is exactly the kind of underlying improvement that makes those agents more reliable in production, not just in a demo.

Frequently Asked Questions

Can I use Google Gemini to transcribe audio?

Yes. Gemini 3.5 Transcribe converts raw audio files and speech into accurate, polished text while automatically detecting multiple languages.

Is there a free way to try Gemini 3.5 Transcribe?

It's in public preview via the Gemini API in Google AI Studio, and also shows up in consumer features like Rambler on Gboard and the Gemini app on macOS. No official pricing for API usage at scale has been disclosed yet.

Does it support word-level timestamps?

Yes, in the pre-recorded audio mode via the Interactions API. This is useful for subtitling, interactive transcripts, and detailed audio analysis.

Can it tell different speakers apart?

Yes, it supports speaker diarization for up to 3 speakers with word-level timestamps in pre-recorded audio. Support for more than 3 speakers is currently experimental.

What's the actual accuracy like?

Google reports a 4.0% Word Error Rate in streaming mode and 2.6% in non-streaming mode, with a 70% faster time-to-final-transcription than its previous model, Chirp 3.

Partner with Kashtbhanjan Digital

Models like Gemini 3.5 Transcribe are exactly what we build on when developing custom AI agents, whether that's a voice-enabled support agent, a meeting-notes automation, or a call-analysis tool for your sales team. If you want help figuring out what a model release like this actually means for your business, get in touch.

Select your location below for country-specific AI agent development services.

Want to Build a Voice-Enabled AI Agent?

We build custom AI agents on top of the latest models, including voice-enabled customer support, meeting transcription, and call analysis agents.

Book a Free AI Agent Development Consultation

Work With Us

Want to grow your business with SEO, AI automation, or a new website?

Kashtbhanjan Digital has been helping businesses rank on Google, automate with AI agents, and build professional websites for over 20 years.

Gemini 3.5 Transcribe: Features, API & AI Speech-to-Text