Skip to content
AZ Labs
Product Updates26 August 2026•3 min read

Gemini 3.5 Transcribe adds streaming and recorded-audio transcription

AZ Labs editorial illustration for Gemini 3.5 Transcribe
Inspect
Original AZ Labs editorial illustration Original AZ Labs editorial illustration • © 2026 AZ Labs

Google's dedicated speech-to-text model has separate routes for live speech and recorded audio. Its release predates the September Gemini audio announcements.

AI Neural Narration

48kHz Studio

Fish Audio Neural Engine · Natural editorial narration

0:000:00
verified

Key Takeaways

  • check_circleThe original release date is 26 August 2026, not September.
  • check_circleStreaming and recorded audio use separate API routes.
  • check_circleCompare transcripts on the terminology and accents your team actually uses.

An August release in the current audio family

Google introduced Transcribe on 26 August 2026 in public preview. It supports custom vocabulary, smart transcription and more than 85 languages. September coverage refers to that earlier release.

Streaming uses gemini-3.5-transcribe-live through the Live API. Recorded audio uses gemini-3.5-transcribe through the Interactions API, with speaker attribution and word-level timestamps. Identification beyond three speakers is experimental.

Check the words that carry the meaning

Our assessment: a transcript can read fluently while getting the most consequential detail wrong. Include names, dates, amounts and product codes in your evaluation set. Compare the output with a human reference and record errors in those fields separately from filler words. For a sales call, a wrong follow-up date may matter more than several missing hesitations.

Build the sample set from the audio conditions your service encounters. Include a noisy room, a poor microphone, regional accents and speakers who correct themselves. Review both raw accuracy and the usefulness of any cleaned-up text. An editor should be able to tell whether a change improved readability or altered what the speaker meant.

Keep transcription and downstream actions separate

Before a transcript updates a customer record, identify which fields require human review. Display uncertain names or numbers beside the relevant audio segment where possible. Do not let a polished sentence conceal uncertainty about an order identifier. A transcript is evidence for the workflow, rather than automatic permission to carry out every request it contains.

For recorded meetings, assess speaker attribution with the kinds of groups you actually record. For streaming captions, measure the delay users experience and what happens when a speaker restarts a sentence. Keep an audio retention policy and an approved vocabulary list alongside the integration. Re-evaluate the same sample set when a preview model or its configuration changes, so quality drift is visible.

Frequently Asked Questions

Was Gemini 3.5 Transcribe released in September?

No. Google's original announcement is dated 26 August 2026; September coverage discusses it as part of the broader audio suite.

Which model route handles recorded audio?

Google lists gemini-3.5-transcribe through the Interactions API. Streaming speech uses gemini-3.5-transcribe-live through the Live API.

Primary Sources

Share this articlePost on X
arrow_backBack to all news