Google has expanded Gemini’s voice tools with two models designed to turn audio into text. Gemini 3.5 Transcribe processes recordings, while Gemini 3.5 Transcribe Live produces a transcription as a person speaks.
Two models for different needs
Gemini 3.5 Transcribe is designed for existing audio such as interviews, meetings, classes, podcasts, and voice notes.
The model can detect the language used in each segment, separate different speakers, and provide the exact time at which each word was spoken.
It also supports a custom list of up to 1,000 terms to help guide the recognition of names, abbreviations, and technical vocabulary.
More than 85 languages and multiple speakers
The models support more than 85 languages. Google does not claim identical accuracy across all of them, so results may vary depending on the language, voice clarity, microphone, and background noise.
Speaker separation distinguishes different voices in a conversation, but it does not automatically identify each person’s real name. An application may use labels such as “Speaker 1” and “Speaker 2” to organize the conversation.
Real-time transcription
Gemini 3.5 Transcribe Live maintains a continuous connection to process audio while a person is speaking.
The model can provide preliminary results during an utterance and confirm the transcription after detecting that the speaker has finished. This capability can support live captions, voice assistants, or meetings that are transcribed as they happen.
Availability
Both models are available through the Gemini Developer API and can be tested in Google AI Studio. Google presents them as stable models for developers.
The documentation includes a limited free tier and paid plans for applications with higher usage. Free access is not unlimited, and conditions may vary according to service usage.



