Google unveils Gemini 3.5 Transcribe AI model for audio transcription — The Verge
Google has updated its Gemini Audio lineup and introduced the new Gemini 3.5 Transcribe model for audio transcription. It is designed to automatically recognize specialized terminology and more than 85 languages, format text, and remove filler words such as “uh” and “um,” The Verge reports.
Capabilities of the new model
Google said that Gemini 3.5 Transcribe is a significant step forward compared with the previous Chirp 3 transcription model, particularly in multilingual performance and word error rates. Users can add their own vocabulary so that the model correctly accounts for special word spellings and professional jargon without further manual editing.
The model can also attribute remarks to up to three speakers in pre-recorded audio and add timestamps for every word. According to Google, the tool allows users to edit text with their voice.
More current news is available on the UA.News Telegram channel Telegram.
Gemini Live updates
Alongside Transcribe, the company is launching Gemini 3.5 Live and Gemini 3.5 Live Experimental. These models advance the speech recognition technology used in Gemini’s voice mode. Gemini 3.5 Live is expected to better handle interruptions in the middle of a phrase, recognize languages, and work with visual information in real time.
The experimental version of Gemini 3.5 Live, according to Google’s description, can provide step-by-step spoken explanations of the process of completing more complex reasoning-related tasks. The new capabilities have begun rolling out in English for users of the Gemini app on macOS, as well as for the Rambler dictation feature on Android in selected countries and languages.
For developers, the models are available as a public preview through the Gemini API in AI Studio and Antigravity. Google promised to add Chrome support in the near future.