跳到正文
原文
Google DeepMind·· 2026-08-27精选AI 评分71

Google DeepMind 发布 Gemini 3.5 Transcribe 语音转文字模型

Intelligent transcription with Gemini 3.5 Transcribe

AI 导读

Google DeepMind 发布 Gemini 3.5 Transcribe,称其为目前最精确的语音转文字模型,可将原始音频直接转为准确、格式化文本。据 Artificial Analysis 测量,流式场景平均 WER 为 4.0%,非流式 2.6%;相比此前的 Chirp 3,最终转写时间改善 70%。

推荐理由

官方给出流式与非流式两套 API 的 WER 数据和相对 Chirp 3 的延迟提升,便于判断语音转写能否接入现有工作流。

来源:Google DeepMind · deepmind.google