Google 发布 Gemini 3.5 Transcribe 实时转写模型
Google introduces the Gemini 3.5 Transcribe real-time speech model
Google 发布面向实时语音交互的 Gemini 3.5 Transcribe,可把原始音频直接整理为带格式的文字,并针对噪声、专业术语与口语停顿进行处理。Google 称这是其目前最精确的语音转文字模型;该表述与相关质量结果来自厂商评测,仍需结合语言、口音和场景独立验证。
Google has introduced Gemini 3.5 Transcribe for real-time voice interactions. It converts raw audio into formatted text while handling noise, specialist vocabulary, and spoken disfluencies. Google calls it its most precise speech-to-text model to date; that description and the supporting quality results are vendor evaluations that still need independent testing across languages, accents, and settings.