Google AI に関する報道
関連記事 1 件
Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
AI インサイトGoogle AIは、Gemini 3.5 Transcribeという音声からテキストへの変換モデルを発表しました。このモデルは、85以上の言語で平均WER 2.6%を報告しています。このモデルは、ストリーミングエンドポイントとバッチエンドポイントの2つのエンドポイントに分けられています。ストリーミングエンドポイントはサブ秒単位の転写を提供しますが、話し手の識別と単語のタイムスタンプはサポートしていません。一方、バッチエンドポイントはこれら2つの機能をサポートしていますが、コストはストリーミングエンドポイントの半分です。Googleは、ストリーミングエンドポイントのWERは4.0%、非ストリーミングエンドポイントのWERは2.6%で、Chirp 3よりも70%高速であると報告しています。Google AI releases Gemini 3.5 Transcribe, a speech-to-text model supporting 85+ languages.The release of Gemini 3.5 Transcribe marks a significant breakthrough for Google AI in the field of speech-to-text, particularly in terms of multilingual support. The model's low WER and fast transcription speed will help improve the performance of voice agents and transcription pipelines.- DevelopersWill get access to a better speech-to-text model that supports multiple languages.
Next, we can expect Google AI to continue innovating and improving in the field of speech-to-text, particularly in terms of multilingual support and model efficiency.重要度 85/100