Stories about Google AI
1 related stories
Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
AI InsightGoogle AI has released Gemini 3.5 Transcribe, a speech-to-text model that reports an average WER of 2.6% across 85+ languages. The model is split into two endpoints: streaming and batch. The streaming endpoint provides sub-second transcription but drops speaker diarization and word timestamps. The batch endpoint keeps both, at half the cost. Google reports a WER of 4.0% for streaming and 2.6% for non-streaming, with 70% faster finalization than Chirp 3.Google AI releases Gemini 3.5 Transcribe, a speech-to-text model supporting 85+ languages.The release of Gemini 3.5 Transcribe marks a significant breakthrough for Google AI in the field of speech-to-text, particularly in terms of multilingual support. The model's low WER and fast transcription speed will help improve the performance of voice agents and transcription pipelines.- DevelopersWill get access to a better speech-to-text model that supports multiple languages.
Next, we can expect Google AI to continue innovating and improving in the field of speech-to-text, particularly in terms of multilingual support and model efficiency.Importance 85/100