Gemini-3.5-Transcribe(blog.google)
353 points by k9294 13 days ago | 122 comments
tl;dr: Google introduced Gemini 3.5 Transcribe, a speech-to-text model claiming 4.0% WER for streaming and 2.6% for non-streaming, with 70% better latency than its Chirp 3 predecessor. It handles disfluencies, filler word removal, custom vocabulary, speaker diarization (up to 3 speakers), and supports 85+ languages, and can delegate tasks to other Gemini models via function calling. It's available in public preview through the Gemini API, Google AI Studio, and Enterprise Agent Platform, and powers features like Rambler on Android and voice controls in the Gemini macOS app.
HN Discussion:
  • Gemini's replacement of Google Assistant is broken and cannot perform basic tasks like playing songs or setting alarms
  • In personal benchmarks, other STT models like Voxtral, ElevenLabs, or Soniox outperform Gemini Transcribe on accuracy or latency
  • The model over-simplifies speech and removes meaningful content, breaking user intent
  • ~The article's description of function calling is confusing and misleading about the STT model's capabilities
  • Concerns about hallucination issues carrying over from Chirp, and skepticism about reliability