The Gemini Native Audio Generation Text-to-Speech (TTS) model differentiates itself from traditional TTS models by using a large language model that knows not only what to say, but also how to say it.
Text-to-Speech generates audio with natural, human-like quality, which creates speech that sounds like a real person. To start, specify a voice when sending a synthesis request.
Turn text into natural-sounding speech in 220+ voices across 40+ languages and variants with an API powered by Google’s machine learning technology.
Note: Studio voices support SSML, except for the following tags: <mark>, <emphasis>, <prosody pitch>, and <lang>. Check the table of supported voices for availability of Studio voices in specific...
Get started with Cloud Text-to-Speech in your language of choice. v1 and v1beta1 REST API Reference. v1 and v1beta1 gRPC API Reference. SSML elements supported in Cloud TTS. List of...
Google Cloud Text-to-Speech enables developers to synthesize natural-sounding speech with 30 voices, available in multiple languages and variants. It applies DeepMind’s groundbreaking research in WaveNet and Google’s powerful neural networks t