Last checked: October 8, 2026. Google's official speech guide currently documents Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. This page explains that cloud workflow; VoiceAI Free does not provide a Gemini generator, API proxy or free Google credits.
What the official guide describes
The TTS API takes text and returns audio. It is distinct from the Live API for interactive conversations. The speech guide documents single-speaker and multi-speaker generation, prebuilt voice choices and a WAV response. Sustained delivery instructions belong in structured speech_metadata rather than being mixed into the verbatim transcript. Keep a clean script and the delivery instructions separately so a revision does not accidentally change what is said.
Flash or Flash-Lite?
Google lists two Gemini 3.8 TTS options. Choose using its current model guidance, account access and required workload, not an unverified promise about quality or speed. Rate limits, pricing and availability can change. A demo in AI Studio is not proof of an unlimited API allowance or permission to publish every voice in every context. Read the linked documentation and the terms that apply to your account.
A practical setup workflow
- Open the official TTS documentation and select a supported model.
- Try an original, non-sensitive paragraph in AI Studio if your account offers the feature.
- Choose a prebuilt voice and keep delivery directions separate from the text.
- Review pronunciation, pauses and the actual file format before building automation.
- Set usage limits and keep API credentials on a trusted server, never in a public website bundle.
For two speakers, label turns consistently and map each speaker to a permitted voice. Do not use unauthorized identity replication. A scene with short turns is easier to correct than a long monologue. Keep versioned scripts, generation settings and the checked terms beside the final audio.
Cloud TTS and private local work are different
Gemini generation sends text to Google under its applicable service terms. That may suit an approved production pipeline, but is not browser-local inference. VoiceAI Free's optional Kokoro path downloads model resources and then generates English WAVs on the device. It has eight curated voices and a 1,000-character per-job limit. Device playback is another path and cannot create a downloadable file through Web Speech.
Use local playback to proof a script before cloud rendering when that fits your data policy. Do not paste confidential client material into any cloud demo simply because a comparison calls it free. Browser drafts remain on the current profile and are not a secure shared archive. The privacy decision should be made before the first generation.
What this guide does not claim
We publish no controlled quality ranking, latency number or benchmark comparing Gemini with Kokoro. Official model capabilities are not measurements on your device. We do not claim unlimited free access, a blanket commercial license or a integration that is not present. Verify output rights, input rights and current account conditions before public or commercial use.
FAQ
Can I run Gemini 3.8 inside this browser toolkit?
No. Our local AI engine is Kokoro. Follow Google's own workflow for Gemini TTS.
Can I save generated audio?
Google's documentation shows WAV output. Validate the returned MIME type and save a real WAV, not a renamed compressed file.
Do I need an API key here?
No key is needed for VoiceAI Free. Never paste a Google key into this site.
Official sources and disclosure
Google's speech generation documentation and Gemini API pricing. These are source links, not affiliate links. Review them for changing model access and terms.