Gemini 3.8 Flash-Lite TTS

Gemini 3.8 Flash-Lite TTS
Google · Audio Generation
POST /v1/audio/speech

Cost-efficient speech for voice agents and dubbing, with streaming, vocal tags, and two-speaker dialogue across 101 languages and 2,000+ voices.

At a glance

FieldValue
Model idgemini-3-8-flash-lite-tts
Model release date2026-09-22
Input modalitiesText
Output modalitiesAudio
Context window-
Weight precision-
Featurestext_to_speech, multi_speaker, multilingual, voice_tags, voice_control, streaming, voice_cloning
Native inferenceNo
NewYes
Supported endpointsPOST /v1/audio/speech, POST /v1/audio/speech:stream, GET /v1/voices
Alternate model idsgemini-3.8-flash-lite-tts, google/gemini-3.8-flash-lite-tts

Pricing

ChargeSpecRate
Inputper 1M prompt tokens$2.60
Outputper 1M generated tokens$31.20

Example request

curl https://api.empiriolabs.ai/v1/audio/speech \
-H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
-H 'Content-Type: application/json' \
-d '{"model": "gemini-3-8-flash-lite-tts", "input": "Hello from EmpirioLabs."}'

Parameters

ParameterTypeRequiredDefaultDescription
inputstringyes-Text to convert to speech. For multi-speaker mode, prefix lines with Speaker1: / Speaker2:. Up to 5,000 characters. Put delivery direction in style_prompt, and place these vocal tags inline where the sound should happen: <argh>, <breath>, <heavy breath>, <exhales>, <cackle>, <cheer>, <chuckle>, <chuckles>, <cough>, <cry>, <gasp>, <giggle>, <groan>, <growl>, <grunt>, <grr>, <hiss>, <laugh>, <laughter>, <moan>, <pant>, <pff>, <phew>, <scream>, <shout>, <shriek>, <sigh>, <sighs>, <sneeze>, <snicker>, <snort>, <sob>, <throat-clearing>, <tsk>, <whimper>, <whispers>, <whispering>, <yawn>, <short pause>, <long pause>. Capitalize a word to stress it, and wrap a listener reaction in pipes, for example |mm-hmm|. · Max: 5000
modeenumno"single"single = one voice, multi = two-voice dialogue (uses voice + voice2 + speaker names). · Allowed: single, multi
languagestringno-Optional BCP-47 language tag (en-US, es-ES, etc.). The language of the text is detected automatically. The 30 listed voices speak every supported language; a library voice must be used with its own language.
voiceenumno"Charon"Primary voice name (e.g. Kore, Puck, Aoede). Leave blank for the default. These 30 voices speak every supported language. GET /v1/voices lists the full library of more than 2,000 voices, and any voice id it returns is also accepted here. custom clones a voice from voice_audio_url and voice_consent_audio_url, and the response’s voice_id then works here or in voice2 for 7 days. · Allowed: Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat, custom
voice_audio_urlstringno-voice=custom only: an https URL to a 10 to 30 second recording of clean, natural speech from the voice to clone (WAV, MP3, M4A, OGG, WEBM or FLAC). Record it in a quiet room, on the same microphone as the consent recording.
voice_consent_audio_urlstringno-voice=custom only: an https URL to the same speaker reading this statement aloud: “I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model.” The same statement in any of 30 languages is accepted; consent_statements in this parameter’s metadata lists each one.
voice2enumno"Kore"Second voice name for multi-speaker mode. · Allowed: Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirrhoe, Autonoe, Enceladus, Iapetus, Umbriel, Algieba, Despina, Erinome, Algenib, Rasalgethi, Laomedeia, Achernar, Alnilam, Schedar, Gacrux, Pulcherrima, Achird, Zubenelgenubi, Vindemiatrix, Sadachbia, Sadaltager, Sulafat
speaker1_namestringno"Speaker1"Display name used in the input prefix for speaker 1 (default: Speaker1).
speaker2_namestringno"Speaker2"Display name used in the input prefix for speaker 2 (default: Speaker2).
output_formatenumno"WAV"Audio file format: WAV, MP3, OGG (Opus), or ALAW / MULAW for telephony. · Allowed: WAV, MP3, OGG, ALAW, MULAW
speednumberno1.0Playback rate. 1.0 = natural; <1 slower, >1 faster. · Range: 0.25 – 2.0
volume_gainnumberno0Output gain in dB. 0 = unchanged. · Range: -96 – 16
sample_rateenumno"24000"Output sample rate in Hz (8000, 16000, 22050, 24000, 44100, or 48000). · Allowed: 8000, 16000, 22050, 24000, 44100, 48000
style_promptstringno-Natural-language style direction (e.g. “warm, conversational” or “newscaster, serious”).

Notes

Limits

  • Text: up to 5,000 characters per request
  • Audio billing: 32 output tokens per second of generated audio
  • 101 languages. The language of the text is detected automatically.

Voices

  • The 30 voices in the voice list speak every supported language.
  • GET /v1/voices?model=gemini-3-8-flash-lite-tts lists the full library of more than 2,000 voices across 30 locales, with the language, accent, gender, pitch, and a description of each. Filter it with language, gender, pitch, or q. Any id it returns works as voice or voice2; when you also set language, use the voice’s own language.

Custom voices

  • Set voice to custom and pass two recordings of the same adult speaker as https URLs: voice_audio_url, 10 to 30 seconds of natural speech, and voice_consent_audio_url, the speaker reading “I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model.” The statement is accepted in 30 languages; consent_statements on that parameter lists each wording.
  • Record both in a quiet room on the same microphone. WAV, MP3, M4A, OGG, WEBM and FLAC are accepted, up to 15 MB each.
  • The response includes voice_id and voice_id_expires_at. Pass the voice_id as voice or voice2 to reuse the voice for 7 days without new recordings. Anyone with the id can use the voice until it expires, so keep it private.

Delivery direction

  • Put direction for the whole passage in style_prompt rather than in the text, for example “warm, unhurried, and reassuring”.
  • Place vocal tags inline where the sound should happen. Use these English tags even when the text is in another language:
<argh> <breath> <heavy breath> <exhales> <cackle> <cheer> <chuckle> <chuckles> <cough> <cry> <gasp> <giggle> <groan> <growl> <grunt> <grr> <hiss> <laugh> <laughter> <moan> <pant> <pff> <phew> <scream> <shout> <shriek> <sigh> <sighs> <sneeze> <snicker> <snort> <sob> <throat-clearing> <tsk> <whimper> <whispers> <whispering> <yawn> <short pause> <long pause>
  • Capitalize a word to stress it. Wrap a listener reaction in pipes, for example |mm-hmm|, to add it without starting a new turn.

Multi-speaker

  • Up to two speakers per request. Start each line with a speaker name and a colon, for example Speaker1: and Speaker2:. The names must match speaker1_name and speaker2_name.

Output

  • WAV, MP3, OGG (Opus), ALAW, or MULAW, at 8,000 to 48,000 Hz.
  • POST /v1/audio/speech:stream sends 16-bit PCM chunks at 24,000 Hz while the audio is generated, then the complete file in the requested format. speed is available on POST /v1/audio/speech only.

Machine-readable schema: GET https://api.empiriolabs.ai/v1/models/gemini-3-8-flash-lite-tts.