Products & Services · · Team LATENT
Gemini 3.8 Flash TTS arrives: A comparison with GPT-Live
Compare Gemini 3.8 Flash TTS and GPT-Live with response speed as the priority. We examine first audio versus substantive answers, language support, speech evaluations across seven languages, API pricing, differing measurement conditions, and Google's 2027 price changes.
On September 23, Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS. For conversation and read-aloud applications, the first question is how long users wait to hear speech. A comparison with GPT-Live needs to distinguish first audio, the start of a substantive answer, and interruption handling. We also compare language support, perceived speech quality, and pricing from published sources.
Key points
- Response measurements use different conditions. GPT-Live's measured median onset in an English app test was 1.099 seconds, but this does not establish response times in other languages, API performance, or a head-to-head result against Gemini.
- Google lists 130 languages for Flash and 101 for Lite. Flash leads their Japanese listening comparison, while Lite leads in English and some other languages.
- Flash TTS currently charges $0.0135 per generated audio minute; GPT-Live charges $0.05 per active session minute. These cover different usage, and Google's prices change in 2027.

What happened
- Google's September 23 announcement introduced Flash TTS for expressive speech and Flash-Lite TTS for high-volume, low-latency use.
- Availability includes the Gemini API and Google AI Studio. This article compares the paid Standard tier of the Gemini Developer API with OpenAI API pricing.
- This article compares published specifications and reported evaluations. LATENT has not conducted its own audio measurements or listening tests.
Primary sources: Google's announcement, GPT-Live specifications
Measure first audio separately from the start of an answer
When speed matters, first align what is being timed. TTS onset is the delay between sending text and receiving the first audio. Conversational response time runs from the end of the user's speech to audible output, and can include turn-end detection, recognition, response generation, networking, and playback delays.
A quick acknowledgment such as 'Let me check' can precede the actual answer. Separating first audio, the start of a useful answer, and answer completion distinguishes conversational rhythm from task completion time. OpenAI likewise recommends evaluating useful spoken response time and task success on the same scenarios when optimizing GPT-Live's backend. GPT-Live latency optimization
Published seconds do not form a speed ranking under matched conditions
Concrete response measurements are available. However, provider statistics for Gemini TTS and a smartphone app test using GPT-Live have different measurement boundaries.
| Target | Reported value | What was measured |
|---|---|---|
| Gemini 3.8 Flash TTS | Latency P50 3.47 seconds | Google AI Studio through OpenRouter; displayed for all locations over one week |
| Gemini 3.8 Flash-Lite TTS | Latency P50 2.66 seconds | Same display scope; not a test with matched scripts and languages |
| GPT-Live 1 ChatGPT app | Median 1.099 seconds P90 1.203 seconds | End of English user speech to first audible output, measured by Agora |
ConditionsOpenRouter's Flash and Lite values checked September 30. P50 is the median; P90 indicates a value at or below which approximately 90% of observations fall. GPT-Live was measured by Agora on July 9. These conditions cannot establish cross-model speed ratios or the fastest model.
GPT-Live's roughly one-second result was measured in an English app test
Agora tested an iPhone 13 running ChatGPT v1.2026.183 on a Plus account, in English, with 30 trials per condition. First audio includes acknowledgments, so the substantive answer need not begin at 1.099 seconds. Median time to stop speaking after an interruption was 1.399 seconds. This is a July English-language snapshot from one device, location, and account. Response times in other languages were not measured, and this result does not guarantee API performance in any language, including English. Agora also sells voice communication services. Methods and limitations
OpenRouter's figures aggregate observed latency on that routing path. They do not mean a direct API call from Japan will always take 2.66 or 3.47 seconds from text submission to audible playback. They are not a comparison with matched audio lengths and streaming conditions; relative first-audio performance in Japanese remains unverified.
For conversational responsiveness, a practical first step is to test GPT-Live's complete conversation flow against a pipeline using Flash-Lite TTS for speech output. This follows their roles as a full-duplex conversation system and low-latency TTS component, rather than a measured speed ranking. Google positions Lite for low latency and high throughput, with streaming available in both TTS models. OpenAI's dedicated TTS also streams and recommends WAV or PCM when response speed matters. Lite's positioning / Gemini streaming / OpenAI TTS formats
Language coverage and natural speech quality are different claims
Google lists 130 languages for Flash and 101 for Lite. Both lists include Japanese, English, Chinese, Cantonese, Korean, French, German, Spanish, Portuguese, Arabic, Hindi, and Vietnamese. Language and script variants also appear in the lists, so counts alone do not establish a quality difference between providers.
| Model | Coverage confirmed in official documentation | Limits when assessing naturalness |
|---|---|---|
| Gemini 3.8 Flash TTS | Google lists 130 languages, including Japanese and English as well as Thai, Finnish, and Swedish | Listening preference results for seven languages appear below |
| Gemini 3.8 Flash-Lite TTS | Google lists 101 languages; the three additional languages above were absent from the Lite list reviewed | The count does not guarantee naturalness across every supported language |
| gpt-4o-mini-tts | Publishes a language list including Japanese, English, Chinese, Korean, and Thai | Voices are optimized for English |
| GPT-Live 1 | Documents language prompting; no exhaustive language list or total was found in the API documentation reviewed | The dedicated TTS language list cannot be assumed to apply to GPT-Live |
ConditionsFlash language list / Lite language list / OpenAI TTS languages / GPT-Live language prompting, checked September 30.
Voice design and line-by-line direction in one production workflow
Flash TTS can create a voice from a natural-language description, then change emotion and delivery speed for individual lines while retaining that voice. It supports two-speaker scripts, laughter, sighs, and backchannel responses. One example is a character explaining something calmly in one scene and reacting with surprise in the next. Google's announcement
The official specifications position Flash TTS for voice fidelity, acting, and regional accents, and Flash-Lite TTS for high-volume read-aloud use and the speech output stage of voice agents. They support 130 and 101 languages respectively, including Japanese. Both share the same API schema, making it easier to switch for different workloads. Flash TTS specifications / Flash-Lite TTS specifications
Voice replication is also available, with a process that verifies the speaker's recorded consent. The announcement also mentions Voice remixing for adjusting an existing voice's pitch and timbre, but describes it as coming soon. It should not be confused with features already available. Available and planned features
Japanese speech evaluations outperform OpenAI's dedicated TTS model
Google published evaluation results from Hume AI and Voice Arena. Hume AI assesses multiple dimensions of speech, including naturalness and expressiveness. Voice Arena calculates Elo ratings from blind pairwise listening preferences. Google's methodology says Gemini was tested using production checkpoints, default sampling settings, and single-attempt generation. Evaluation methodology and results
| Metric or language | Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | OpenAI gpt-4o-mini-tts |
|---|---|---|---|
| Hume AI overall quality index | 0.920 | 0.914 | 0.740 |
| Voice Arena Japanese | 1232 | 1152 | 975 |
| Voice Arena English | 1061 | 1087 | 940 |
| Brazilian Portuguese | 1104 | 1134 | 946 |
| Vietnamese | 1135 | 1156 | 839 |
| Modern Standard Arabic | 1204 | 1181 | 911 |
| Hindi | 1106 | 1076 | 843 |
| Mexican Spanish | 1152 | 1146 | 880 |
ConditionsThird-party evaluations reproduced in Google's September 2026 materials. Language rows use Voice Arena Elo; higher scores indicate stronger listening preferences. These do not isolate naturalness alone and may reflect pronunciation and voice preferences. Hume AI's index uses a different scale, and scores should not be compared across languages. GPT-Live is not included.
These results do not establish an audio-quality winner against GPT-Live
Flash TTS scores above Flash-Lite TTS and gpt-4o-mini-tts in Japanese. In English, however, Lite scores above Flash, so the more expensive model does not always lead. Elo measures listening preferences: 1232 versus 975 does not mean Japanese capability is approximately 26% better.
Within these seven languages, Flash scores above Lite in Japanese, Modern Standard Arabic, Hindi, and Mexican Spanish. Lite scores above Flash in English, Brazilian Portuguese, and Vietnamese. Chinese, Korean, French, and other languages are absent from this published table: support can be confirmed without establishing a naturalness ranking.
The OpenAI model evaluated here is gpt-4o-mini-tts, not GPT-Live 1. The materials reviewed do not provide a matched evaluation of Gemini 3.8 TTS against GPT-Live or side-by-side response-latency measurements. They cannot establish a winner against GPT-Live or a speed multiplier.
The published charts also do not specify per-model vote counts, confidence intervals, or detailed OpenAI model versions. Hume AI itself explains that naturalness, expressiveness, identity stability, and reliability should be assessed separately. These results can help select a Japanese narration model, but do not guarantee pronunciation of proper nouns or stability across long scripts. Hume AI's evaluation framework
GPT-Live keeps conversing while listening to the user
GPT-Live 1 is a voice conversation model that can listen and speak simultaneously. It handles conversations in which users add details or correct themselves, and can delegate reasoning and tool execution to another model or agent. Suitable uses include changing booking requirements during a conversation or continuing to talk while information is being retrieved.
Gemini TTS handles the stage that turns supplied text into audio. Streaming lets playback begin before generation finishes, but it does not include understanding incoming speech and deciding how to respond. A conversational application needs separate speech recognition, response generation, and handling such as stopping playback when the user interrupts. Gemini speech generation guide
| Comparison | Gemini 3.8 TTS | GPT-Live 1 |
|---|---|---|
| Core task | Generate speech from text | Two-way voice conversation |
| Voice control | Voice design and script direction | Instructions for conversational style |
| When the user starts speaking | The app combines recognition and playback interruption handling | Listening and speaking are handled simultaneously |
| Reasoning and external tools | Combine with another model or application logic | Delegate to a backend agent |
ConditionsFeature comparison based on Google's specifications and OpenAI's GPT-Live guide. This is not a measured ranking of audio quality or latency.
For reading a script aloud, include gpt-4o-mini-tts in the comparison
OpenAI, the company behind ChatGPT, also offers a dedicated text-to-speech API separate from GPT-Live. Its Text to speech guide describes gpt-4o-mini-tts, which generates audio from text and a selected voice, with instructions for speed, emotion, intonation, and whispering. It also supports streaming. Its voices are optimized for English, so Japanese users should evaluate it with their own scripts.
Per-request input limits also differ: 2,000 tokens for gpt-4o-mini-tts and 8,192 for Gemini 3.8 TTS. Google's developer specifications list an audio output serving limit of 16,384 tokens, equivalent to about 10.9 minutes at 25 tokens per second. Producing a long audiobook therefore requires splitting the script and checking for changes in voice or volume at the boundaries. OpenAI model specifications / Gemini model specifications / Audio token conversion
Gemini Flash TTS audio output currently costs $0.0135 per minute
In the Gemini Developer API's paid Standard tier, Flash TTS audio output costs $9 per million tokens and Flash-Lite TTS costs $6. Audio is billed at 25 tokens per second, or 1,500 tokens per minute. The following calculations cover audio output only; the input script and instructions cost extra. Official pricing
| Model and period | Input text per million tokens | Audio output per million tokens | Audio output calculated per minute |
|---|---|---|---|
| Flash TTS Through December 31, 2026 | $0.50 | $9 | $0.0135 |
| Flash-Lite TTS Through December 31, 2026 | $0.50 | $6 | $0.009 |
| Flash TTS From January 1, 2027 | $1 | $18 | $0.027 |
| Flash-Lite TTS From January 1, 2027 | $1 | $12 | $0.018 |
ConditionsGemini Developer API prices, checked September 30. Standard pricing in US dollars; excludes free allowances, Batch, Flex, Priority, and caching. Per-minute calculation: audio token price x 25 x 60 / 1,000,000.
OpenAI's dedicated TTS costs $0.60 for input and $12 for audio output
The official gpt-4o-mini-tts specifications list $0.60 per million input text tokens and $12 per million audio output tokens. For example, 100,000 input tokens plus one million audio output tokens would cost $0.06 + $12 = $12.06.
This is not a comparison at equal audio duration. Providers tokenize audio differently, so Google's 25-tokens-per-second conversion cannot be applied to OpenAI. The OpenAI pricing and model pages reviewed do not give an equivalent per-minute conversion, so we have not assigned gpt-4o-mini-tts a flat minute rate. A rigorous cost comparison requires generating the same script and measuring audio duration and actual usage.
A GPT-Live minute measures session time, not generated audio length
GPT-Live 1 costs $0.05 per voice session minute, billed per second. According to OpenAI's cost guide, this includes user speech, assistant speech, silence, and time spent waiting for the backend. Muting the microphone does not end the session. Backend reasoning models and tools are billed separately.
Suppose a 10-minute conversation includes three minutes of assistant speech. GPT-Live's session fee would be $0.50. Generating three minutes of audio with Gemini TTS at current prices would cost $0.0405 for Flash or $0.027 for Lite in audio output fees. The latter figures exclude speech recognition, response generation, input text, and conversation management. The difference cannot simply be treated as savings for an entire voice application.
At scale, the scheduled price change also matters. Audio output for 1,000 generated minutes costs $13.50 for Flash or $9 for Lite in 2026, rising to $27 and $18 in 2027. GPT-Live costs $50 for 1,000 session minutes, plus backend charges. Budgets need separate totals for generated audio duration and active session duration.
Measure conversational delay first, then compare quality and cost
For Japanese narration or character dialogue, the published results and direction features give a reason to try Flash TTS first. For large volumes of standard messages or announcements, compare the same script with the lower-output-cost Flash-Lite TTS and decide whether the quality difference matters. When adding read-aloud features to an OpenAI-based application, include gpt-4o-mini-tts as well.
For responsiveness, test the same short and long Japanese utterances and time first audio, substantive-answer onset, and stopping after an interruption. Check P90 as well as the median to identify occasional long waits. Once candidates meet the required speed on the intended device and network, compare pronunciation errors, naturalness, and costs including retries.
Sources and references
Prices and specifications checked September 30, 2026. Evaluation scores come from Google's September 2026 materials.
- Gemini 3.8 TTS announcement / Google, September 23
- Gemini 3.8 TTS evaluation methodology and results / Google DeepMind, September 2026
- Real World VoiceEQ Bench Hume AI
- Gemini 3.8 Flash TTS model specifications / Google
- Gemini 3.8 Flash-Lite TTS model specifications / Google
- Gemini speech generation guide / Google
- Gemini Developer API pricing / Google
- GPT-Live 1 specifications and pricing / OpenAI
- Voice session billing scope / OpenAI
- GPT-4o mini TTS specifications and pricing / OpenAI
- Text to speech guide / OpenAI
- GPT-Live response measurements on a real device / Agora, published July 10
- Flash TTS latency by serving route / OpenRouter, checked September 30
- Flash-Lite TTS latency by serving route / OpenRouter, checked September 30
- GPT-Live backend latency optimization / OpenAI
- GPT-Live language and pronunciation prompting / OpenAI