AI news, with the context that matters.

Products & Services · ·

Gemini 3.8 Flash TTS arrives: A comparison with GPT-Live

Compare Gemini 3.8 Flash TTS and GPT-Live with response speed as the priority. We examine first audio versus substantive answers, language support, speech evaluations across seven languages, API pricing, differing measurement conditions, and Google's 2027 price changes.

On September 23, Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS. For conversation and read-aloud applications, the first question is how long users wait to hear speech. A comparison with GPT-Live needs to distinguish first audio, the start of a substantive answer, and interruption handling. We also compare language support, perceived speech quality, and pricing from published sources.

Key points

  1. Response measurements use different conditions. GPT-Live's measured median onset in an English app test was 1.099 seconds, but this does not establish response times in other languages, API performance, or a head-to-head result against Gemini.
  2. Google lists 130 languages for Flash and 101 for Lite. Flash leads their Japanese listening comparison, while Lite leads in English and some other languages.
  3. Flash TTS currently charges $0.0135 per generated audio minute; GPT-Live charges $0.05 per active session minute. These cover different usage, and Google's prices change in 2027.
Official announcement image for Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS
Source: Google

What happened

  • Google's September 23 announcement introduced Flash TTS for expressive speech and Flash-Lite TTS for high-volume, low-latency use.
  • Availability includes the Gemini API and Google AI Studio. This article compares the paid Standard tier of the Gemini Developer API with OpenAI API pricing.
  • This article compares published specifications and reported evaluations. LATENT has not conducted its own audio measurements or listening tests.

Primary sources: Google's announcement, GPT-Live specifications

Measure first audio separately from the start of an answer

When speed matters, first align what is being timed. TTS onset is the delay between sending text and receiving the first audio. Conversational response time runs from the end of the user's speech to audible output, and can include turn-end detection, recognition, response generation, networking, and playback delays.

A quick acknowledgment such as 'Let me check' can precede the actual answer. Separating first audio, the start of a useful answer, and answer completion distinguishes conversational rhythm from task completion time. OpenAI likewise recommends evaluating useful spoken response time and task success on the same scenarios when optimizing GPT-Live's backend. GPT-Live latency optimization

Published seconds do not form a speed ranking under matched conditions

Concrete response measurements are available. However, provider statistics for Gemini TTS and a smartphone app test using GPT-Live have different measurement boundaries.

Published response measurements and their scope
TargetReported valueWhat was measured
Gemini 3.8 Flash TTSLatency P50
3.47 seconds
Google AI Studio through OpenRouter; displayed for all locations over one week
Gemini 3.8 Flash-Lite TTSLatency P50
2.66 seconds
Same display scope; not a test with matched scripts and languages
GPT-Live 1
ChatGPT app
Median 1.099 seconds
P90 1.203 seconds
End of English user speech to first audible output, measured by Agora

ConditionsOpenRouter's Flash and Lite values checked September 30. P50 is the median; P90 indicates a value at or below which approximately 90% of observations fall. GPT-Live was measured by Agora on July 9. These conditions cannot establish cross-model speed ratios or the fastest model.

GPT-Live's roughly one-second result was measured in an English app test

Agora tested an iPhone 13 running ChatGPT v1.2026.183 on a Plus account, in English, with 30 trials per condition. First audio includes acknowledgments, so the substantive answer need not begin at 1.099 seconds. Median time to stop speaking after an interruption was 1.399 seconds. This is a July English-language snapshot from one device, location, and account. Response times in other languages were not measured, and this result does not guarantee API performance in any language, including English. Agora also sells voice communication services. Methods and limitations

OpenRouter's figures aggregate observed latency on that routing path. They do not mean a direct API call from Japan will always take 2.66 or 3.47 seconds from text submission to audible playback. They are not a comparison with matched audio lengths and streaming conditions; relative first-audio performance in Japanese remains unverified.

For conversational responsiveness, a practical first step is to test GPT-Live's complete conversation flow against a pipeline using Flash-Lite TTS for speech output. This follows their roles as a full-duplex conversation system and low-latency TTS component, rather than a measured speed ranking. Google positions Lite for low latency and high throughput, with streaming available in both TTS models. OpenAI's dedicated TTS also streams and recommends WAV or PCM when response speed matters. Lite's positioning / Gemini streaming / OpenAI TTS formats

Language coverage and natural speech quality are different claims

Google lists 130 languages for Flash and 101 for Lite. Both lists include Japanese, English, Chinese, Cantonese, Korean, French, German, Spanish, Portuguese, Arabic, Hindi, and Vietnamese. Language and script variants also appear in the lists, so counts alone do not establish a quality difference between providers.

Flash and Lite also differ in language coverage
ModelCoverage confirmed in official documentationLimits when assessing naturalness
Gemini 3.8 Flash TTSGoogle lists 130 languages, including Japanese and English as well as Thai, Finnish, and SwedishListening preference results for seven languages appear below
Gemini 3.8 Flash-Lite TTSGoogle lists 101 languages; the three additional languages above were absent from the Lite list reviewedThe count does not guarantee naturalness across every supported language
gpt-4o-mini-ttsPublishes a language list including Japanese, English, Chinese, Korean, and ThaiVoices are optimized for English
GPT-Live 1Documents language prompting; no exhaustive language list or total was found in the API documentation reviewedThe dedicated TTS language list cannot be assumed to apply to GPT-Live

ConditionsFlash language list / Lite language list / OpenAI TTS languages / GPT-Live language prompting, checked September 30.

Voice design and line-by-line direction in one production workflow

Flash TTS can create a voice from a natural-language description, then change emotion and delivery speed for individual lines while retaining that voice. It supports two-speaker scripts, laughter, sighs, and backchannel responses. One example is a character explaining something calmly in one scene and reacting with surprise in the next. Google's announcement

The official specifications position Flash TTS for voice fidelity, acting, and regional accents, and Flash-Lite TTS for high-volume read-aloud use and the speech output stage of voice agents. They support 130 and 101 languages respectively, including Japanese. Both share the same API schema, making it easier to switch for different workloads. Flash TTS specifications / Flash-Lite TTS specifications

Voice replication is also available, with a process that verifies the speaker's recorded consent. The announcement also mentions Voice remixing for adjusting an existing voice's pitch and timbre, but describes it as coming soon. It should not be confused with features already available. Available and planned features

Japanese speech evaluations outperform OpenAI's dedicated TTS model

Google published evaluation results from Hume AI and Voice Arena. Hume AI assesses multiple dimensions of speech, including naturalness and expressiveness. Voice Arena calculates Elo ratings from blind pairwise listening preferences. Google's methodology says Gemini was tested using production checkpoints, default sampling settings, and single-attempt generation. Evaluation methodology and results

Seven-language listening evaluations show different strengths for Flash and Lite
Metric or languageGemini 3.8
Flash TTS
Gemini 3.8
Flash-Lite TTS
OpenAI
gpt-4o-mini-tts
Hume AI overall quality index0.9200.9140.740
Voice Arena Japanese12321152975
Voice Arena English10611087940
Brazilian Portuguese11041134946
Vietnamese11351156839
Modern Standard Arabic12041181911
Hindi11061076843
Mexican Spanish11521146880

ConditionsThird-party evaluations reproduced in Google's September 2026 materials. Language rows use Voice Arena Elo; higher scores indicate stronger listening preferences. These do not isolate naturalness alone and may reflect pronunciation and voice preferences. Hume AI's index uses a different scale, and scores should not be compared across languages. GPT-Live is not included.

These results do not establish an audio-quality winner against GPT-Live

Flash TTS scores above Flash-Lite TTS and gpt-4o-mini-tts in Japanese. In English, however, Lite scores above Flash, so the more expensive model does not always lead. Elo measures listening preferences: 1232 versus 975 does not mean Japanese capability is approximately 26% better.

Within these seven languages, Flash scores above Lite in Japanese, Modern Standard Arabic, Hindi, and Mexican Spanish. Lite scores above Flash in English, Brazilian Portuguese, and Vietnamese. Chinese, Korean, French, and other languages are absent from this published table: support can be confirmed without establishing a naturalness ranking.

The OpenAI model evaluated here is gpt-4o-mini-tts, not GPT-Live 1. The materials reviewed do not provide a matched evaluation of Gemini 3.8 TTS against GPT-Live or side-by-side response-latency measurements. They cannot establish a winner against GPT-Live or a speed multiplier.

The published charts also do not specify per-model vote counts, confidence intervals, or detailed OpenAI model versions. Hume AI itself explains that naturalness, expressiveness, identity stability, and reliability should be assessed separately. These results can help select a Japanese narration model, but do not guarantee pronunciation of proper nouns or stability across long scripts. Hume AI's evaluation framework

GPT-Live keeps conversing while listening to the user

GPT-Live 1 is a voice conversation model that can listen and speak simultaneously. It handles conversations in which users add details or correct themselves, and can delegate reasoning and tool execution to another model or agent. Suitable uses include changing booking requirements during a conversation or continuing to talk while information is being retrieved.

Gemini TTS handles the stage that turns supplied text into audio. Streaming lets playback begin before generation finishes, but it does not include understanding incoming speech and deciding how to respond. A conversational application needs separate speech recognition, response generation, and handling such as stopping playback when the user interrupts. Gemini speech generation guide

Reading a script and holding a live conversation require different features
ComparisonGemini 3.8 TTSGPT-Live 1
Core taskGenerate speech from textTwo-way voice conversation
Voice controlVoice design and script directionInstructions for conversational style
When the user starts speakingThe app combines recognition and playback interruption handlingListening and speaking are handled simultaneously
Reasoning and external toolsCombine with another model or application logicDelegate to a backend agent

ConditionsFeature comparison based on Google's specifications and OpenAI's GPT-Live guide. This is not a measured ranking of audio quality or latency.

For reading a script aloud, include gpt-4o-mini-tts in the comparison

OpenAI, the company behind ChatGPT, also offers a dedicated text-to-speech API separate from GPT-Live. Its Text to speech guide describes gpt-4o-mini-tts, which generates audio from text and a selected voice, with instructions for speed, emotion, intonation, and whispering. It also supports streaming. Its voices are optimized for English, so Japanese users should evaluate it with their own scripts.

Per-request input limits also differ: 2,000 tokens for gpt-4o-mini-tts and 8,192 for Gemini 3.8 TTS. Google's developer specifications list an audio output serving limit of 16,384 tokens, equivalent to about 10.9 minutes at 25 tokens per second. Producing a long audiobook therefore requires splitting the script and checking for changes in voice or volume at the boundaries. OpenAI model specifications / Gemini model specifications / Audio token conversion

Gemini Flash TTS audio output currently costs $0.0135 per minute

In the Gemini Developer API's paid Standard tier, Flash TTS audio output costs $9 per million tokens and Flash-Lite TTS costs $6. Audio is billed at 25 tokens per second, or 1,500 tokens per minute. The following calculations cover audio output only; the input script and instructions cost extra. Official pricing

Gemini TTS prices change in January 2027
Model and periodInput text
per million tokens
Audio output
per million tokens
Audio output
calculated per minute
Flash TTS
Through December 31, 2026
$0.50$9$0.0135
Flash-Lite TTS
Through December 31, 2026
$0.50$6$0.009
Flash TTS
From January 1, 2027
$1$18$0.027
Flash-Lite TTS
From January 1, 2027
$1$12$0.018

ConditionsGemini Developer API prices, checked September 30. Standard pricing in US dollars; excludes free allowances, Batch, Flex, Priority, and caching. Per-minute calculation: audio token price x 25 x 60 / 1,000,000.

OpenAI's dedicated TTS costs $0.60 for input and $12 for audio output

The official gpt-4o-mini-tts specifications list $0.60 per million input text tokens and $12 per million audio output tokens. For example, 100,000 input tokens plus one million audio output tokens would cost $0.06 + $12 = $12.06.

This is not a comparison at equal audio duration. Providers tokenize audio differently, so Google's 25-tokens-per-second conversion cannot be applied to OpenAI. The OpenAI pricing and model pages reviewed do not give an equivalent per-minute conversion, so we have not assigned gpt-4o-mini-tts a flat minute rate. A rigorous cost comparison requires generating the same script and measuring audio duration and actual usage.

A GPT-Live minute measures session time, not generated audio length

GPT-Live 1 costs $0.05 per voice session minute, billed per second. According to OpenAI's cost guide, this includes user speech, assistant speech, silence, and time spent waiting for the backend. Muting the microphone does not end the session. Backend reasoning models and tools are billed separately.

Suppose a 10-minute conversation includes three minutes of assistant speech. GPT-Live's session fee would be $0.50. Generating three minutes of audio with Gemini TTS at current prices would cost $0.0405 for Flash or $0.027 for Lite in audio output fees. The latter figures exclude speech recognition, response generation, input text, and conversation management. The difference cannot simply be treated as savings for an entire voice application.

At scale, the scheduled price change also matters. Audio output for 1,000 generated minutes costs $13.50 for Flash or $9 for Lite in 2026, rising to $27 and $18 in 2027. GPT-Live costs $50 for 1,000 session minutes, plus backend charges. Budgets need separate totals for generated audio duration and active session duration.

Measure conversational delay first, then compare quality and cost

For Japanese narration or character dialogue, the published results and direction features give a reason to try Flash TTS first. For large volumes of standard messages or announcements, compare the same script with the lower-output-cost Flash-Lite TTS and decide whether the quality difference matters. When adding read-aloud features to an OpenAI-based application, include gpt-4o-mini-tts as well.

For responsiveness, test the same short and long Japanese utterances and time first audio, substantive-answer onset, and stopping after an interruption. Check P90 as well as the median to identify occasional long waits. Once candidates meet the required speed on the intended device and network, compare pronunciation errors, naturalness, and costs including retries.

Sources and references

Prices and specifications checked September 30, 2026. Evaluation scores come from Google's September 2026 materials.