Products & Services · · Team LATENT
Gemini 3.8 Live launches, keeping conversation going while work runs
The two models differ in handling complex requests, acknowledgments and interruptions. We explain Japanese support, the cost of a whole conversation and availability across apps.
Can AI reduce the silence while it researches a spoken request? Google's new Live models handle conversation and work at the same time. We consider how conversational fluency and task completion differ, and when to choose each model.
Key takeaways
- Both can continue talking while external services process requests; Extended Thinking also handles complex reasoning in the background.
- In independent tests, Extended Thinking completed more multi-step customer-service tasks, while standard Live performed better on acknowledgments and interruptions.
- Both support 97 languages, including Japanese, but availability varies by app and subscription.
What happened
- Google announced Gemini 3.8 Live and Extended Thinking on September 15, at 2 a.m. September 16 in Japan.
- Both are available in the Gemini API and Google AI Studio. Consumer apps are rolling out gradually; Gemini Enterprise receives a limited preview.
- Audio API rates for both models are $3 per million input tokens and $12 per million output tokens.
Primary sources: Google announcement, Developer announcement, Pricing
Keep talking while research and booking operations run
Google announced two models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both accept and produce audio, and can continue talking to users while waiting for external queries or processing.
Standard Live is designed for fast exchanges and simple instructions; Extended Thinking targets multi-step work. Google's developer documentation lists voice search and language practice for the former, and travel planning or technical support that draws on multiple sources for the latter.
For example, an assistant could discuss a trip while checking flights and hotels, then compare options when the results arrive. Extended Thinking can reason and process tasks in the background while giving spoken progress updates. Reducing silent waits distinguishes this from simply transcribing a question before answering it.
The model cannot, by itself, operate every service or booking site. Apps must provide integrations and action tools for real reservations or internal-data queries. Announcement demonstrations should be distinguished from features immediately available in consumer apps.
Extended Thinking scores higher on complex requests
In Artificial Analysis's voice-model evaluation, Live scores 76.0 and Extended Thinking 82.6 on the Speech to Speech Index. Both exceed the 71.5 score of the previous Gemini 3.1 Flash Live at high reasoning.
The index combines four results, including solving spoken questions, completing customer-service tasks and human conversational preferences. A score of 82.6 does not mean success on 82.6% of everyday requests.
The largest gap appears on tau-Voice, which simulates customer-service tasks such as changing flights or returning products. Task completion is 30.1% for Live and 68.6% for Extended Thinking. The test checks whether the required operations follow policy and leave data in the correct final state, not merely whether the model answers.
On Full Duplex Bench, covering acknowledgments, turn-taking and interruptions, Live scores 96.1% and Extended Thinking 91.9%. The model stronger on complex work does not win every aspect of conversation. Extended Thinking was measured at high reasoning.
Higher is better in each metric. Extended Thinking was measured at high reasoning. Which model leads depends on the metric.Note 1
Composite indexSpeech to Speech Index
3.8 Live76.0
Extended Thinking82.6
Customer-service task completiontau-Voice
3.8 Live30.1%
Extended Thinking68.6%
Acknowledgments and interruptionsAggregate of four Full Duplex Bench measures
3.8 Live96.1%
Extended Thinking91.9%
Extended Thinking completes more complex requests, while standard Live scores higher on conversational behavior. These are not Japanese-only evaluations.
Independent measurementsBased on Artificial Analysis evaluations, checked in the September 15 archive. The composite index is not a success rate. Test definitions follow the evaluator's documentation.
Conversation in 97 languages, including Japanese
Google says users can switch languages during a conversation. The Live API supported-language list includes Japanese among 97 languages.
Both models accept text, images and video as well as audio, enabling conversations grounded in the surroundings, such as questions about what a camera sees. Google demonstrates playing chess from a view of the board and creating interfaces from hand-drawn sketches and spoken instructions.
Language support is not the same as accuracy on tasks in that language. The evaluations above do not establish performance on Japanese names and addresses, honorifics or reservations involving proper nouns. Japanese-language services need tests using their actual vocabulary and background noise.
Identical API rates do not mean identical conversation costs
Google's pricing table gives Live and Extended Thinking identical API rates. Standard paid-tier prices per million tokens are $3 for audio input and $12 for audio output. Text input costs $0.75, and text output, including reasoning tokens, costs $4.50.
The developer announcement also gives approximate audio rates of $0.005 per input minute and $0.018 per output minute. These convert the amount of input and output audio separately into price. Adding them and multiplying by call duration does not give the full bill.
One reason is conversation-history billing. At each turn, the Live API also processes the conversation history currently retained as input. Audio kept in that history is billed again as input, so the cost of each response can grow as a conversation lengthens. Enabling audio transcription also incurs separate text-output charges.
Both models always enable Proactive Audio, under which the model decides whether incoming audio warrants a response. Input charges apply while the Live API listens, even when it does not respond. Output is charged when it responds.
Total cost varies with reasoning effort, how much the AI speaks and how much history is retained. At identical unit prices, Extended Thinking can still cost more if it processes more. Compare the cost of completing the intended task, not just an indicative per-minute rate.
US dollars per million tokens. Standard paid-tier rates shared by Live and Extended Thinking.
AudioThe user speaks and the AI responds with audio
Input$3.00
Output$12.00
TextOutput includes reasoning tokens
Input$0.75
Output$4.50
Charges depend on usage. Reprocessing conversation history and transcription also affect cost, so call duration alone is insufficient.
Official pricingBased on Gemini API pricing and Live API billing rules. Image and video input and search charges are not included in this table.
Availability differs between consumer and enterprise services
According to the announcement-day guidance, developers can use both models through the Gemini API and Google AI Studio. Gemini Enterprise receives a limited preview, not general availability for every enterprise customer.
For consumers, standard Live begins rolling out in Search Live and Extended Thinking in Gemini Live. Both are gradual rollouts, not a claim that everyone switched to the new models on announcement day.
In Google Workspace, Docs is available to Google AI Pro and Ultra users, while Gmail and Keep target paid Google AI subscribers. Availability for Workspace business customers and Gemini Enterprise for Customer Experience is planned later.
Even under the same Live name, apps built by developers through the API and Google's own apps differ in available actions and terms. Japanese API support alone does not mean every account in Japan can use the new features.
A spoken reply does not necessarily mean the background task is finished
When conversation and processing run in parallel, the end of a reply must be distinguished from the end of a task. Google's Extended Thinking specification says reasoning or external processing may continue after a spoken turn ends.
An acknowledgment of a booking request may arrive while availability is still being checked. Users need not only quick responses but also clarity about what has progressed and what is confirmed. Keeping conversation flowing and communicating completion accurately must go together.
The two models broaden the work that can be requested by voice, but different settings have different needs. For language practice, compare pacing and interruptions; for reservation services, compare checking conditions and finishing procedures. These are better guides to fit than an overall ranking alone.
Sources and references
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking Google / September 15, 2026 / Archived version
- Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe Google / September 15, 2026 / Archived version
- Speech to Speech AI Model & Provider Leaderboard Artificial Analysis / September 15 measurements / Archived version
- Thinking in the Live API Google / Parallel conversation and processing / Archived version
- Gemini 3.8 Live Google / Model specifications / Archived version
- Gemini 3.8 Live Extended Thinking Google / Model specifications / Archived version
- Live API capabilities Google / Languages and features / Archived version
- Gemini Developer API pricing Google / API pricing / Archived version
- Live API best practices Google / Billing for history and audio / Archived version
Update history
- Published.
- Corrected the scope of the tests cited in the key takeaways and added the conditions under which audio input is billed even without a response.