Google Releases Gemini 3.8 Live for Real-Time Voice Agents That Execute Tasks During Conversations
Summary
Google releases Gemini 3.8 Live and Extended Thinking for developers to build real-time voice agents that execute tasks mid-conversation, adding asynchronous function calling, live visual context and support for more than 97 languages.
Key Points
- Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking in the Gemini API and Google AI Studio for real-time voice agents that can execute tasks during conversations.
- Gemini 3.8 Live supports 97+ languages, asynchronous function calling, live visual context and alphanumeric parsing; audio input costs $0.005 per minute and output costs $0.018 per minute.
- Gemini 3.5 Transcribe supports 85+ languages, records 4.0% streaming and 2.6% non-streaming word error rates, and transcribes files up to one hour with timestamps and speaker labels via the Interactions API.