Skip to main content
Breaking News

Google unveils new Gemini 3.8 Live models to power real-time voice AI

Google has launched two new artificial intelligence models — Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — designed to strengthen real-time

2 min read5 views
Google unveils new Gemini 3.8 Live models to power real-time voice AI
Sharefin

The tech giant has rolled out two upgraded AI systems built for instant reasoning and natural voice conversations, aiming to help businesses build reliable voice assistants while enhancing speech features across Google's own consumer products.

Google has launched two new artificial intelligence models — Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — designed to strengthen real-time reasoning, support enterprise-grade voice assistants, and make spoken interactions with AI feel more human-like.

The company says the releases give developers and businesses the essential building blocks to create dependable, ready-to-deploy voice applications. Beyond enterprise use, Google noted that the underlying technology will also improve voice-based features within its own ecosystem, including the Gemini app, Google Workspace, and Search, allowing users to work through complex, multi-step tasks using only their voice.

Two models built for different needs

Google has split the launch into two distinct versions, each suited to different operational demands. Gemini 3.8 Live is designed for large-scale use at a lower cost, combining fluid voice conversation with the ability to understand visual context in real time. Gemini 3.8 Live Extended Thinking, on the other hand, is built for more demanding tasks that require deeper reasoning and the ability to work through problems in sequential steps.

According to Google, both models performed strongly on major industry benchmarks for voice and audio quality, while remaining more cost-efficient than comparable models from competitors. The Extended Thinking version reportedly ranked first globally on the Artificial Analysis Speech to Speech Quality Index, scoring 82.6. It also topped rankings for handling complex, task-based conversations, posting a 68.6 percent score on the tau-Voice benchmark and 35.1 percent on a banking-specific version of the same test developed by Sierra. Google further stated that the model achieved a 97.7 percent score on the Big Bench Audio benchmark.

Multilingual capabilities and seamless background tasks

The standard Gemini 3.8 Live model can process live video input almost instantly, using visual cues to generate more accurate and context-aware spoken responses. It is capable of recognising and switching between 97 languages mid-conversation without interruption. The model can also activate connected tools and run background processes — such as API calls — without pausing the conversation, allowing it to confirm a user's request verbally while completing related digital tasks in the background.

For more complex queries, the Extended Thinking model can reason through a problem while continuing to speak. Google explained that to avoid awkward silences during heavier computational tasks, the system uses natural conversational fillers — such as saying "Let me check that" — and offers real-time verbal updates as it works through multi-step processes.