Alphabet Inc. (NASDAQ:GOOG) has introduced two new Gemini voice models and highlighted a benchmark result that puts its Extended Thinking model ahead of competing voice agents from OpenAI and xAI, strengthening its position in the real-time voice AI era. The benchmark lead is real and measurable. However, voice AI benchmark rankings have shifted repeatedly over the past year, with rival labs releasing competing models only weeks apart and each reaching the top on different preferred benchmarks. Against that backdrop, the Extended Thinking model’s 1.1-point lead today may reveal more about the rapid pace of competition than it does about any lasting technological advantage for Google.

Google Unveils New Gemini Live Models and Claims a Benchmark Win
On September 15, Google launched Gemini 3.8 Live alongside Gemini 3.8 Live Extended Thinking. The two new models are designed to support more reliable and production-ready voice agents. The Extended Thinking model posted an 82.6 score on the Artificial Analysis’ Speech to Speech Quality Index, narrowly beating OpenAI’s GPT-Live-1 Astra at 81.5 and xAI’s Grok Voice Think Fast 2.0 at 81.3. The Extended Thinking model is designed to reason and speak at the same time, allowing it to use natural verbal cues while handling complex, multi-step tasks. The company is also partnering with Salesforce, Lumeris, and Genspark to develop enterprise use cases for the technology. Both models are available to developers through the Gemini API and Google AI Studio. Gemini 3.8 Live is rolling out in Search Live, while Gemini 3.8 Live Extended Thinking is rolling out in Gemini Live and selected Workspace experiences, including Docs, Gmail, and Keep. Both are also available in private preview through Gemini Enterprise.
Leadership That Doesn’t Stick
The Extended Thinking model’s current lead is roughly one point, but the benchmark results show how quickly leadership can change in voice AI. Several months earlier, xAI’s Grok Voice Agent claimed the top spot on a different benchmark with a score of 92.3%. The pace of competition has been equally fast. Independent comparisons have since described the race between OpenAI and xAI as “essentially a tie.” At the same time, OpenAI’s GPT-Realtime-2 has been described as the most enterprise-validated voice AI product on the market. It was used in early testing by companies including Zillow and Foundation Health, covering use cases in sectors such as real estate and healthcare.
According to our database, the number of hedge funds holding GOOGL rose from 265 at the end of the first quarter of 2026 to 275 at the end of Q2 2026. While the increase was modest, it reflects a continued growth in institutional interest.
Gemini’s benchmark performance is a genuine result backed by independent measurement. However, the frequent changes at the top of the voice AI leaderboard show how quickly competitive positions can shift. For Google, long-term enterprise adoption and customer retention may therefore be more important to sustaining its position than any single benchmark result.
READ NEXT: Zscaler’s AI Story Is Accelerating, Its Growth Guide Isn’t, and Google’s Traffic Numbers Look Fine, Its Search Economics Tell Another Story




