Sep 21 edition/Reporting & analysis
ModelsAgentsBusinessMultimodal

ModelsArchitectures & capability

GPT-Live-1 narrowly tops Grok Voice on speech benchmark, but Gemini leads the table

Artificial Analysis scores put GPT-Live-1 slightly ahead of Grok Voice Think Fast 2.0, but Gemini 3.8 Live Extended Thinking ranks higher. For builders, the bigger shift is full-duplex voice architectures that pair live conversation with backend reasoning.

THE CORE IDEAS4 TAKEAWAYS
01

The benchmark result is a close comparison, not a decisive win: GPT-Live-1 is reported at 81.5 versus Grok Voice Think Fast 2.0 at 81.3, while Gemini 3.8 Live Extended Thinking leads at 82.6. [7] [13]

02

OpenAI positions GPT-Live-1 as a full-duplex voice layer that can keep conversation flowing while delegating harder reasoning, tools, or business logic to backend systems. [4] [6]

03

For production teams, the practical evaluation is broader than perceived humanness: latency, interruption handling, task completion, human preference, escalation, and cost all shape whether a voice agent works in real workflows. [7] [9]

04

A separate research warning tempers naturalness claims: realtime voice systems may respond to spoken words while missing emotional cues such as distress, fear, or sarcasm in delivery. [12]

WHY IT MATTERS

Evidence shows speech-to-speech models are now being compared on task and interaction quality, not only audio realism. The implication is that buyers should test voice agents against real workflows before choosing on demos or headline scores.

Executive brief

The “GPT‑Live‑1 beat Grok Voice” headline is technically true but incomplete: Artificial Analysis lists GPT‑Live‑1 (Astra, medium) at 81.5 vs Grok Voice Think Fast 2.0 at 81.3, while Gemini 3.8 Live Extended Thinking leads both at 82.6 on the same speech-to-speech index. The real story is not a knockout; it is the arrival of production-grade, full-duplex voice agents that separate a live conversational layer from deeper backend reasoning. For builders, the shift is architectural: lower-latency interaction, asynchronous delegation, and explicit evaluation of turn-taking, task success, and human preference.

What changed and event timeline

  1. OpenAI introduces GPT‑Live

    OpenAI said GPT‑Live‑1 and GPT‑Live‑1 mini would roll out to ChatGPT Voice, built on full-duplex listening/speaking and backend delegation for harder work.

  2. Grok Voice Think Fast 2.0 becomes current Grok voice model

    xAI docs state grok-voice-latest points to grok-voice-think-fast-2.0, positioning Grok as the low-latency rival in this comparison.

  3. GPT‑Live‑1 reaches the OpenAI API

    OpenAI launched GPT‑Live‑1 for developers at $0.05/minute for the front-end voice layer, with WebRTC, WebSocket, telephony/SIP, and backend delegation.

  4. Gemini 3.8 overtakes the headline fight

    Artificial Analysis lists Gemini 3.8 Live Extended Thinking at 82.6, above GPT‑Live‑1 Astra at 81.5 and Grok Voice Think Fast 2.0 at 81.3.

  5. Reddit commentary amplifies the GPT‑Live vs Grok framing

    The linked post and video emphasize full duplex, interruption handling, and a narrow GPT‑Live edge over Grok, while the video later notes Gemini’s higher score.

Capabilities and access

Exact OpenAI model: GPT‑Live‑1; evaluated configuration: GPT‑Live‑1 (Astra, medium). OpenAI says GPT‑Live handles spoken conversation, listens while speaking, and delegates backend work to OpenAI Responses models or a developer-controlled agent/service. Sessions can use WebRTC, WebSockets, telephony/SIP, or sideband server control. API pricing is duration-based; OpenAI’s launch post states $0.05/minute for the front-end voice layer, with backend model/tool usage billed separately.

Read the full section

Exact OpenAI model: GPT‑Live‑1; evaluated configuration: GPT‑Live‑1 (Astra, medium). OpenAI says GPT‑Live handles spoken conversation, listens while speaking, and delegates backend work to OpenAI Responses models or a developer-controlled agent/service. Sessions can use WebRTC, WebSockets, telephony/SIP, or sideband server control. API pricing is duration-based; OpenAI’s launch post states $0.05/minute for the front-end voice layer, with backend model/tool usage billed separately. Getting started with GPT‑Live, OpenAI API launch

Technical analysis for researchers and developers

GPT‑Live is a two-part system: a native speech front end for conversation control and a backend for reasoning/tool work. Artificial Analysis’ index currently weights Speech Reasoning, Agentic Performance, Arena Preference, and Task Success equally; earlier versions weighted Full Duplex Bench directly. Reproducibility depends on public benchmark harnesses plus opaque proprietary model endpoints.

Read the full section

GPT‑Live is a two-part system: a native speech front end for conversation control and a backend for reasoning/tool work. That design removes the classic STT→LLM→TTS cascade from the critical interaction loop, but it does not remove application responsibilities: permissions, confirmations, function execution, state saving, and cancellation policy remain developer-owned. Artificial Analysis’ index currently weights Speech Reasoning, Agentic Performance, Arena Preference, and Task Success equally; earlier versions weighted Full Duplex Bench directly. Reproducibility depends on public benchmark harnesses plus opaque proprietary model endpoints. Methodology

Claims and evidence

  • Independent: Artificial Analysis lists GPT‑Live‑1 Astra at 81.5, Grok Voice Think Fast 2.0 at 81.3, and Gemini 3.8 Live Extended Thinking at 82.6. Speech to Speech Models
  • Vendor-reported: OpenAI says GPT‑Live‑1 listens and speaks simultaneously and can delegate reasoning/tool calls to backends such as Astra or third-party systems. OpenAI API launch
  • Vendor/customer-reported: Speak says tutor-specific prompting cut GPT‑Live‑1 interruption rates from 27.6% to 13.6%, and delivered 476/477 authored lesson lines in 27 evaluation sessions. Speak blog
Read the full section
  • Independent: Artificial Analysis lists GPT‑Live‑1 Astra at 81.5, Grok Voice Think Fast 2.0 at 81.3, and Gemini 3.8 Live Extended Thinking at 82.6. Speech to Speech Models
  • Vendor-reported: OpenAI says GPT‑Live‑1 listens and speaks simultaneously and can delegate reasoning/tool calls to backends such as Astra or third-party systems. OpenAI API launch
  • Vendor/customer-reported: Speak says tutor-specific prompting cut GPT‑Live‑1 interruption rates from 27.6% to 13.6%, and delivered 476/477 authored lesson lines in 27 evaluation sessions. Speak blog
  • Commentary/transcript: The video’s core argument is that GPT‑Live‑1 feels more human because of full duplex, but Grok remains faster and Gemini leads overall. 00:24, 03:06, 07:10

Context and prior work

This is part of a broader move from turn-based voice bots to continuous-time interaction. Artificial Analysis’ June index formalized speech-to-speech evaluation across reasoning, conversational dynamics, and agentic customer-service tasks. Realtime-Venus documents a shared timeline plus dual-loop runtime; “Never Stop Thinking” reports latency cuts from an interrupt-and-resume orchestrator.

Read the full section

This is part of a broader move from turn-based voice bots to continuous-time interaction. Artificial Analysis’ June index formalized speech-to-speech evaluation across reasoning, conversational dynamics, and agentic customer-service tasks. Research prototypes are converging on similar designs: continuous thinking while listening/speaking, asynchronous delegation, and foreground/background loops. Realtime-Venus documents a shared timeline plus dual-loop runtime; “Never Stop Thinking” reports latency cuts from an interrupt-and-resume orchestrator. Artificial Analysis announcement, Realtime‑Venus, Never Stop Thinking

Limitations, safety and contested findings

The “OpenAI beat Grok” claim is narrow: GPT‑Live‑1 leads Grok by 0.2 on one composite score, while Grok is faster on time-to-first-audio and stronger on some listed submetrics; Gemini leads the current overall table. Separately, voice naturalness is not emotional understanding: a July 2026 paper found leading realtime systems often acted on words while ignoring distress, fear, or sarcasm in vocal delivery.

Read the full section

The “OpenAI beat Grok” claim is narrow: GPT‑Live‑1 leads Grok by 0.2 on one composite score, while Grok is faster on time-to-first-audio and stronger on some listed submetrics; Gemini leads the current overall table. Separately, voice naturalness is not emotional understanding: a July 2026 paper found leading realtime systems often acted on words while ignoring distress, fear, or sarcasm in vocal delivery. Artificial Analysis, Real-Time Voice AI Hears but Does Not Listen

Business and practitioner implications

Voice-agent procurement should move from “sounds human” demos to scored task workflows: interruption handling, latency, task completion, escalation quality, and cost per live minute. GPT‑Live‑1 is attractive where natural turn-taking and backend flexibility matter; Grok remains relevant where first-audio latency is decisive; Gemini currently tops the composite index.

Read the full section

Voice-agent procurement should move from “sounds human” demos to scored task workflows: interruption handling, latency, task completion, escalation quality, and cost per live minute. GPT‑Live‑1 is attractive where natural turn-taking and backend flexibility matter; Grok remains relevant where first-audio latency is decisive; Gemini currently tops the composite index. Production teams should prototype against real calls, not scripts, and keep tool execution, approvals, monitoring, and human handoff outside the voice model.

Sources

Read the full section
FOLLOW THE EVIDENCE

The source trail.

Sources (13)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief