SAN FRANCISCO — xAI has released Grok Voice Think Fast 2.0, its most capable speech-to-speech voice model yet, pairing sharper transcription and smoother conversation with a response time fast enough to feel genuinely live.
Announced on July 29, the model reasons while it speaks — listening, thinking, and talking at the same time — which lets it stay quick even as it works through complex, multi-step tasks in the background.
Faster, Sharper, and Benchmark-Leading
The headline number is speed. Grok Voice Think Fast 2.0 delivers audio in 0.70 seconds, down from 1.25 seconds on version 1.0 and well ahead of competing systems. On the Artificial Analysis Speech-to-Speech Quality Index it scored 82.9%, topping both its predecessor at 75.7% and rival models from other major labs.
Accuracy climbed too. In evaluations across thousands of short phrases in 24 languages, xAI reported a 1.5 to 2.0 times improvement in transcription accuracy over dedicated speech-to-text systems from Deepgram and ElevenLabs, with the gap widening to roughly 10 times in noisy, real-world conditions. The gains extend a rapid cadence of releases from the Grok team, which earlier this summer added 21 flagship voices and a voice agent builder.
Built for Real Conversations and Real Work
xAI trained the model to talk more like a person — shorter sentences, one question at a time, and less filler — while quietly guiding users through complex workflows. It also became far more efficient with reasoning, using roughly 60% fewer reasoning tokens than version 1.0, so tool calls often execute before the agent finishes its first sentence.





