Grok Voice Transcribe 2.0 doubles accuracy at the same price
xAI's new speech-to-text model is twice as accurate as the version it replaces for the same money, and the company says it now leads 32 streaming models on accuracy.

xAI has released Grok Voice Transcribe 2.0, a speech-to-text model the company says is twice as accurate as its predecessor without a price rise. The model arrived on 18 September, built on the same audio foundation model that powers Grok Voice, the assistant xAI says handles tens of thousands of customer-support calls a day, transcribes millions of hours of video narration and runs voice agents inside physical products, including the Grok assistant in Tesla vehicles.
Where the accuracy actually improves
Transcription is easy on clean, single-speaker audio. xAI's argument for 2.0 is the difficult end of the market: crackling phone lines, people talking over one another, regional accents, and account codes or email addresses read aloud. The model was trained on live, noisy, multilingual audio gathered across varied environments and then refined with post-training. On xAI's short-phrase test set, the kind of clipped in-car command that gives a model almost no context for identifying the language, word error rate drops from 20.6 per cent to 6.8 per cent. The company says multilingual work is the biggest single improvement over version 1.0, and that on 8kHz telephony audio it leads every model it tested.
Same price, more plumbing
Pricing sits at $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming, unchanged from the previous model, and existing Speech-to-Text API integrations pick up the accuracy improvement without code changes. The feature list is aimed squarely at builders: word-level timestamps with confidence scores, speaker diarization at no extra cost, as many as eight independent channels, key-term biasing with up to 100 domain terms per request, automatic formatting for numbers, dates, currencies, phone numbers and email addresses, filler-word removal, and smart turn detection for voice agents.
Why the Loom example matters
The customer xAI puts at the centre of the announcement is Atlassian, which says it found 2.0 more accurate than its previous system for transcribing Loom screen recordings. Atlassian's Sanchan Saxena, senior vice president of its Teamwork Collection, describes recording an action plan in Loom, pushing the transcript into a coding agent and having the work done, which is a workflow pitch rather than a dictation pitch.
Our opinion
The doubling is the headline; the flat price is the actual news. Transcription has spent two years being commoditised, and when the strongest claim a company can make about a new model is twice the accuracy for the same money, it is telling you where the fight is: not capability, but the cost of processing every hour of audio a business produces. The features give the game away too. Diarization at no extra charge, timestamps, channel separation and key-term biasing exist to push transcripts into automation, and the Atlassian quote confirms xAI is selling plumbing for agents rather than a better notes app. Our caution is about provenance. The ranking comes from xAI's own reading of a public leaderboard, and the customer praising the model in the launch post runs a product already wired into Grok. The numbers are credible. They are simply not independent yet.