Industry News

Grok Transcribe 2.0: Same Price, Far Fewer Errors

xAI released Grok Voice Transcribe 2.0 on 18/09/2026, ranking first for streaming accuracy at the same $0.10/hour batch price. Here is what that means for UK firms that record calls, meetings, or voice notes.

xAI has shipped a speech to text upgrade that keeps the price fixed and moves the accuracy needle where voice agents actually fail.

What happened

On 18/09/2026 SpaceXAI, the name xAI now trades under after the SpaceX acquisition, released Grok Voice Transcribe 2.0.

The company says the new model is twice as accurate as Grok Voice Transcribe 1.0 on its own production evaluations, at the same price.

Batch transcription stays at 0.10 US dollars per hour of audio.

Streaming stays at 0.20 US dollars per hour.

Speaker diarization, word level timestamps, confidence scores, and key term biasing are included rather than billed as extras.

On the public Artificial Analysis AA-WER Streaming board, Grok Voice Transcribe 2.0 ranks first among streaming models for final transcript accuracy at 2.7% word error rate.

The more important number for live agents is the first partial transcript, which fell from 18.3% word error rate on version 1.0 to 3.4% on version 2.0.

That is the text a voice agent gets while the caller is still speaking, and it is what decides whether the agent can interrupt, route, or answer in time.

Existing Speech to Text API integrations pick up the accuracy gain with no code change once 2.0 becomes the default.

Until then you opt in by selecting grok-voice-transcribe-2.0, and you can pin grok-voice-transcribe-1.0 if you need the older timing.

xAI says version 1.0 will be deprecated in the coming weeks.

Why this matters for UK businesses

Plenty of UK firms still rely on noisy phone audio: customer support queues, field engineer notes, Zoom and Teams recordings, and Loom style walkthroughs.

Clean studio audio is not the hard case.

Flaky lines, overlapping speakers, regional accents, and account codes read aloud are.

xAI tuned 2.0 on those conditions, including telephony at 8 kHz, spoken credentials such as phone numbers and email addresses, and short multilingual commands.

On its short phrase set across 19 languages, word error rate dropped from 20.6% to 6.8%.

Atlassian is already using the model for Loom video transcription, and has said the higher accuracy opens a path from a recorded action plan straight into Cursor for code changes.

For a British SME, the commercial point is simpler than the leaderboard.

You get a measurable accuracy step without a price rise, and a flat per hour rate that is easier to forecast than token billed audio.

If you already transcribe support calls for QA, compliance, or coaching, fewer credential and number errors means less manual cleanup.

If you are building a voice agent that must act on partial transcripts, the jump from roughly one error in five words to one in thirty is the difference between a demo and something you can put on a live line.

How we see it at Adevious AI

We would treat this as an A/B test, not an automatic flip.

xAI bought accuracy with a little latency.

Independent board numbers show the new model is about a third of a second slower to return its first partial and its final transcript than version 1.0.

For recorded meeting notes that is invisible.

For a barge in voice agent that needs to speak back in under a second, it is a real trade.

Pin version 1.0 on any latency sensitive path until you have timed 2.0 on your own audio.

Then run both models on a week of real calls from your UK numbers, and score the mistakes that actually cost money: wrong account codes, misheard postcodes, and missed language switches.

Vendor leaderboards are English weighted and agent talk heavy.

They do not prove how the model handles your Scottish engineer on a site radio, or your Indian support contractor on a softphone.

Also watch where the audio is processed.

The Speech to Text service currently runs in a single US region.

xAI cites SOC 2 Type II, BAAs, and EU data residency options, but the network hop still sits inside your latency and data transfer budget if callers sit in Britain.

Confirm retention settings and any UK GDPR processing terms before you pipe regulated call audio into a new default.

What to do this week

If you already call the Grok Speech to Text API, switch one non critical batch job to grok-voice-transcribe-2.0 and compare word error on spoken numbers and emails.

If you do not, pick one recorded channel you already keep, a Loom library, a support sample, or a Teams recording set, and price a week of batch transcription at 0.10 US dollars per hour.

Decide before the default flips whether you want 2.0, a pinned 1.0, or a different vendor for latency sensitive paths.

Bottom line

Grok Voice Transcribe 2.0 is not a flashy new product line.

It is a quieter kind of release: same API, same price, much better partial transcripts, and a deprecation clock on the older model.

For UK teams that live on phone and meeting audio, that is worth a controlled trial this week rather than a surprise when the default changes.

Sources

SpaceXAI / xAI, Introducing Grok Voice Transcribe 2.0, 18/09/2026: https://x.ai/news/grok-voice-transcribe-2

OrcaRouter analysis of Artificial Analysis AA-WER Streaming board, 18/09/2026: https://www.orcarouter.ai/blog/grok-voice-transcribe-2-0-release

The Next Web, coverage of Transcribe 2.0 pricing and SpaceXAI naming, 19/09/2026: https://thenextweb.com/news/spacexai-transcribe-2-ten-cents-production-audio-credentials

Acronyms

API: Application Programming Interface

BAA: Business Associate Agreement

GDPR: General Data Protection Regulation

SME: Small and Medium sized Enterprise

SOC 2: Service Organisation Control 2

UK: United Kingdom

WER: Word Error Rate