Grok releases Voice Transcribe 2.0 with improved accuracy
Get Alerts TSLA Hot Sheet
Join SI Premium – FREE
Investing.com -- Grok launched Voice Transcribe 2.0 on Friday, a speech-to-text model that delivers twice the accuracy of its predecessor at the same price point. The company said the new model ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
The updated transcription tool builds on the audio foundation model that powers Grok Voice, which currently handles tens of thousands of customer-support calls daily and transcribes millions of hours of video narration. The technology also operates voice agents in physical products, including the Grok assistant in Tesla (NASDAQ: TSLA) vehicles.
Grok Voice Transcribe 2.0 shows improvements in processing real-world audio conditions such as poor phone connections, multiple speakers, regional accents, and spoken credentials. The company tested the model against seven competitors across four categories: telephony audio from customer-support calls, conversations with Grok, spoken credentials, and short multilingual voice commands.
In telephony tests using 8 kHz audio, the new model achieved a 7.1% word error rate, compared to 10.6% for the previous version. For conversational audio, the error rate dropped from 8.7% to 3.3%. The model reduced errors in transcribing credentials such as phone numbers and email addresses from 7.2% to 3.2%.
The model supports dozens of languages with automatic detection and can follow mid-recording language switches. For short phrases in 19 languages, the word error rate decreased from 20.6% to 6.8%.
Atlassian Loom adopted the new model for transcribing all its video recordings. "We've always believed the best way to move work forward is to capture context once and let it flow everywhere," said Sanchan Saxena, SVP of Teamwork Collection at Atlassian.
Pricing remains unchanged from the first version, with batch transcription at $0.10 per hour of audio and streaming at $0.20 per hour. The service includes speaker diarization, word-level timestamps, and key term biasing at no additional cost. The model will become the default in the Speech-to-Text API in the coming weeks.
Serious News for Serious Traders! Try StreetInsider.com Premium Free!
You May Also Be Interested In
- Freedom Broker Downgrades Nucor (NUE) to Hold, 'Guides Conservative Earnings Yet Again'
- LNG tanker hit by suspected cyberattack in Mediterranean Sea
- Gauzy wins court approval for debt restructuring, SPD film output to resume
Create E-mail Alert Related Categories
InvestingRelated Entities
Tesla, Maynard Um, Mark Zuckerberg, ARKSign up for StreetInsider Free!
Receive full access to all new and archived articles, unlimited portfolio tracking, e-mail alerts, custom newswires and RSS feeds - and more!



Tweet
Share