Human speech becomes the next frontier of artificial intelligence.
San Francisco, United States.
Elon Musk’s artificial intelligence company xAI has launched Grok Voice Transcribe 2.0, a speech recognition model designed to convert recorded audio and live conversations into text with improved accuracy. Announced on September 18, the technology reportedly delivers twice the transcription accuracy of its predecessor while maintaining the same pricing structure. The release targets businesses, software developers and organizations seeking reliable transcription in environments where background noise, multiple speakers and language differences complicate automated communication. Its arrival intensifies competition in an expanding market where spoken language is becoming an essential interface between humans and intelligent systems.
The model builds on the audio technology underlying Grok Voice, which xAI says already processes thousands of customer service calls daily and supports voice assistants in Tesla vehicles. Its development involved recordings captured in noisy, multilingual environments, followed by additional training intended to improve performance under realistic conditions. According to the company, the system ranks first for accuracy among 32 streaming transcription models evaluated by Artificial Analysis. These benchmark results offer a comparative reference, although performance may vary according to recording quality, language and application requirements.

Multilingual recognition represents one of the principal advances. Grok Voice Transcribe 2.0 automatically identifies spoken languages and can follow conversations when speakers switch between them. In company testing involving short voice commands across 19 languages, its word error rate declined from 20.6% to 6.8% compared with the previous generation. The improvement is particularly relevant for international customer service operations, multilingual meetings and applications requiring accurate interpretation of brief instructions.
The technological architecture also introduces capabilities for more complex workflows. Developers can process prerecorded files or stream audio in real time, obtaining transcripts with individual word timestamps, confidence indicators and speaker identification. The system supports up to eight independent audio channels and allows users to specify as many as 100 specialized terms to improve recognition of technical vocabulary. Additional features include formatting spoken numbers, currencies and telephone information, removing filler words and detecting when a speaker has finished talking.

Pricing forms an important component of xAI’s commercial strategy. Batch transcription costs $0.10 per hour of audio, while real-time processing is priced at $0.20 per hour, with several advanced functions included. The company intends to make the new model the default option in its speech recognition interface and gradually discontinue the previous version. Existing integrations can adopt the improved model without extensive software modifications, potentially reducing migration costs for developers.
Commercial adoption has already begun. Atlassian selected the technology for Loom after reporting improved transcription accuracy compared with its previous solution. The integration allows recorded explanations and screen demonstrations to become structured text that can subsequently support development workflows through tools such as Cursor. This application illustrates how transcription is evolving from a documentation function into an intermediary between human instructions and automated software execution.

The development nevertheless introduces questions about information security and reliability. Conversations may contain confidential business information, personal identifiers or sensitive customer data, making responsible handling of audio recordings essential. Furthermore, greater benchmark accuracy does not eliminate transcription errors, particularly when technical instructions or consequential decisions depend on individual words.
Grok Voice Transcribe 2.0 represents another step toward making spoken language a practical interface for artificial intelligence. Its broader significance lies in connecting conversations with digital processes, potentially allowing human instructions to move directly into automated workflows while preserving the need for verification and human accountability.
Phoenix24: claridad en la zona gris. / Phoenix24: clarity in the grey zone.