Recent advancements in artificial intelligence (AI) are significantly enhancing the efficiency and precision of voice transcription. On Tuesday, September 1, tech leader Meta unveiled a groundbreaking speech-to-text model known as Muse Voice Transcribe, capable of transcribing dialogues in real-time. This innovative model is also the first of its kind developed by Meta Superintelligence Labs (MSL) to offer real-time audio perception.
According to Meta, Muse Voice Transcribe supports more than 70 languages, which include five prominent Indian languages: Hindi, Tamil, Telugu, Malayalam, and Kannada. The company asserts that this model is adept at managing code-switching—where speakers interchange languages during conversations—without requiring separate models or additional processing steps.
In the Indian context, where users frequently alternate between English and their native languages, Muse Voice Transcribe could serve as a valuable tool, enhancing the practicality of real-time transcription. The model generates text dynamically as an individual speaks, eliminating the need to wait until an entire recording is completed. Furthermore, it can differentiate between speakers in recordings containing over 20 voices and can handle audio lasting more than an hour, as reported by Meta.
Meta emphasized that these functionalities are integrated within a single model without necessitating post-processing. Muse Voice Transcribe has been trained on over 70 languages, with 25 of these languages validated at the time of its launch. The company also noted that the model achieved the top position on the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026.
A notable characteristic of Muse Voice Transcribe is its ability to balance transcription speed and accuracy. Rather than adhering to a fixed duration for listening before generating each word, the model evaluates this timing on a word-by-word basis. This allows it to quickly respond to clear speech while dedicating more time to words that are more challenging to decipher.
The Muse Voice Transcribe model is now accessible via Meta’s Model API. Meta mentioned that the model is already being utilized for dictation in Meta AI for Mac and Muse Code. The pricing for the API is set at $3 for every 1,000 audio minutes, translating to approximately $0.18 per hour.
This launch coincides with the growing significance of speech-based AI systems for diverse applications, including transcription, dictation, coding, and voice assistance. By accommodating code-switching and a broader array of regional languages, these systems can better understand communication patterns in multilingual markets such as India. Additionally, Meta has shared further technical details about Muse Voice Transcribe in a research blog released alongside the announcement.




















