Meta’s Muse Voice Transcribe supports 5 Indian languages and 70+ languages globally

0
111
Meta brings real time voice transcription to 5 Indian languages and 70+ languages
Meta brings real time voice transcription to 5 Indian languages and 70+ languages

Real time speech processing is getting a major upgrade with Meta’s new Muse Voice Transcribe model. Developed by Meta Superintelligence Labs, it is the company’s first real time audio perception model and supports more than 70 languages, including Hindi, Tamil, Telugu, Kannada and Malayalam.

Real time transcription with speaker separation

Muse Voice Transcribe converts speech into text as it happens. It can also identify individual speakers and detect when a person starts or stops speaking. These functions work together as the audio arrives, without requiring separate post processing.

The model can distinguish more than 20 speakers in a recording and handle audio longer than 1 hour. It processes sound in 80 millisecond chunks and decides when there is enough information to transcribe each word. It can delay harder words while committing simpler ones faster. Meta says reinforcement learning helps the model manage these delays while keeping transcription errors low.

Meta trained the model across more than 70 languages and extensively validated 25 languages for the initial release. Its supported languages include 5 major Indian languages: Hindi, Tamil, Telugu, Kannada and Malayalam.

Supports multilingual conversations

The model can also handle code switching, including language changes within the same sentence. Users do not need to manually change the language whenever speakers switch languages.

Muse Voice Transcribe also supports language, keyword and context biasing. This helps the model recognise words based on the audio and the wider conversation.

Availability and pricing

Meta has made Muse Voice Transcribe available through the Meta Model API, allowing developers to use it for speech transcription in their own applications. Meta AI for Mac and Muse Code already use the model for dictation.

The API costs $3 (roughly Rs. 300) per 1,000 audio minutes, which is about $0.18 (roughly Rs. 17) per hour.

Also read: Viksit Workforce for a Viksit Bharat

Do Follow: The Mainstream LinkedIn | The Mainstream Facebook | The Mainstream Youtube | The Mainstream Twitter

About us:

The Mainstream is a premier platform delivering the latest updates and informed perspectives across the technology business and cyber landscape. Built on research-driven, thought leadership and original intellectual property, The Mainstream also curates summits & conferences that convene decision makers to explore how technology reshapes industries and leadership. With a growing presence in India and globally across the Middle East, Africa, ASEAN, the USA, the UK and Australia, The Mainstream carries a vision to bring the latest happenings and insights to 8.2 billion people and to place technology at the centre of conversation for leaders navigating the future.