Meta has introduced Muse Voice Transcribe, a new artificial intelligence model designed to listen, transcribe and distinguish between speakers in real time. It can handle streaming transcription, identify more than 20 speakers in a conversation and detect when a person has finished speaking while the audio is being captured. Meta says that the model can also switch between languages mid-conversation and fine-tuned to recognise specific names, places and other terms.
Meta stated that the model processes incoming audio in 80-millisecond slices and uses an adaptive delay system to determine when it has enough information to transcribe a word. The easier segments are processed almost instantly, while more complex phrases receive slightly more processing time.
The company trained the system using reinforcement learning, balancing transcription accuracy with lower latency. Meta claims this helps the model improve the trade-off between speed and accuracy. The same architecture also handles speaker changes and detects conversational endpoints, allowing transcription and speaker tracking to work within a single system.
In a blog post, Meta stated that the model was trained on over 70 languages with 25 undergoing extensive validation. These include Chinese, French, Hindi, Japanese, Spanish and Vietnamese.
Meta also showcased the model switching between English and Mandarin within the same sentence. The company also showcased its ability to handle long recordings involving multiple speakers without requiring manual cleanup.
Meta sees the technology as a building block for more personalised AI assistants capable of understanding accents, interruptions, overlapping speech and multilingual conversations.
Muse Voice Transcribe is rolling out immediately across several Meta products, including voice dictation in Meta AI and the company’s Muse Code tool. The model will also be available through Meta’s Model API and Meta AI for Mac.