HologramHologramModels Hub
Index Sep 17, 2026
Models

facebook

mms-300m

Verified
Address
Identical bytes on
Status
Verified
Trending
#24
Downloads, 30 days
22k
Weights
1.3 GB
Sources
2
Revision
Manifest

MMS is Facebook AI's massive multilingual pretrained model for speech ("MMS"). It is pretrained in with Wav2Vec2's self-supervised training objective on about 500,000 hours of speech data in over 1,400 languages.

At a glance

Architecture
Wav2vec2
Library
transformers
License
CC BY NC 4.0Non commercial
Languages
ab, af, ak, am, ar, as, av, ay +150
Released
May 2023
Updated
Jun 2023
Likes
554
Downloads, all time
3,570,168

Architecture

Layers
24
Hidden size
1,024
Attention
16 heads
Vocabulary
32

Family

Models built on mms-300m.

Run it

Loads with Transformers AutoModelForPreTraining and AutoProcessor, pinned to the indexed revision.

from transformers import AutoModelForPreTraining, AutoProcessor

model = AutoModelForPreTraining.from_pretrained("facebook/mms-300m", revision="4ee317ce793c53dbc041fc4376c7558292dd38dc")
processor = AutoProcessor.from_pretrained("facebook/mms-300m", revision="4ee317ce793c53dbc041fc4376c7558292dd38dc")

Papers

Datasets

Spaces

Used in 25 Spaces.

Read the full model card

Massively Multilingual Speech (MMS) - 300m

Facebook's MMS counting 300m parameters.

MMS is Facebook AI's massive multilingual pretrained model for speech ("MMS"). It is pretrained in with Wav2Vec2's self-supervised training objective on about 500,000 hours of speech data in over 1,400 languages.

When using the model make sure that your speech input is sampled at 16kHz.

Note: This model should be fine-tuned on a downstream task, like Automatic Speech Recognition, Translation, or Classification. Check out the **How-to-fine section or this blog for more information about ASR.

Table Of Content

How to finetune

Coming soon...

Model details

  • Developed by: Vineel Pratap et al.

  • Model type: Multi-Lingual Automatic Speech Recognition model

  • Language(s): 1000+ languages

  • License: CC-BY-NC 4.0 license

  • Num parameters: 300 million

  • Cite as:

    @article{pratap2023mms,
      title={Scaling Speech Technology to 1,000+ Languages},
      author={Vineel Pratap and Andros Tjandra and Bowen Shi and Paden Tomasello and Arun Babu and Sayani Kundu and Ali Elkahky and Zhaoheng Ni and Apoorv Vyas and Maryam Fazel-Zarandi and Alexei Baevski and Yossi Adi and Xiaohui Zhang and Wei-Ning Hsu and Alexis Conneau and Michael Auli},
    journal={arXiv},
    year={2023}
    }
    

Additional Links

Derived on Sep 16, 2026 from Hugging Face at revision 4ee317ce, README.md , config.json .