MuSE (Multimodal Speech Embedding) is not typically classified as a general-purpose foundation model. It’s a specialized model for aligning audio and text representations. True foundation models support broader tasks and modalities with high generalizability.











