Xiaomi's MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layer Multi-Token Prediction (MTP) introduced in MiMo-V2-Flash.
Xiaomi's MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layer Multi-Token Prediction (MTP) introduced in MiMo-V2-Flash.