NVIDIA's efficient hybrid MoE model combining Mamba-2 and Transformer layers, with 30B total and 3.5B active parameters. Reasoning can be toggled on or off, and it handles English, code and several European and Asian languages.
NVIDIA's efficient hybrid MoE model combining Mamba-2 and Transformer layers, with 30B total and 3.5B active parameters. Reasoning can be toggled on or off, and it handles English, code and several European and Asian languages.