NVIDIA's hybrid Mamba-Transformer MoE model (31B total, 3B active) with a multi-token prediction speculative decoding head for low-latency serving. Tuned for high-throughput reasoning and agentic workloads.
NVIDIA's hybrid Mamba-Transformer MoE model (31B total, 3B active) with a multi-token prediction speculative decoding head for low-latency serving. Tuned for high-throughput reasoning and agentic workloads.