Alibaba's Qwen3 vision-language model with 30B total / 3B active parameters. Supports text and image inputs.