Nemotron 3 nano: Open, efficient mixture-of-experts hybrid mamba-transformer model for agentic reasoning
A Blakeman, A Grattafiori, A Basant, A Gupta… - arXiv preprint arXiv …, 2025 - arxiv.org
arXiv preprint arXiv:2512.20848, 2025•arxiv.org
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer
language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more
than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and
large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than
our previous generation Nemotron 2 Nano while activating less than half of the parameters
per forward pass. It achieves up to 3.3 x higher inference throughput than similarly-sized …
language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more
than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and
large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than
our previous generation Nemotron 2 Nano while activating less than half of the parameters
per forward pass. It achieves up to 3.3 x higher inference throughput than similarly-sized …
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than our previous generation Nemotron 2 Nano while activating less than half of the parameters per forward pass. It achieves up to 3.3x higher inference throughput than similarly-sized open models like GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507, while also being more accurate on popular benchmarks. Nemotron 3 Nano demonstrates enhanced agentic, reasoning, and chat abilities and supports context lengths up to 1M tokens. We release both our pretrained Nemotron 3 Nano 30B-A3B Base and post-trained Nemotron 3 Nano 30B-A3B checkpoints on Hugging Face.
arxiv.org