What is North-Mini-Code-1.0?
North-Mini-Code-1.0 is an open weights research release from Cohere and Cohere Labs, a 30B parameter Mixture-of-Experts model tuned for code generation, agentic software engineering and terminal tasks. Only 3B parameters are active per token: each token passes through 8 of the 128 experts, so the whole model has to fit in memory but the per-token compute stays small. Cohere published the weights on June 5, 2026 under Apache 2.0.
| Specification | North-Mini-Code-1.0 |
|---|---|
| Total parameters | 30B total, 3B active per token |
| Architecture | Decoder-only sparse Mixture-of-Experts, 128 experts, 8 active |
| Attention | Sliding-window with RoPE, interleaved 3:1 with global attention without positional embeddings |
| Context window | 256K tokens |
| Max output | 64K tokens |
| Modalities | Text input, text output |
| Post-training | Two-stage SFT, then reinforcement learning with verifiable rewards (RLVR) |
| Release date | June 5, 2026 |
| License | Apache 2.0 |
Each expert is an FFN block with SwiGLU activation, the router applies a sigmoid to the logits before top-k selection, and a single dense layer sits ahead of the sparse ones. Post-training ran in two cascaded stages, supervised fine-tuning followed by reinforcement learning with verifiable rewards, focused on agentic coding. The model supports interleaved thinking and, per Cohere, works best with it turned on: pass the generated thinking along to later agentic steps and chat turns. Cohere recommends temperature 1.0 and top_p 0.95 for generation.
What North-Mini-Code-1.0 is good at
Cohere publishes its launch comparison as a chart rather than a score table, so there are no vendor numbers to reprint here. The model card does say exactly what the model was measured on: SWE-Bench Verified, SWE-Bench Pro, Terminal-Bench v2 and Terminal-Bench Hard for agentic coding, plus SciCode and LiveCodeBench v6 for complex code generation outside tool use, each run with 3 seeds and averaged. That list is a good map of what the model is built for: multi-step software engineering with a terminal in the loop, not general chat.
Tool use is trained in and supported through standard chat templates with JSON schema tool descriptions. Cohere documents running the model in OpenCode against a local vLLM server with interleaved reasoning enabled, and you can try it in OpenCode or the hosted Hugging Face Space before downloading anything. Serving from the original weights needs vLLM main plus Cohere's melody library for response parsing.
North-Mini-Code-1.0 hardware requirements
The system requirement to check is memory: the whole 30B has to load even though only 3B are active per token. Cohere ships safetensors; the GGUF builds below, with their real file sizes, come from the community repo unsloth/North-Mini-Code-1.0-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 12 GB | UD-IQ2_M | 9.86 GB |
| 16 GB | UD-IQ3_S | 12.77 GB |
| 24 GB | UD-Q4_K_XL | 19.25 GB |
| 32 GB | UD-Q5_K_XL | 23.00 GB |
| 48 GB and up | UD-Q8_K_XL | 33.21 GB |
Neighbouring builds differ by a gigabyte or two, so when two builds both fit, take the larger one. A 9.38 GB UD-IQ1_M exists below this ladder for squeezed machines, but low-bit quants give up the most quality. If the format is new to you, start with what GGUF is.
How to run North-Mini-Code-1.0 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for North-Mini-Code-1.0 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For a same-shape comparison point, see Qwen3 Coder 30B-A3B, another 30B-A3B coding MoE, or browse every model you can run locally.
North-Mini-Code-1.0 license
North-Mini-Code-1.0 is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can run it in a commercial coding stack, fine-tune it, and ship products on top of it without a usage fee.