North-Mini-Code-1.0

Updated
05.10.2026
Tools
Thinking
Embedding
Reasoning
Code

North-Mini-Code-1.0 is Cohere’s 30B-A3B MoE for agentic coding and terminal tasks: 256K context, Apache 2.0.

At a glance

  • License: Apache 2.0
  • Parameters: 30B total, 3B active per token
  • Context length: 256K tokens, 64K max output
  • Modalities: Text input, text output
  • Minimum hardware: 12 GB memory (smallest GGUF is 9.38 GB)

What is North-Mini-Code-1.0?

North-Mini-Code-1.0 is an open weights research release from Cohere and Cohere Labs, a 30B parameter Mixture-of-Experts model tuned for code generation, agentic software engineering and terminal tasks. Only 3B parameters are active per token: each token passes through 8 of the 128 experts, so the whole model has to fit in memory but the per-token compute stays small. Cohere published the weights on June 5, 2026 under Apache 2.0.

SpecificationNorth-Mini-Code-1.0
Total parameters30B total, 3B active per token
ArchitectureDecoder-only sparse Mixture-of-Experts, 128 experts, 8 active
AttentionSliding-window with RoPE, interleaved 3:1 with global attention without positional embeddings
Context window256K tokens
Max output64K tokens
ModalitiesText input, text output
Post-trainingTwo-stage SFT, then reinforcement learning with verifiable rewards (RLVR)
Release dateJune 5, 2026
LicenseApache 2.0

Each expert is an FFN block with SwiGLU activation, the router applies a sigmoid to the logits before top-k selection, and a single dense layer sits ahead of the sparse ones. Post-training ran in two cascaded stages, supervised fine-tuning followed by reinforcement learning with verifiable rewards, focused on agentic coding. The model supports interleaved thinking and, per Cohere, works best with it turned on: pass the generated thinking along to later agentic steps and chat turns. Cohere recommends temperature 1.0 and top_p 0.95 for generation.

What North-Mini-Code-1.0 is good at

Cohere publishes its launch comparison as a chart rather than a score table, so there are no vendor numbers to reprint here. The model card does say exactly what the model was measured on: SWE-Bench Verified, SWE-Bench Pro, Terminal-Bench v2 and Terminal-Bench Hard for agentic coding, plus SciCode and LiveCodeBench v6 for complex code generation outside tool use, each run with 3 seeds and averaged. That list is a good map of what the model is built for: multi-step software engineering with a terminal in the loop, not general chat.

Tool use is trained in and supported through standard chat templates with JSON schema tool descriptions. Cohere documents running the model in OpenCode against a local vLLM server with interleaved reasoning enabled, and you can try it in OpenCode or the hosted Hugging Face Space before downloading anything. Serving from the original weights needs vLLM main plus Cohere's melody library for response parsing.

North-Mini-Code-1.0 hardware requirements

The system requirement to check is memory: the whole 30B has to load even though only 3B are active per token. Cohere ships safetensors; the GGUF builds below, with their real file sizes, come from the community repo unsloth/North-Mini-Code-1.0-GGUF.

MemoryBuild to pickFile size
12 GBUD-IQ2_M9.86 GB
16 GBUD-IQ3_S12.77 GB
24 GBUD-Q4_K_XL19.25 GB
32 GBUD-Q5_K_XL23.00 GB
48 GB and upUD-Q8_K_XL33.21 GB

Neighbouring builds differ by a gigabyte or two, so when two builds both fit, take the larger one. A 9.38 GB UD-IQ1_M exists below this ladder for squeezed machines, but low-bit quants give up the most quality. If the format is new to you, start with what GGUF is.

How to run North-Mini-Code-1.0 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for North-Mini-Code-1.0 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For a same-shape comparison point, see Qwen3 Coder 30B-A3B, another 30B-A3B coding MoE, or browse every model you can run locally.

North-Mini-Code-1.0 license

North-Mini-Code-1.0 is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can run it in a commercial coding stack, fine-tune it, and ship products on top of it without a usage fee.

Get the weights from Hugging Face

huggingface-cli download CohereLabs/North-Mini-Code-1.0
from transformers import AutoModel
model = AutoModel.from_pretrained("CohereLabs/North-Mini-Code-1.0")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

North-Mini-Code-1.0 is CohereLabs' first model aimed at developers, a 30.5B-parameter Mixture-of-Experts model with about 3B active parameters built for agentic software engineering. It has native tool use and interleaved thinking, so it can plan, edit files, and run terminal commands as a coding agent. It's open-weight under the Apache 2.0 license and can run fully on your own hardware through Atomic Chat.

The full-precision model targets a single H100 80GB GPU in FP8, or 2x A100 40GB in BF16. Because it's a sparse MoE with only ~3B active parameters, quantized GGUF builds shrink the footprint a lot, ranging from roughly 9GB up to full BF16, so a consumer GPU can load a lower-bit quant. In Atomic Chat you choose the quant that fits your available VRAM.

Yes. It's released under the Apache 2.0 license, which allows commercial use, modification, and redistribution at no cost. Running it locally through Atomic Chat means there are no API fees or per-token charges. Cohere does ask that you also follow its Acceptable Use Policy.

Yes. Once you download the weights, inference runs entirely on your machine with no internet connection required. You can pull the files on a connected computer, move them to an air-gapped environment, and run the model there. That keeps proprietary source code on your own hardware, which is the main reason to use it locally in Atomic Chat.

Download the weights with huggingface-cli download CohereLabs/North-Mini-Code-1.0, or grab a quantized GGUF build for smaller hardware. For serving, vLLM and SGLang support its cohere2moe architecture today, while llama.cpp and Ollama need a build that includes the 128-expert support. The simplest path is Atomic Chat, where you pick a quant and load it with one click, then use it for chat or wire it into a coding agent.