Cohere Releases North Mini Code: A 30B Agentic Coding Model Optimized for Single-GPU Inference
Cohere has launched North Mini Code, a 30-billion parameter mixture-of-experts model designed for agentic software engineering. The Apache 2.0-licensed model achieves high output speeds while fitting on a single H100 GPU.

Key takeaways · 3
- 01
North Mini Code activates only 3 billion of its 30 billion parameters per token.
- 02
The model achieves a time to first token of 0.25 seconds, beating the 1.95-second median.
- 03
During tests, the system produced 75 million output tokens, highlighting its extreme verbosity.
Architecture and Deployment Specs
Cohere released North Mini Code on June 9, 2026, positioning the system as a 30 billion parameter agentic coding agent. [3] The model operates as a sparse mixture-of-experts Transformer featuring 128 total experts, with eight activated per token. [3] This specific architecture gives the system the effective compute profile of a 3 billion parameter dense model during inference, despite its 30 billion total parameters. [3] Cohere co-founder Nick Frosst demonstrated the model running locally on a Mac Studio utilizing MLX with approximately 20 gigabytes of RAM. [3]
Cohere made the model freely available to developers under an Apache 2.0 license. [9] The release directly targets developers seeking open-source coding agents that can operate efficiently on a single NVIDIA H100 GPU. [3] North Mini Code supports a 256,000-token total context window alongside a 64,000-token maximum generation length. [9] The model weights can be downloaded from Hugging Face in BF16, FP8, and W4A16 formats. [9]
Benchmark Performance Metrics
Independent testing conducted by Artificial Analysis ranks North Mini Code eighth out of 127 comparable open-weight models on output speed. [3] The model achieved an output rate of 210 tokens per second during these third-party tests. [3] Furthermore, Artificial Analysis recorded a time to first token of 0.25 seconds for North Mini Code, establishing a notable speed advantage compared to the class median of 1.95 seconds. [3]
The system achieved a score of 33.4 on the Artificial Analysis Coding Index. [9] In addition to the downloadable weights, Cohere made the model accessible via its own API, the Cohere Model Vault, and OpenRouter. [9] For inference hardware, the minimum requirement to run the system is one H100 at either FP8 or FP4. [9]
Specialization and Verbosity Characteristics
Unlike standard AI models that fine-tune a general-purpose base model on code data, Cohere built North Mini Code from scratch specifically for agentic software engineering tasks. [3] The system is highly optimized for multi-file code review, architecture mapping, sub-agent orchestration, and terminal tasks. [3] It is also integrated into existing scaffolds including OpenCode, SWE-Agent, and mini-SWE-Agent. [3]
However, the system exhibits notable verbosity during operational output. [3] During the Artificial Analysis benchmarking process, the model generated 75 million output tokens. [3] This volume represents three times the output of comparable models, which achieved a class median of 25 million output tokens in the same tests. [3]
Real-World Agentic Automation
Cohere has applied its AI agent platforms to internal operations by connecting the system to the cloud security platform Wiz through a custom Model Context Protocol (MCP) server. [7] The resulting security agent automates incident response workflows, handling tasks from triaging critical findings to drafting reports and updating Wiz status. [7] Prior to implementing this agentic system, manually investigating an affected asset and searching for existing tracking tickets could take between 30 minutes to two hours per critical finding. [7]
In a separate engineering application, Cohere utilized AI coding agents to automate the feedback loop required for maintaining a long-lived software fork. [8] The company applied this method to its internal fork of vLLM after a routine upstream release broke its cohere-transcribe-03-2026 ASR model. [8] The agentic workflow compressed the time required to absorb a new upstream release from weeks of intermittent developer attention to just days of mostly unattended agent time. [8]
What it means
The release of North Mini Code introduces the clearest open-source contender to emerge in 2026 for developers deciding whether to route agentic coding workloads through a self-hosted open-source model or a managed API. [3] Its time to first token of 0.25 seconds significantly undercuts the 1.95-second class median of comparable open-weight models evaluated by Artificial Analysis. [3] By utilizing a 128-expert architecture that activates only eight experts per token, Cohere matches the inference compute profile of a 3 billion parameter dense model. [3] What the sources don't address: How the model's generation of 75 million output tokens impacts downstream hosting costs compared to the 25 million token class median.
The release demonstrates a structural shift toward highly specialized, efficient AI agents. By utilizing a mixture-of-experts architecture, organizations can deploy complex reasoning models on commodity hardware like single GPUs.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
2 September 2026
Archived
29 August 2026
Event evidence refreshed from source cluster.
28 August 2026
Archived
26 August 2026
Cohere Open-Sources North Mini Code, a 30B Agentic Coding Model Optimized for Single-GPU Inference
26 August 2026
Event evidence refreshed from source cluster.
26 June 2026
Event evidence refreshed from source cluster.
24 June 2026
Event evidence refreshed from source cluster.
18 June 2026
Event evidence refreshed from source cluster.
16 June 2026
Event evidence refreshed from source cluster.
13 June 2026
Event evidence refreshed from source cluster.
9 June 2026
Event created from source cluster.
Sources
- Introducing North Mini Code: Cohere’s First Model For Developershuggingface.co
- LLM Serving FairnessCohere Search
- Why Cultural Awareness is Essential for Global AICohere Search
- Creating a Security Agent with Cohere North and Wiz | CohereCohere Search
- Automating fork maintenance with AI agents | CohereCohere Search
- Skip to contentCohere Search
- Generative AI for business: Use cases, benefits, and adoptionCohere Search
- Meet ‘North Mini Code’: Cohere’s 30B Open-Weight Mixture-of-Experts Model With 3B Active Parameters for Agentic CodingMarkTechPost
- Cohere North Mini Code Gives AI Developers More ControlAI Business
- Cohere Open-Sources North Mini Code: A 30B Coding Agent That Runs on a Single H100 - Bind AIblog.getbind.co