IBM Launches Granite 4.2: Open-Source Reasoning Models with Agentic RL
IBM has released Granite 4.2, a family of open-source reasoning language models in 3B, 8B, and 30B parameter sizes. The models feature switchable thinking modes and native tool calling, with the larger sizes trained in live software environments.

Key takeaways · 3
- 01
Models include a switchable thinking mode and a low-effort mode for easy questions.
- 02
The 8B and 30B models underwent agentic reinforcement learning in live software environments.
- 03
All Granite 4.2 models are released under the commercial-friendly Apache 2.0 license.
New Reasoning Architecture
IBM has released Granite 4.2, its first family of dense, decoder-only reasoning language models, in three sizes: 3B, 8B, and 30B parameters. [2] Each model is pre-trained from scratch on roughly 15T tokens with a five-phase strategy that extends the context window to 512K tokens. [1] Unlike earlier Granite releases, which were instruction-following assistants, Granite 4.2 is built around explicit reasoning. [3]
Every model can emit a chain of thought before answering, and every model exposes a thinking / non-thinking switch plus a low-effort mode that spends a short reasoning budget on easy questions. [3] Supervised fine-tuning followed pre-training, utilizing about 7.2 million samples of chain-of-thought, reasoning, and agentic-trajectory data. [2]
Agentic Reinforcement Learning
The multi-stage reinforcement learning pipeline features separate training runs targeting specific capabilities. [2] All three sizes run foundational RL on verifiable rewards, such as math problems with checkable answers and code graded by hidden unit tests. [2]
Only the 8B and 30B models continue into the agentic block, a three-stage process where the model acts inside a live environment and receives rewards based on task completion. [2] In the software-engineering stage, driven by the OpenHands harness, the model edits real repositories and passes only if the hidden test suite passes. [2]
What it means
The release of Granite 4.2 positions IBM strongly in the enterprise AI space by offering models with native tool calling and explicit reasoning under an Apache 2.0 license. [1][3] The agentic RL approach for the 8B and 30B models, utilizing real environments like OpenHands and live shells, demonstrates a commitment to practical software engineering and DevOps automation applications. [2][3] What the sources don't address: How the reasoning performance and execution speed of Granite 4.2's low-effort mode compare to similar tiered reasoning features from competitors like OpenAI or Anthropic.
The release of Apache 2.0-licensed models with built-in reasoning and native tool calling lowers the barrier for enterprises to deploy advanced AI agents. The unique agentic training environment approach may offer improved reliability for specific coding and automation tasks.
Why it matters
Test yourself on this story — 4 questions.
Create a free account to take the quiz, earn XP, and get a daily session built for your industry.
Take the quizHow this developed
1 September 2026
Archived
26 August 2026
Event evidence refreshed from source cluster.
26 August 2026
Updated with 1 new source — now corroborated by 2 sources.
26 August 2026
Event evidence refreshed from source cluster.
25 August 2026
Event evidence refreshed from source cluster.
25 August 2026
Updated with 1 new source (tech_press) — now corroborated by 2 sources.
25 August 2026
IBM Unveils Granite 4.2 Reasoning Models Featuring Agentic RL and 512K Context
25 August 2026
Event created from source cluster.
Sources
- Granite 4.2 LLMs: How They're BuiltHugging Face - Blog
- IBM’s Granite 4.2 Models Learn to Think and Act Inside EnvironmentsUnite.AI
- IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise ModelsMarkTechPost
- IBM's new Granite 4.2 models ride the wave of interest in local LLMsArs Technica AI