Skip to main content

IBM Launches Granite 4.2: Open-Source Reasoning Models with Agentic RL

25 AUGUST 2026·2 MIN READ·2 SOURCES·Official source plus independent coverage

IBM has released Granite 4.2, a family of open-source reasoning language models in 3B, 8B, and 30B parameter sizes. The models feature switchable thinking modes and native tool calling, with the larger sizes trained in live software environments.

IBM Launches Granite 4.2: Open-Source Reasoning Models with Agentic RL

Key takeaways · 3

  • 01

    Models include a switchable thinking mode and a low-effort mode for easy questions.

  • 02

    The 8B and 30B models underwent agentic reinforcement learning in live software environments.

  • 03

    All Granite 4.2 models are released under the commercial-friendly Apache 2.0 license.

New Reasoning Architecture

IBM has released Granite 4.2, its first family of dense, decoder-only reasoning language models, in three sizes: 3B, 8B, and 30B parameters. [2] Each model is pre-trained from scratch on roughly 15T tokens with a five-phase strategy that extends the context window to 512K tokens. [1] Unlike earlier Granite releases, which were instruction-following assistants, Granite 4.2 is built around explicit reasoning. [3]

Every model can emit a chain of thought before answering, and every model exposes a thinking / non-thinking switch plus a low-effort mode that spends a short reasoning budget on easy questions. [3] Supervised fine-tuning followed pre-training, utilizing about 7.2 million samples of chain-of-thought, reasoning, and agentic-trajectory data. [2]

Agentic Reinforcement Learning

The multi-stage reinforcement learning pipeline features separate training runs targeting specific capabilities. [2] All three sizes run foundational RL on verifiable rewards, such as math problems with checkable answers and code graded by hidden unit tests. [2]

Only the 8B and 30B models continue into the agentic block, a three-stage process where the model acts inside a live environment and receives rewards based on task completion. [2] In the software-engineering stage, driven by the OpenHands harness, the model edits real repositories and passes only if the hidden test suite passes. [2]

What it means

The release of Granite 4.2 positions IBM strongly in the enterprise AI space by offering models with native tool calling and explicit reasoning under an Apache 2.0 license. [1][3] The agentic RL approach for the 8B and 30B models, utilizing real environments like OpenHands and live shells, demonstrates a commitment to practical software engineering and DevOps automation applications. [2][3] What the sources don't address: How the reasoning performance and execution speed of Granite 4.2's low-effort mode compare to similar tiered reasoning features from competitors like OpenAI or Anthropic.

The release of Apache 2.0-licensed models with built-in reasoning and native tool calling lowers the barrier for enterprises to deploy advanced AI agents. The unique agentic training environment approach may offer improved reliability for specific coding and automation tasks.

Why it matters
Story quiz

Test yourself on this story — 4 questions.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 1 September 2026

    Archived

  2. 26 August 2026

    Event evidence refreshed from source cluster.

  3. 26 August 2026

    Updated with 1 new source — now corroborated by 2 sources.

  4. 26 August 2026

    Event evidence refreshed from source cluster.

  5. 25 August 2026

    Event evidence refreshed from source cluster.

  6. 25 August 2026

    Updated with 1 new source (tech_press) — now corroborated by 2 sources.

  7. 25 August 2026

    IBM Unveils Granite 4.2 Reasoning Models Featuring Agentic RL and 512K Context

  8. 25 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.