Granite 4.2 LLMs: How They're Built

2026-08-26 · Hugging Face

Overview

IBM’s Granite team introduced Granite 4.2 as the reasoning-focused generation of the Granite model family. It is the first Granite release centered on explicit reasoning and includes three dense decoder-only models: 3B, 8B, and 30B.

Compared with earlier Granite versions that emphasized instruction following, Granite 4.2 adds:

  • chain-of-thought generation before answers,
  • a thinking / non-thinking switch,
  • a low-effort thinking mode for simpler queries,
  • native tool-calling support.

All models are released under the Apache 2.0 license.

End-to-End Training Pipeline

All three model sizes share the same overall recipe:

1. pre-training from scratch,

2. supervised fine-tuning (SFT),

3. multi-stage reinforcement-learning post-training.

The clearest capability split appears in post-training:

  • 8B and 30B receive an additional agentic RL stage,
  • they learn to operate in real sandboxed environments by calling tools, editing/running code, using terminals, and searching the web.

For serving, Granite 4.2 can run behind OpenAI-compatible endpoints (e.g., vLLM), emits tool calls in OpenAI function-calling format, and is also supported in SGLang.

Model Architecture

Granite 4.2 uses a dense decoder-only Transformer with the following core choices:

  • Attention: Grouped Query Attention (GQA)
  • Positional encoding: RoPE with θ = 10,000,000
  • Feed-forward: MLP with SwiGLU activation
  • Normalization: RMSNorm (ε = 1e-5)
  • Embeddings: untied input/output embeddings
  • Precision: bfloat16

Selected configuration highlights:

  • Layers: 40 (3B), 40 (8B), 64 (30B)
  • Base sequence length: 131072 for all three
  • KV heads: 8 across models
  • MLP hidden size: 8192 / 12800 / 32768

Pre-Training Strategy

Granite 4.2 is trained from scratch on approximately 15 trillion tokens using a five-phase schedule:

  • Phases 1–2: foundational pre-training
  • Phases 3–4: mid-training with progressive annealing toward higher-quality data
  • Phase 5: long-context training, extending context to 512K tokens

Each phase uses its own data mixture and learning-rate schedule, moving gradually from broad web-scale corpora to more curated, higher-quality sources.

SFT: Data Mix and Quality Control

SFT is designed to turn base models into reliable assistants for instruction following, reasoning, and tool use.

Scale and composition:

  • about 7.2 million samples,
  • roughly 100B tokens total,
  • around 65B trainable tokens,
  • 31.6% agentic and 68.4% non-agentic data.

Agentic data coverage

Agentic data spans:

  • software engineering (69%),
  • tool calling (12.1%),
  • terminal use (8.0%),
  • math (3.5%),
  • search (0.8%),
  • action (0.2%).

Trajectories are generated with diverse agent scaffolds/harnesses, including OpenHands, OpenCode, Terminus-2, SWE-agent, OpenResearcher, MiniSWE, OpenSeeker, EnvScaler, Gemini CLI, Hermes, Codex, and Goose. The corpus combines open-source datasets with IBM’s synthetic RL environments.

Non-agentic categories

Major categories include:

  • instruction following (18.8%),
  • coding (18.8%),
  • math (14.6%),
  • multilingual (7.0%),
  • science (5.4%),
  • reasoning (3.0%),
  • safety (0.8%).

Quality-control pipeline

Before entering the final SFT mixture, samples go through multiple filters:

1. normalization into a consistent OpenAI Chat format,

2. LLM-as-judge scoring using GPT-OSS-120B and Gemma 4,

3. removal of low-quality items, hallucinated/fabricated content, invalid tool interactions, and calls to undefined functions,

4. targeted dataset-specific heuristic cleaning,

5. local and global deduplication (the source text is truncated after this point).

Takeaway

Granite 4.2 combines large-scale from-scratch pre-training, reasoning- and tool-centric SFT, and staged RL post-training under one architecture family. The added agentic RL for 8B/30B is a key differentiator for real-environment operation. Overall, the release emphasizes both open licensing and practical deployment compatibility.

Source