Granite 4.2 LLMs: How They're Built
2026-08-26 · Hugging Face
Overview
IBM’s Granite team introduced Granite 4.2 as the reasoning-focused generation of the Granite model family. It is the first Granite release centered on explicit reasoning and includes three dense decoder-only models: 3B, 8B, and 30B.
Compared with earlier Granite versions that emphasized instruction following, Granite 4.2 adds:
- chain-of-thought generation before answers,
- a thinking / non-thinking switch,
- a low-effort thinking mode for simpler queries,
- native tool-calling support.
All models are released under the Apache 2.0 license.
End-to-End Training Pipeline
All three model sizes share the same overall recipe:
1. pre-training from scratch,
2. supervised fine-tuning (SFT),
3. multi-stage reinforcement-learning post-training.
The clearest capability split appears in post-training:
- 8B and 30B receive an additional agentic RL stage,
- they learn to operate in real sandboxed environments by calling tools, editing/running code, using terminals, and searching the web.
For serving, Granite 4.2 can run behind OpenAI-compatible endpoints (e.g., vLLM), emits tool calls in OpenAI function-calling format, and is also supported in SGLang.
Model Architecture
Granite 4.2 uses a dense decoder-only Transformer with the following core choices:
- Attention: Grouped Query Attention (GQA)
- Positional encoding: RoPE with θ = 10,000,000
- Feed-forward: MLP with SwiGLU activation
- Normalization: RMSNorm (ε = 1e-5)
- Embeddings: untied input/output embeddings
- Precision: bfloat16
Selected configuration highlights:
- Layers: 40 (3B), 40 (8B), 64 (30B)
- Base sequence length: 131072 for all three
- KV heads: 8 across models
- MLP hidden size: 8192 / 12800 / 32768
Pre-Training Strategy
Granite 4.2 is trained from scratch on approximately 15 trillion tokens using a five-phase schedule:
- Phases 1–2: foundational pre-training
- Phases 3–4: mid-training with progressive annealing toward higher-quality data
- Phase 5: long-context training, extending context to 512K tokens
Each phase uses its own data mixture and learning-rate schedule, moving gradually from broad web-scale corpora to more curated, higher-quality sources.
SFT: Data Mix and Quality Control
SFT is designed to turn base models into reliable assistants for instruction following, reasoning, and tool use.
Scale and composition:
- about 7.2 million samples,
- roughly 100B tokens total,
- around 65B trainable tokens,
- 31.6% agentic and 68.4% non-agentic data.
Agentic data coverage
Agentic data spans:
- software engineering (69%),
- tool calling (12.1%),
- terminal use (8.0%),
- math (3.5%),
- search (0.8%),
- action (0.2%).
Trajectories are generated with diverse agent scaffolds/harnesses, including OpenHands, OpenCode, Terminus-2, SWE-agent, OpenResearcher, MiniSWE, OpenSeeker, EnvScaler, Gemini CLI, Hermes, Codex, and Goose. The corpus combines open-source datasets with IBM’s synthetic RL environments.
Non-agentic categories
Major categories include:
- instruction following (18.8%),
- coding (18.8%),
- math (14.6%),
- multilingual (7.0%),
- science (5.4%),
- reasoning (3.0%),
- safety (0.8%).
Quality-control pipeline
Before entering the final SFT mixture, samples go through multiple filters:
1. normalization into a consistent OpenAI Chat format,
2. LLM-as-judge scoring using GPT-OSS-120B and Gemma 4,
3. removal of low-quality items, hallucinated/fabricated content, invalid tool interactions, and calls to undefined functions,
4. targeted dataset-specific heuristic cleaning,
5. local and global deduplication (the source text is truncated after this point).
Takeaway
Granite 4.2 combines large-scale from-scratch pre-training, reasoning- and tool-centric SFT, and staged RL post-training under one architecture family. The added agentic RL for 8B/30B is a key differentiator for real-environment operation. Overall, the release emphasizes both open licensing and practical deployment compatibility.