IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B parameter sizes. Unlike earlier Granite releases, which were instruction-following assistants, Granite 4.2 is built around explicit reasoning. Every model can emit a chain of thought before answering, and every model exposes a thinking / non-thinking switch plus a low-effort mode that spends a short reasoning budget on easy questions. The models are decoder-only dense transformers, pre-trained from scratch on roughly 15 trillion tokens, then post-trained through a multi-stage reinforcement learning chain. For the 8B and 30B, that chain includes an agentic RL block where the model learns to edit code, drive a terminal, and run web searches inside real sandboxed environments. All three ship under Apache 2.0. IBM also released two 470M-parameter Granite Speech 5.0 Turbo CTC models alongside the LLMs.
‘+rows[i][0]+’‘+txt+’
‘;
body
.g42v===undefined)?’NA’:v.toFixed(2);” h=”” class=”"bar"”>
‘+rows[i][0]+’‘+txt+’
‘;
.g42 .hd
.g42 .kick{font-size:11px;letter-spacing:.14em;text-transform:uppercase;color:#78a9ff;font-weight:600}
.g42 h3{font-size:20px;line-height:1.25;margin:6px 0 4px;color:#f4f4f4;font-weight:600}
.g42 .sub{font-size:13px;color:#a8a8a8;line-height:1.5}
.g42 .sizes{display:flex;gap:8px;margin-top:14px;flex-wrap:wrap}
.g42 .sz{background:#262626;border:1px solid #393939;color:#c6c6c6;font-size:12px;font-weight:600;padding:7px 14px;border-radius:2px;cursor:pointer;transition:all .18s}
.g42 .sz:hover{border-color:#78a9ff;color:#f4f4f4}
.g42 .sz.on{background:#0f62fe;border-color:#0f62fe;color:#fff}
.g42 .tabs{display:flex;border-bottom:1px solid #393939;background:#1c1c1c;overflow-x:auto}
.g42 .tb{flex:1;min-width:118px;background:none;border:none;border-bottom:2px solid transparent;color:#8d8d8d;font-size:12px;font-weight:600;padding:13px 8px;cursor:pointer;white-space:nowrap;font-family:inherit;transition:all .18s}
.g42 .tb:hover{color:#f4f4f4;background:#262626}
.g42 .tb.on{color:#78a9ff;border-bottom-color:#0f62fe;background:#161616}
.g42 .pane{display:none;padding:20px}
.g42 .pane.on{display:block;animation:fi .3s ease}
@keyframes fi{from{opacity:0;transform:translateY(6px)}to{opacity:1;transform:none}}
.g42 .lbl{font-size:11px;letter-spacing:.1em;text-transform:uppercase;color:#8d8d8d;font-weight:600;margin-bottom:10px}
.g42 .row{display:flex;gap:8px;flex-wrap:wrap;margin-bottom:16px}
.g42 .btn{background:#262626;border:1px solid #393939;color:#c6c6c6;font-size:12px;font-weight:600;padding:8px 14px;border-radius:2px;cursor:pointer;font-family:inherit;transition:all .18s}
.g42 .btn:hover{border-color:#78a9ff;color:#f4f4f4}
.g42 .btn.on{background:#0f62fe;border-color:#0f62fe;color:#fff}
.g42 .term{background:#0b0b0b;border:1px solid #393939;border-radius:2px;padding:14px;font-family:’IBM Plex Mono’,ui-monospace,Menlo,Consolas,monospace;font-size:12px;line-height:1.65;color:#c6c6c6;min-height:150px;white-space:pre-wrap;word-break:break-word}
.g42 .th{color:#78a9ff}
.g42 .an{color:#42be65}
.g42 .cur{display:inline-block;width:7px;height:14px;background:#0f62fe;vertical-align:-2px;animation:bl .9s step-end infinite}
@keyframes bl{50%{opacity:0}}
.g42 .meter{height:6px;background:#262626;border-radius:3px;overflow:hidden;margin:14px 0 6px}
.g42 .meter i{display:block;height:100%;background:linear-gradient(90deg,#0f62fe,#78a9ff);width:0;transition:width .9s cubic-bezier(.2,.8,.2,1)}
.g42 .mtx{display:flex;justify-content:space-between;font-size:11px;color:#8d8d8d}
.g42 .pipe{display:flex;gap:6px;overflow-x:auto;padding-bottom:8px;margin-bottom:14px}
.g42 .stg{flex:1;min-width:76px;background:#262626;border:1px solid #393939;border-radius:2px;padding:11px 6px;text-align:center;cursor:pointer;transition:all .2s;position:relative}
.g42 .stg:hover{border-color:#78a9ff}
.g42 .stg .n{font-size:11px;font-weight:700;color:#f4f4f4;display:block}
.g42 .stg .t{font-size:9px;color:#8d8d8d;margin-top:3px;display:block;letter-spacing:.06em}
.g42 .stg.on{background:#0f62fe;border-color:#0f62fe}
.g42 .stg.on .n,.g42 .stg.on .t{color:#fff}
.g42 .stg.off{opacity:.28}
.g42 .stg.pulse{animation:pl .6s ease}
@keyframes pl{0%{box-shadow:0 0 0 0 rgba(15,98,254,.7)}100%{box-shadow:0 0 0 12px rgba(15,98,254,0)}}
.g42 .card{background:#1c1c1c;border:1px solid #393939;border-left:3px solid #0f62fe;border-radius:2px;padding:14px}
.g42 .card h4{font-size:13px;color:#f4f4f4;margin-bottom:6px}
.g42 .card p{font-size:12.5px;color:#a8a8a8;line-height:1.6}
.g42 .kv{display:grid;grid-template-columns:repeat(auto-fit,minmax(92px,1fr));gap:8px;margin-top:12px}
.g42 .kv div{background:#262626;border:1px solid #393939;border-radius:2px;padding:8px}
.g42 .kv b{display:block;font-size:14px;color:#78a9ff;font-family:’IBM Plex Mono’,monospace}
.g42 .kv span{font-size:10px;color:#8d8d8d;letter-spacing:.06em;text-transform:uppercase}
.g42 .loop{display:flex;gap:10px;align-items:center;justify-content:center;flex-wrap:wrap;margin:6px 0 16px}
.g42 .nd{background:#262626;border:1px solid #393939;border-radius:2px;padding:12px 14px;font-size:12px;font-weight:600;color:#8d8d8d;transition:all .3s;min-width:96px;text-align:center}
.g42 .nd.hot{background:#0f62fe;border-color:#0f62fe;color:#fff;transform:translateY(-3px)}
.g42 .ar{color:#525252;font-size:15px}
.g42 .bar{margin-bottom:11px}
.g42 .bar .bt{display:flex;justify-content:space-between;font-size:11.5px;color:#c6c6c6;margin-bottom:4px}
.g42 .bar .bt em{font-style:normal;font-family:’IBM Plex Mono’,monospace;color:#78a9ff}
.g42 .bar .bw{height:9px;background:#262626;border-radius:2px;overflow:hidden}
.g42 .bar .bw i{display:block;height:100%;background:linear-gradient(90deg,#0f62fe,#78a9ff);width:0;transition:width 1s cubic-bezier(.2,.8,.2,1)}
.g42 .note{font-size:11px;color:#6f6f6f;margin-top:12px;line-height:1.55}
.g42 .ft{border-top:1px solid #393939;padding:12px 20px;display:flex;justify-content:space-between;align-items:center;flex-wrap:wrap;gap:8px;background:#1c1c1c}
.g42 .ft span{font-size:11px;color:#6f6f6f}
.g42 .ft b{color:#0f62fe;font-weight:600}
.g42 .ft a{color:#78a9ff;text-decoration:none;font-size:11px}
.g42 .ft a:hover{text-decoration:underline}
@media(max-width:640px){
.g42 h3{font-size:17px}
.g42 .pane{padding:15px}
.g42 .tb{min-width:96px;font-size:11px;padding:11px 6px}
.g42 .nd{min-width:80px;font-size:11px;padding:10px}
.g42 .ar{display:none}
}
</style>
<div class="g42">
<div class="hd">
<div class="kick">Interactive Explainer</div>
<h3>IBM Granite 4.2: reasoning modes, staged RL, and agentic behavior</h3>
<div class="sub">Pick a model size, then step through how Granite 4.2 thinks, how it was trained, and how it scores.</div>
<div class="sizes">
<button class="sz" data-s="3">3B Dense</button>
<button class="sz on" data-s="8">8B Dense</button>
<button class="sz" data-s="30">30B Dense</button>
</div>
</div>
<div class="tabs">
<button class="tb on" data-p="p1">Thinking switch</button>
<button class="tb" data-p="p2">RL curriculum</button>
<button class="tb" data-p="p3">Agentic loop</button>
<button class="tb" data-p="p4">Benchmarks</button>
</div>
<div class="pane on" id="p1">
<div class="lbl">Choose a reasoning budget</div>
<div class="row">
<button class="btn on md" data-m="think">Thinking</button>
<button class="btn md" data-m="low">Low effort</button>
<button class="btn md" data-m="non">Non-thinking</button>
</div>
<div class="term" id="t1"></div>
<div class="meter"><i id="m1"></i></div>
<div class="mtx"><span id="mL">Reasoning budget</span><span id="mR"></span></div>
<div class="note">Every Granite 4.2 model exposes the same switch through the chat template: <code>enable_thinking</code> and <code>low_effort</code>. Prior turns are stripped by default via <code>truncate_history_thinking=True</code>.</div>
</div>
<div class="pane" id="p2">
<div class="lbl">Post-training chain <span id="p2s"></span></div>
<div class="pipe" id="pipe"></div>
<div class="row"><button class="btn" id="play">Run the curriculum</button></div>
<div class="card" id="sc"></div>
<div class="note">Stage hyperparameters shown are the published 30B settings. Each stage is a separate GRPO run that warm-starts from the previous checkpoint. The 3B model skips the agentic block entirely.</div>
</div>
<div class="pane" id="p3">
<div class="lbl">One agentic turn, step by step</div>
<div class="loop">
<div class="nd" data-i="0">Reason</div><span class="ar">→</span>
<div class="nd" data-i="1">Call tool</div><span class="ar">→</span>
<div class="nd" data-i="2">Observe</div><span class="ar">→</span>
<div class="nd" data-i="3">Answer</div>
</div>
<div class="row"><button class="btn" id="play3">Play the loop</button><button class="btn" id="step3">Step</button></div>
<div class="term" id="t3"></div>
<div class="note">Granite 4.2 emits OpenAI-format function calls over an OpenAI-compatible endpoint, so it drops into agent harnesses without adapters.</div>
</div>
<div class="pane" id="p4">
<div class="lbl">Reported scores <span id="p4s"></span></div>
<div class="row">
<button class="btn on cat" data-c="code">Agentic coding</button>
<button class="btn cat" data-c="gen">Agentic / tools</button>
<button class="btn cat" data-c="reason">Reasoning</button>
<button class="btn cat" data-c="long">Long context</button>
</div>
<div id="bars"></div>
<div class="note">Numbers are IBM-reported figures from the Granite 4.2 technical write-up. NA means the size was not evaluated on that benchmark. GDPval is an index score, not a percentage, and is excluded from these bars.</div>
</div>
<div class="ft">
<span>Built by <b>Marktechpost</b></span>
<a href="https://huggingface.co/collections/ibm-granite/granite-42-language-models" target="_blank" rel="noopener">Granite 4.2 on Hugging Face →</a>
</div>
</div>
<script>
(function(){
var size=” tabs=”” var=”” tbs=”document.querySelectorAll(‘.g42″ .tb=”” for=”” i=”0;i<tbs.length;i++){tbs[i].onclick=function(){” j=”0;j<tbs.length;j++){tbs[j].classList.remove(‘on’);}” this.classlist.add=”” ps=”document.querySelectorAll(‘.g42″ .pane=”” k=”0;k<ps.length;k++){ps[k].classList.remove(‘on’);}” document.getelementbyid=”” if=”” res=”” size=”” selector=”” szs=”document.querySelectorAll(‘.g42″ .sz=”” s=”0;s<szs.length;s++){szs[s].onclick=function(){” drawpipe=”” thinking=”” switch=”” modes=”{” think:=”” to=”” new=”” tokens=”” let=”” see.=”” the=”” problem=”” is=”” find=”” how=”” many=”” are=”” in=”” it=”” out:=”” t=”” r=”” a=”” w=”” b=”” e=”” y.=”” r.=”” position=”” ans:=”” word=”” low:=”” low_effort=”True’,pct:34,budget:’short,” capped=”” reasoning=”” budget=”” answer.=”” non:=”” chain=”” of=”” thought=”” emitted=”” capital=”” france=”” paris.=”” t1=”document.getElementById(‘t1’),m1=document.getElementById(‘m1’),mR=document.getElementById(‘mR’),tmr=null;” function=”” typemode=”” d=”MODES[k];clearTimeout(tmr);” head=”<span style="color:#6f6f6f">> “>\n\n’;
var full=(d.think?’<think>\n’+esc(d.think)+’\n</think>\n\n’:’<think></think>\n\n’)+’‘+esc(d.ans)+’‘;
var plain=(d.think?’
var n=0,step=Math.max(2,Math.ceil(plain.length/110));
m1.style.width=”0%”;setTimeout(function(){m1.style.width=d.pct+’%’;},40);
mR.textContent=d.budget;
(function tick(){
n+=step;
if(n>=plain.length){t1.innerHTML=head+full;res();return;}
t1.innerHTML=head+esc(plain.slice(0,n))+’‘;
tmr=setTimeout(tick,16);
})();
}
function esc(x){return x.replace(/&/g,’&’).replace(//g,’>’);}
var mds=document.querySelectorAll(‘.g42 .md’);
for(var a=0;a
e.onclick=function(){sel(i);};
pipe.appendChild(e);
})(i);
}
sel(0);
}
function sel(i){
var ns=pipe.children;
for(var j=0;j
‘;}
sc.innerHTML=’
‘+st.n+(skip?’ · not run for 3B’:”)+’
‘+st.d+’
‘+kv+’
‘;
res();
}
document.getElementById(‘play’).onclick=function(){
var i=0;var b=this;b.disabled=true;
var iv=setInterval(function(){
if(i>=STAGES.length){clearInterval(iv);b.disabled=false;return;}
if(!(size===’3’&&STAGES[i].agentic)){
sel(i);pipe.children[i].classList.add(‘pulse’);
(function(n){setTimeout(function(){pipe.children[n].classList.remove(‘pulse’);},600);})(i);
}
i++;
},700);
};
drawPipe();
/* ———- 3. agentic loop ———- */
var LOOP=[
‘> user: What\u0027s the weather like in Boston right now?\n\n<think>\nThe user is asking for the weather in Boston. Let me check the\ntools available. There is get_current_weather, which takes a\ncity parameter. I need to call it with city set to Boston.\n</think>‘,
‘> tool call emitted\n\n<tool_call>\n\n Boston\n\n</tool_call>‘,
‘> role: tool\n\n{“temperature”: “72°F”,\n “condition”: “Partly cloudy”,\n “humidity”: “65%”}’,
‘<think>\nThe tool returned the weather data for Boston. I need to\npresent this clearly to the user.\n</think>\n\nThe current weather in Boston is 72°F, partly cloudy, with\n65% humidity.‘
];
var t3=document.getElementById(‘t3’),nds=document.querySelectorAll(‘.g42 .nd’),cur=0,iv3=null;
function show3(i){
cur=i;
for(var j=0;j
};
show3(0);
/* ———- 4. benchmarks ———- */
var B={
code:[[‘SWE-Bench Verified’,null,47.67,57.00],[‘SWE-Bench Multilingual’,null,30.78,41.89],[‘SWE-Bench Pro’,null,19.11,33.29],[‘Terminal-Bench 2.1’,null,20.56,29.24]],
gen:[[‘tau3-bench’,50.99,66.34,68.05],[‘BFCL (v4)’,52.41,50.29,61.39],[‘ProfBench’,32.10,41.20,42.90],[‘BirdBench’,null,41.07,41.85]],
reason:[[‘AIME25’,78.33,86.67,89.17],[‘HMMT Feb25’,66.67,78.33,89.17],[‘GPQA’,54.80,64.14,66.41],[‘LiveCodeBench v6’,69.71,73.24,75.77],[‘SciCode’,24.11,36.09,38.76],[‘MMLU-Pro’,67.84,74.04,77.60],[‘Arena-Hard-V2’,34.96,65.19,67.93]],
long:[[‘RULER 64K’,67.52,80.99,89.96],[‘RULER 128K’,55.30,71.41,81.38]]
};
var cat=”code”,bars=document.getElementById(‘bars’),p4s=document.getElementById(‘p4s’);
var cts=document.querySelectorAll(‘.g42 .cat’);
for(var c=0;c
‘+rows[i][0]+’‘+txt+’
‘;
}
bars.innerHTML=h;
var is=bars.querySelectorAll(‘i’);
setTimeout(function(){for(var k=0;k
“>
Is it deployable?
Yes, All three Granite 4.2 language models ship under Apache 2.0, so download, fine-tuning, and commercial production use carry no licensing gate.
- Which companies: The 3B fits solo developers and startups running on a laptop through Ollama or LM Studio, especially with the released GGUF quants down to Q4_K_M. The 8B suits mid-market teams on a single modern GPU. The 30B targets enterprises with A100/H100-class capacity, or FP8/NVFP4 serving on vLLM. Regulated organizations get the additional benefit of on-prem weights.
- Industries: Software and developer tooling, financial services, healthcare, telecom, public sector, and contact centers, which is where the new speech models land.
- Applications: Software engineering agents, terminal and DevOps automation, deep-research and search agents, long-document RAG, structured tool calling, and high-volume transcription.
Architecture
Granite 4.2 is a decoder-only dense transformer, not a hybrid or MoE design. Core components are Grouped Query Attention with 8 KV heads, RoPE with θ = 10,000,000, SwiGLU MLPs, RMSNorm (ε = 1e-5), untied input/output embeddings, and bfloat16 precision.
The 3B uses 40 layers at embedding size 2560. The 8B uses 40 layers at 4096. The 30B goes to 64 layers with an MLP hidden size of 32,768. The published architecture table lists a 131,072-token (128K) sequence length, while the five-phase pre-training run includes a long-context phase extending to 512K tokens. Pre-training covers roughly 15 trillion tokens from scratch.
The training pipeline is the actual story
Supervised fine-tuning uses about 7.2 million samples, roughly 100B tokens with ~65B trainable. The mixture is 31.6% agentic and 68.4% non-agentic, and software engineering is 69% of the agentic slice. Trajectories were generated across harnesses including OpenHands, SWE-agent, Terminus-2, MiniSWE, Codex, and Goose. Quality control used GPT-OSS-120B and Gemma 4 as judges, plus SHA-256 deduplication over the tools and messages fields.
Post-training is a multi-stage, multi-environment RL chain, not a single pass. Each stage is a separate asynchronous GRPO run that warm-starts from the previous checkpoint, with a leave-one-out baseline instead of a value network and truncated importance sampling to bound off-policy drift. The order is RLVR, then skill boosters, then SWE, Terminal, Search, then RLHF.
The agentic RL block runs only on the 8B and 30B. The 3B takes foundational RL and alignment only. That single design choice explains most of the capability gap across sizes. Training ran on NeMo-RL and NeMo-Gym over an NVIDIA GB200 NVL72 cluster hosted by CoreWeave.
Two supporting pieces matter: 1 trillion tokens of synthetic code from IBM’s CodeAlchemy pipeline, and a speculative decoding layer for faster serving.
Reported results
IBM’s numbers, by size (3B / 8B / 30B):
| Benchmark | 3B | 8B | 30B |
|---|---|---|---|
| SWE-Bench Verified | NA | 47.67 | 57.00 |
| Terminal-Bench 2.1 | NA | 20.56 | 29.24 |
| τ³-bench | 50.99 | 66.34 | 68.05 |
| BFCL (v4) | 52.41 | 50.29 | 61.39 |
| AIME25 | 78.33 | 86.67 | 89.17 |
| GPQA | 54.80 | 64.14 | 66.41 |
| MMLU-Pro | 67.84 | 74.04 | 77.60 |
| RULER 128K | 55.30 | 71.41 | 81.38 |
Speech: 470M parameters, no LLM backbone
The Turbo CTC models come in at 470 million parameters and drop the LLM backbone entirely, using connectionist temporal classification to map audio to text. IBM reports an RTFx throughput near 12,600 on a single H200, against roughly 6,000 for current speed leaders on the Open ASR leaderboard. A WebGPU demo is live.
Key Takeaways
- Granite 4.2 ships dense 3B/8B/30B reasoning models under Apache 2.0.
- A thinking / low-effort / non-thinking switch is exposed in the chat template.
- Agentic RL (SWE, Terminal, Search) trains only the 8B and 30B.
- The 30B hits 57.00 on SWE-Bench Verified and 29.24 on Terminal-Bench 2.1.
- Granite Speech 5.0 Turbo CTC is 470M parameters with no LLM backbone.
Check out the IBM Research blog, the technical write-up, and the GitHub repo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.



