Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
2026-08-26 · Hugging Face
Quantization-Aware Healing: A Compressed, 4-Bit Model That Outperforms Its Full-Precision Original
Making a large language model smaller almost always incurs a cost. The standard recipe for efficient deployment compresses the architecture first—cutting parameters by removing layers, heads, or neurons—and then quantizes the remaining weights down to 4 bits to shrink memory and compute further. While both steps save resources, they systematically degrade the capabilities users care about most: reasoning, mathematical problem-solving, and code generation. Consequently, serious deployment pipelines add a recovery step, known as "healing," before production. Recent open-weight releases like gpt-oss, NVIDIA's Nemotron family, and Hypernova 60B all rely on this compress-then-heal approach.
Why Usual Healing Methods Fall Short
Most efficiency pipelines follow three steps: structural compression, quantization, and healing. The differences lie entirely in that final step:
- Quantization-Aware Training (QAT): The dominant method inserts fake-quantization operators into the forward pass and continues fine-tuning on a task loss. In practice, this requires re-running an expensive multi-stage post-training process (SFT, RLHF, agentic tuning) through a noisier, lower-precision forward pass. It is costly and, as results show, can become unstable if training continues past its optimal point.
- Quantization-Aware Distillation (QAD): This alternative avoids re-running training history. It distills a frozen full-precision teacher directly into the quantized student via KL-divergence loss on output logits. This works well if only quantization is applied. However, once structural compression occurs, there is no independently trained full-precision version of the smaller architecture. The only candidate teacher is the recovered bfloat16 checkpoint, which is itself a degraded approximation of the original model. Distilling from it anchors the student to a degraded target, capping its accuracy at the recovered checkpoint's ceiling.
The QAH Approach
Quantization-Aware Healing (QAH) removes that accuracy ceiling with one fundamental change: it distills directly from the original, pre-compression model rather than the recovered one.
- Architecture-Agnostic Distillation: The teacher and student do not even share an architecture. The teacher is full-size and full-precision; the student is half the size and running in MXFP4. Because a teacher's output distribution is architecture-agnostic, the size or shape mismatch does not prevent transfer. The student only sees the teacher's output distribution, matched via KL divergence on the logits.
- Reframing the Quantization Stage: Under QAH, quantization is no longer a lossy postprocessing step applied after healing. It becomes a second, full pass of distillation against the original teacher, providing supervision the bfloat16 checkpoint never received. The 4-bit student is not just compensating for information lost to quantization; it is picking up information the earlier recovery stage lacked the time or data to transfer.
- Stability Benefits: Because KL distillation ties the student to a fixed teacher distribution, there is no pressure to drift once the student catches up. In contrast, cross-entropy task loss pushes the student toward hard labels indefinitely, harming accuracy and stability.
- Long Context Support: To support healing corpora with documents up to 32k tokens, QAH reuses a memory-efficient chunked KL-divergence loss. This computes KL one sequence slice at a time, never materializing the full vocabulary-by-sequence grid, fitting 32k-token healing within a fixed GPU memory budget.
Results
The researchers applied QAH to a GPT-OSS 120B model, which was compressed to 60B parameters, recovered in bfloat16, and then re-quantized to MXFP4 under QAH. The resulting 4-bit model beats its own full-precision (bfloat16) version on 7 out of 9 benchmarks. The 4-bit model ends up smaller, cheaper to run, and more accurate than the checkpoint it was quantized from, effectively inverting the usual relationship between a 4-bit model and its 16-bit origin.