LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation

2026-08-20 · Hugging Face

LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation

Introduction

Liquid AI today released updated Q4_0 GGUF checkpoints for its LFM2.5 model series. The new checkpoints cover four model sizes: LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. These models were trained using Quantization-Aware Distillation (QAD), a technique that enables 4-bit quantized models to maintain significantly higher quality than traditional post-training quantization (PTQ) methods while preserving the exact same memory footprint and inference speed of native Q4_0.

What is Quantization-Aware Distillation?

QAD works by distilling knowledge from a high-precision (typically BF16) teacher model directly into a quantized student model during the training process. This approach allows the student to learn to compensate for quantization errors from the beginning, rather than attempting to recover quality after the fact.

The resulting QAD Q4_0 checkpoints offer three major benefits:

  • Same memory usage and throughput as standard Q4_0 GGUF files
  • Recovery of approximately 97% of the accuracy lost due to quantization
  • Drop-in compatibility with any runtime that supports GGUF Q4_0 artifacts, particularly llama.cpp

Benchmark Results

The team evaluated both PTQ-generated Q4_0 checkpoints and the new QAD Q4_0 checkpoints across a comprehensive suite of benchmarks covering reasoning, instruction following, tool use, and agentic capabilities:

  • GPQA Diamond
  • MMLU-Pro
  • IFEval
  • IFBench
  • Multi-IF
  • BFCLv4

Scale-appropriate math evaluations were also included: GSM8K for the 230M and 350M models, and AIME25 for the 1.2B and 2.6B models. All scores are averaged over five repeats, with BF16 GGUF serving as the upper-bound reference.

Across all four models, the QAD checkpoints substantially outperform their PTQ counterparts. They recover the following percentages of their respective BF16 baseline performance:

  • LFM2.5-230M: 97.1%
  • LFM2.5-350M: 96.5%
  • LFM2.5-1.2B-Instruct: 97.4%
  • LFM2.5-2.6B: 96.6%

Performance on Real Edge Hardware

To demonstrate practical value, decode throughput was measured on four representative edge devices:

  • MacBook Pro (GPU inference)
  • NucBox EVO-X2 (GPU inference)
  • Samsung Galaxy S26 Ultra (Arm CPU inference)
  • Raspberry Pi 5 (Arm CPU inference)

Results show that the QAD Q4_0 checkpoints deliver strong efficiency-quality tradeoffs:

  • The 230M and 350M models match Q5_K_M quality (within evaluation variance) while offering 4-33% higher decode throughput.
  • The 1.2B and 2.6B models match Q4_K_M quality at 3-14% higher throughput.

The QAD checkpoints also perform on par with Unsloth’s strong UD-Q4_K_XL post-training quantization baseline where applicable.

How to Use QAD GGUFs

Using the models is straightforward with any GGUF-compatible runtime. Example command using llama.cpp:


llama-cli -hf LiquidAI/LFM2.5-350M \
  --hf-file LFM2.5-350M-QAD-Q4_0.gguf \
  -p "What is C. elegans?"

Availability and Citation

All four QAD Q4_0 GGUF files are available immediately on Hugging Face under the LiquidAI organization for the respective model repositories.

For citation, please use:

Liquid AI, "LFM2.5 Q4_0: Quantization-Aware Distillation for Edge Deployment", Liquid AI Blog, August 2026.

The release represents a significant step forward in making high-quality language models practical for widespread edge deployment.

Source