Up to 3.2x Faster Inference with LFM2.5-DSpark

2026-08-21 · Hugging Face

Up to 3.2x Faster Inference with LFM2.5-DSpark

Release Overview

Today, Liquid AI releases DSpark draft model checkpoints for three models from the LFM2.5 family: LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These add a speculative decoding path that trades a minimal memory increase for a large decoding speedup without changing output quality.

Key benefits include:

  • Up to 3.18x throughput improvement on GPU
  • Up to 2.87x speedup on-device
  • 57% average reduction in function-calling latency for LFM2.5-2.6B
  • Day-one support for llama.cpp and SGLang with upstream open-sourced integration

How DSpark Works

The decode phase in LLM inference is traditionally memory-bound. Most latency comes from streaming weights from DRAM into SRAM rather than computation. Speculative decoding addresses this by using a lightweight draft model to produce candidate tokens, which the target model then verifies in a single forward pass, amortizing the weight loading cost.

DSpark combines three components:

  • DFlash-style parallel backbone conditioned on the target model’s context features, producing hidden states for all draft tokens in one forward pass.
  • Lightweight sequential head modeled as a Markov chain between neighboring tokens, adding inter-token dependency to raise acceptance rates at later positions.
  • Confidence-scheduled verifier that predicts each token’s survival probability and prunes low-confidence suffixes when verification cost exceeds benefit.

Training and Architecture

The team followed the DSpark recipe with a larger, more diverse data mix covering SFT, chat, code, and function-calling data. Based on ablations, they used simplified attention-only draft models with 5 layers and a block size of 9. Models were trained for 15 epochs on the full dataset; the checkpoint with the highest acceptance rate was selected.

The resulting draft models are compact. Parameter counts are as follows:

Parameter Breakdown

| Component | LFM2.5-1.2B-Instruct | LFM2.5-2.6B | LFM2.5-8B-A1B |

|-----------|----------------------|-------------|---------------|

| Decoder stack (5 layers) | 241.2M | 241.2M | 241.2M |

| Hidden-state projection | 21.0M | 21.0M | 21.0M |

| Markov head | 33.6M | 65.5M | 65.5M |

| Norms + confidence head | 0.0275M | 0.0275M | 0.0275M |

| Total | 295.7M | 327.7M | 327.7M |

Quality Parity

Under greedy decoding, a draft token is accepted only if it matches the target model’s distribution. On rejection, the target model’s own token replaces it. The emitted sequence is therefore identical to baseline greedy decoding by construction, leaving benchmark accuracy unchanged.

Inference Speedup on GPU and On-Device

The models ship with day-one support for llama.cpp (using experimental Metal kernels) and SGLang. Measurements were taken on an M4 Max MacBook Pro with FP16 GGUF weights (llama.cpp) and a single H100 80GB in BF16 (SGLang). Both use DSpark block size 9, batch size 1, and temperature 0.

LFM2.5-2.6B Results

| Dataset | Acceptance (/10) | H100 Speedup | M4 Max Speedup |

|---------|------------------|--------------|----------------|

| MATH500 | 5.42 | 3.06x (326→1000 tok/s) | 2.25x (61→137 tok/s) |

| HumanEval | 4.54 | 2.56x (326→835 tok/s) | 2.63x (61→161 tok/s) |

| MBPP | 4.71 | 2.64x (326→861 tok/s) | 2.11x (62→132 tok/s) |

| GSM8K | 4.32 | 2.22x (312→693 tok/s) | 2.36x (60→143 tok/s) |

| MT-Bench | 5.07 | 2.87x (325→933 tok/s) | 1.99x (62→123 tok/s) |

| Mean | 4.81 | 2.67x | 2.27x |

Function-calling latency is reduced by 57% on average in multi-tool scenarios.

LFM2.5-1.2B-Instruct Results

Mean acceptance 5.02/10. Average speedup reaches 2.10x on H100 (656→1384 tok/s) and 2.54x on M4 Max (138→350 tok/s). Acceptance rates show higher variance across datasets.

LFM2.5-8B-A1B Results

Mean acceptance improves to 6.95/10. GPU speedup averages 2.54x, but on-device improvement is limited to 1.18x due to current MoE implementation constraints in llama.cpp’s Metal backend.

How to Use LFM2.5-DSpark

Running DSpark draft models with SGLang requires a build with DSpark support for LFM targets (PR #31041). The integration is open-sourced and ready for immediate use.

Source