Deploy local agents everywhere with LFM2.5-2.6B

2026-08-20 · Hugging Face

Deploy local agents everywhere with LFM2.5-2.6B

Liquid AI has introduced LFM2.5-2.6B, a 2.6B parameter model built specifically to power capable agents entirely on-device. It supports tool calling and multi-step workflows while remaining small and fast enough to run on everyday hardware ranging from laptops to phones. This enables developers to deploy agents anywhere, keep data private on the device, and scale usage without cloud inference costs.

Key Strengths

  • Best-in-class agent performance: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks.
  • Agentic Reinforcement Learning: Trained inside popular agentic harnesses for improved compatibility.
  • Efficient inference: 220 tokens/s on Apple M5 Max, 113 tokens/s on AMD Ryzen CPU, using under 2.5 GB of memory.

How the Model Was Built

The model was pre-trained on approximately 34T tokens, with a mid-training phase that extends the context window to 128K. Post-training transforms the base model into a reliable agent through four stages:

1. Supervised Fine-Tuning (SFT): Two rounds of SFT, heavily weighted toward agentic data including tool use, web search, and harness trajectories.

2. Teacher Specialization: One specialist teacher trained per domain (math, code, tool use, and others).

3. Multi-domain On-Policy Distillation (MOPD): Distills the specialist teachers into a single student model.

4. Agentic Reinforcement Learning (Agentic RL): Performs multi-turn RL inside real agent harnesses, enabling the model to work across different tools, system prompts, and multi-turn environments.

The Agentic RL Pipeline

The pipeline separates model optimization, inference, and environment execution:

  • The Training Engine optimizes the model.
  • The Rollout Engine generates actions using the latest policy.
  • The RL framework orchestrates the loop by launching rollouts, collecting trajectories and rewards, and updating the model.

Actions are executed within a Sandbox Service. The Blackbox Harness (e.g., OpenClaw or Hermes Agent) hosts the agent and coordinates with the task environment. The Harness Proxy treats agentic harnesses as black boxes without modification while transparently capturing token-level trajectories for RL sample reconstruction and validation.

Benchmark Results

LFM2.5-2.6B was evaluated against models up to 4x its size on STEM, instruction following, tool use, and agentic tasks. Despite being the smallest model in the comparison group, it competes with and frequently outperforms larger counterparts.

Selected Benchmark Highlights (full table available in original release):

  • Tops all instruction-following benchmarks (IFBench, Multi-IF, IFStruct).
  • Leads most tool-use benchmarks (ToolSandbox, BFCLv4 second only to 9.7B Qwen).
  • Strong performance on agentic tasks (Claw-Eval, PinchBench, BrowseComp+).
  • Competitive on knowledge (AA Omniscience) and math (AIME25).
  • Coding remains an area where larger models hold a clearer advantage.

Inference Performance

LFM2.5-2.6B ships with day-one support across the inference ecosystem, including llama.cpp, MLX, vLLM, SGLang, and ONNX.

CPU Inference: Thanks to its efficient LFM2 architecture, it achieves 220 tokens/s on an Apple M5 Max and 113 tokens/s on Ryzen AI Max+ 395. Even at 30 tokens/s, it can run capable agents on phones.

GPU Inference: The fastest in its size class, reaching nearly 15K output tokens per second at high concurrency, equating to roughly 1.3B tokens per day on a single H100.

How to Use LFM2.5-2.6B


pip install -U transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "LiquidAI/LFM2.5-2.6B"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="bfloat16",
    # attn_implementation="flash_attention_2"
)
tokenizer = AutoTokenizer.from_pretrained(model_id)

prompt = "What is C. elegans?"
input_ids = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True,
    return_tensors="pt"
).to(model.device)

output = model.generate(
    input_ids,
    do_sample=True,
    temperature=0.2,
    top_k=80,
    repetition_penalty=1.05,
    max_new_tokens=512,
)
print(tokenizer.decode(output[0], skip_special_tokens=False))

Getting Started

Both LFM2.5-2.6B and LFM2.5-2.6B-Base are available on Hugging Face. A browser-based WebGPU demo is provided, along with guides for integrating with local agent harnesses such as OpenClaw, Hermes Agent, and Pi.

LFM2.5-2.6B advances the vision of AI that can run anywhere — private, cost-effective, and universally deployable.

Source