LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

2026-08-20 · Hugging Face

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

Model Overview

LFM2.5-VL-3B is LiquidAI's most capable vision-language model that can run on your own hardware. It understands both documents and screens, performs object grounding, and supports tool calling. The model answers directly instead of performing step-by-step reasoning, which keeps responses fast for real-time and on-device applications.

It extends previous vision-language capabilities with four major improvements:

  • Screen/UI Understanding: Strong understanding of digital screens across different devices.
  • Grounding: Improved object grounding and detection using natural language queries.
  • Multi-image Input: Enhanced reasoning across multiple images.
  • Function Calling: Significantly stronger function calling performance in both text-only and vision-text scenarios.

Training Approach

LFM2.5-VL-3B pairs a SigLIP2 400M NaFlex vision encoder with the same pre-trained backbone used in the LFM2.5-2.6B text model. It was pre-trained on approximately 34T tokens, using 4x more vision data than previous versions. The vision data was drawn from curated and synthetic image-caption, OCR, grounding, and instruction-following datasets.

To better support non-Latin scripts, the vocabulary was doubled to 128K by extending the existing tokenizer in place rather than retraining from scratch.

Post-training consists of two stages: First, supervised fine-tuning (SFT) that incorporates knowledge distillation from a larger teacher model along with Antidoom training. Second, multi-reward reinforcement learning (RL).

Benchmark Results

LFM2.5-VL-3B was evaluated on both vision and text benchmarks. It leads its size class on real-world image tasks while maintaining strong performance reading digital content such as documents, charts, and on-screen UI elements. All evaluations used non-reasoning mode with prompts instructing the model to answer directly.

Vision Benchmarks

The model was tested on multilingual visual comprehension, instruction following, visual math and scientific reasoning, document understanding, object detection, multi-image understanding, screen understanding, and hallucination benchmarks.

It achieves leading scores in its class on tasks including:

  • MMStar: 63.3
  • RealWorldQA: 73.1
  • MMBench (dev EN): 81.0
  • MathVista (mini): 68.5
  • ChartQA: 81.3
  • DocVQA: 91.1
  • RefCOCO-avg: 87.9
  • ScreenSpot-v2 (Desktop: 78.7, Mobile: 81.2, Web: 82.2)

The overall average score across vision benchmarks is 69.4, outperforming the previous LFM2-VL-3B (57.2) and several larger models in key areas such as grounding, multi-image understanding, and GUI tasks.

Text-only Benchmarks

On text benchmarks focused on instruction following and tool use, LFM2.5-VL-3B shows clear gains. Instruction following improves across IFEval, IFBench, and Multi-IF. Tool use sees sharp improvement on ToolSandbox (59.5) and BFCL V4 (32.5), placing it on par with Gemma-4-E2B-it and Qwen3.5-2B.

These results confirm that LFM2.5-VL-3B is a strong general-purpose vision-language model. It handles everyday tasks such as captioning, visual question answering, and document understanding particularly well, while excelling at grounding objects, reading screens and documents, and calling tools.

Inference Speed on CPU, GPU and Mobile

LFM2.5-VL-3B includes day-one support across major inference ecosystems: llama.cpp, MLX, vLLM, SGLang, and ONNX.

On-device Performance:

  • 228 tokens/s on M5 Max
  • 116 tokens/s on Ryzen AI Max+ 395
  • Fits in approximately 3 GB of memory
  • Reaches 20 tokens/s on a Galaxy S26 Ultra, enabling fully on-device execution

GPU Performance:

The model maintains consistently low latency and is the fastest among tested models on multi-frame inputs. It also delivers the highest output throughput, reaching approximately 11K tokens per second at high concurrency. This represents roughly 2× the performance of larger 4B-class models and surpasses many smaller 2B-class models, enabling nearly 1B output tokens per day on a single H100.

How to Use LFM2.5-VL-3B

Use LFM2.5-VL-3B when you need on-device intelligence for high-volume workloads. Install the latest version of transformers for compatibility.

Source