Liquid AI ships LFM2.5-VL-3B, a 3B vision model built for the edge

Liquid AI has released LFM2.5-VL-3B, a roughly three billion parameter vision-language model designed to run on phones, laptops, and other edge hardware instead of in a data center. The model pairs the LFM2.5-2.6B text backbone with a SigLIP2 NaFlex image encoder. Unlike many recent models, it does not reason step by step. It answers directly, which keeps latency low for real-time and on-device use, and it handles a 32,000 token context. What it is built to do Liquid says the release adds fo

2 min
Liquid AI ships LFM2.5-VL-3B, a 3B vision model built for the edge

Liquid AI has released LFM2.5-VL-3B, a roughly three billion parameter vision-language model designed to run on phones, laptops, and other edge hardware instead of in a data center.

The model pairs the LFM2.5-2.6B text backbone with a SigLIP2 NaFlex image encoder. Unlike many recent models, it does not reason step by step. It answers directly, which keeps latency low for real-time and on-device use, and it handles a 32,000 token context.

What it is built to do

Illustration of a phone screen with an AI marking objects on it

Liquid says the release adds four capabilities over its earlier LFM2-VL-3B: reading digital screens across mobile, web, and desktop; grounding objects to coordinates from a text query; taking multiple images as input; and triggering actions from text or image prompts. Those are the skills that let a model act as an interface layer between what a device sees and what software should do next.

How it scores

On ScreenSpot-v2, which measures how well a model understands UI screens, it averages 80.7, ahead of Google's Gemma-4-E4B at 51.2 and Qwen 3.5 4B at 78.5, and close to the larger InternVL-3.5-4B at 84.1, according to Liquid's own benchmarks. It reaches 87.9 on RefCOCO-avg for grounding, up from 57.1 in the prior release, and its function-calling scores more than double on ToolSandbox and BFCL v4.

Liquid reports the model runs in about three gigabytes of memory, decoding 228 tokens per second on an M5 Max and 20 tokens per second on a Galaxy S26 Ultra, so it can run fully offline on consumer hardware.

Availability

LFM2.5-VL-3B is available on Hugging Face and Liquid's Playground under the LFM Open License v1.0, which allows free commercial use only for companies under 10 million dollars in annual revenue. The release follows a wave of small, efficient vision models aimed at on-device assistants and private inference.

Sources

Liquid AI blog, LFM2.5-VL-3B: https://www.liquid.ai/blog/lfm2-5-vl-3b

Hugging Face blog: https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b

Liquid Docs: https://docs.liquid.ai/lfm/models/lfm25-vl-3b

MarktechPost: https://www.marktechpost.com/2026/08/13/liquid-ai-lfm2-5-vl-3b-on-device-vision-language-model

Written by

More to read

  • Speech-to-Text Serving in Production: Comparing Faster-Whisper, Moonshine, SenseVoice, and NeMo Canary Architecture, Streaming Latency, and GPU Economics

    In conversational voice AI and real-time agentic workflows, the speech-to-text (STT) layer sets the hard lower bound on system responsiveness. Human conversational cadence expects turn-taking latencies between 200ms and 500ms. When an AI pipeline must accommodate downstream large language model (LLM) time-to-first-token generation (100ms to 250ms) and text-to-speech (TTS) audio synthesis (100ms to 200ms), the automatic speech recognition (ASR) stage cannot exceed 100ms to 150ms of processing ove

    1 min
  • Writer Releases Palmyra X6 Flagship Agentic Model with Rebuilt Enterprise Agent Harness

    Enterprise generative AI platform Writer has launched Palmyra X6, its new flagship agentic foundation model, alongside a rebuilt runtime harness engineered for multi-step workflow execution and governance. The model release introduces substantial latency and efficiency improvements over previous Palmyra iterations, cutting inference costs by 52% while accelerating output generation by 48%. Writer reported average generation speeds of 82 tokens per second and a mean task completion time of 26 se

    1 min
  • River AI Secures .1B Led by General Catalyst to Build Open-Weight Model Infrastructure

    River AI, an artificial intelligence startup founded by former xAI co-founder Igor Babuschkin, has secured $1.1 billion in early-stage funding to build an open-weight model stack and decentralized AI infrastructure platform. The financing round was led jointly by General Catalyst and public benefit corporation AMP PBC, with strategic participation from NVIDIA, AMD Ventures, Y Combinator, and Singapore sovereign fund Temasek. Babuschkin, whose prior engineering background spans OpenAI, Google De

    1 min