Liquid AI Unveils LFM2.5-VL-3B: Edge Vision-Language Model with Enhanced Screen Understanding and Tool Use
Liquid AI has introduced LFM2.5-VL-3B, a vision-language model designed for edge deployment that improves understanding of screens and user interfaces, grounding, multi-image input, and function calling. This model integrates a SigLIP2 400M NaFlex vision encoder with the LFM2.5-2.6B backbone, trained on an impressive 34 trillion tokens and utilizing four times the vision data compared to previous iterations. The post-training phase involved supervised fine-tuning, knowledge distillation, Antidoom training, and multi-reward reinforcement learning. Benchmark tests reveal its strong performance in multilingual visual comprehension, document understanding, and object detection, rivaling Gemma-4-E2B and Qwen3.5-2B in tool utilization. It processes at 228 tokens/s on M5 Max, 116 tokens/s on Ryzen AI Max+ 395, and 20 tokens/s on Galaxy S26 Ultra, achieving a high concurrency output of 11K tokens/s on GPU. The model is accessible on Hugging Face and supports llama.cpp, MLX, vLLM, SGLang, and ONNX.
Key facts
- LFM2.5-VL-3B is a vision-language model for edge deployment.
- It improves screen/UI understanding, grounding, multi-image input, and function calling.
- Uses SigLIP2 400M NaFlex vision encoder with LFM2.5-2.6B backbone.
- Pre-trained on ~34T tokens with 4x more vision data.
- Post-training includes SFT with knowledge distillation and multi-reward RL.
- Leads size class on real-world image tasks and excels in document/screen reading.
- On par with Gemma-4-E2B and Qwen3.5-2B on tool use benchmarks.
- Inference speeds: 228 tok/s on M5 Max, 116 tok/s on Ryzen AI Max+ 395, 20 tok/s on Galaxy S26 Ultra.
- GPU throughput ~11K tokens/s, 2x larger 4B-class models, ~1B tokens/day on H100.
- Available on Hugging Face with support for llama.cpp, MLX, vLLM, SGLang, ONNX.
- Browser demo via WebGPU.
- Part of Liquid AI's vision of AI that runs anywhere.
Entities
Institutions
- Liquid AI
- Hugging Face