AI Potluck
Back to Gap Map Model components / Multimodal models

LFM-VL

Liquid AI
open weights / Overall score: 2.3

LFM-VL is Liquid AI's family of small vision-language models for on-device use, from 450M to 3B parameters, built on its LFM2 hybrid language models. The models read text and images, including several images at once, and handle OCR, document comprehension, object grounding and screen understanding. Liquid AI ships them with GGUF, ONNX and MLX builds for local inference.

Openness

3 high confidence
3.0
weights
open(model.safetensors on the Hub for every LFM2.5-VL size, ungated)
data
described(the release post describes a mixture of curated and synthetic caption, OCR, grounding and instruction data
code
partial(LoRA fine-tuning notebooks in the Liquid4All cookbook
license
LFM-Open-License-1.0(Apache-2.0-based

LFM2.5-VL's weights can be used commercially by companies with under 10 million dollars in annual revenue; larger companies need a paid license from Liquid AI. Liquid publishes fine-tuning notebooks but neither its training data nor its training code. LFM-Open-License-1.0 allows commercial use only within a bound, so the license tier is use_bounded.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across Liquid AI's LFM2-VL and LFM2.5-VL checkpoints, including its own GGUF, ONNX and MLX builds. More than a third of it is those quantized builds, which suit local apps.

Capability

2 medium confidence
2.0

LFM2.5-VL reads full pages, charts and screens and answers over several images, and the 3B model reports strong chart and OCR results for its size. Liquid demonstrates video captioning but reports no video benchmark, so it sits level with Moondream rather than with MiniCPM-V.

Verified 2026-09-27