AI Potluck
Back to Gap Map Model components / Multimodal models

Perception LM

Meta
restricted / Overall score: 2.4

Perception LM (PLM) is Meta FAIR's family of vision-language models for detailed image and video understanding, in 1B, 3B and 8B sizes that pair the Perception Encoder with Llama 3.2 and Llama 3.1 language models. It answers questions about images, documents, charts and video, including fine-grained and spatio-temporally grounded questions about what happens in a clip. Meta released it with its human-labeled video training data and the PLM-VideoBench benchmark.

The GitHub repository, facebookresearch/perception_models, is not declared: it also releases the Perception Encoder vision models, a separate entry under Apache-2.0 whose license does not cover the PLM checkpoints.

Openness

2 high confidence
2.0
weights
open(safetensors for all three sizes on the Hub, behind a manual access request)
data
partial(Meta releases its human-labeled PLM-Video-Human set (CC-BY-4.0) and the synthetic PLM-Image-Auto and PLM-Video-Auto sets (Llama 3.2 license)
code
open(the training script, with stage 1, 2 and 3 configs for each of the 1B, 3B and 8B models)
license
FAIR-Noncommercial-Research(the FAIR Noncommercial Research License on every checkpoint, which bars any commercial use of the models or their outputs)

Meta publishes the weights, its own training data and the training code, but the checkpoints carry a research license that forbids any commercial use of the models or their outputs, and each download needs Meta's approval. The open recipe does not lift that restriction.

Adoption

1 medium confidence
1.0

Hugging Face downloads over the trailing 30 days, summed across the 1B, 3B and 8B checkpoints, most of them for the 1B model.

Capability

3 high confidence
3.0

Perception LM reads documents, charts and video and answers fine-grained and grounded questions about what happens in a clip, level with MiniCPM-V.

Verified 2026-09-28