Perception LM
MetaPerception LM (PLM) is Meta FAIR's family of vision-language models for detailed image and video understanding, in 1B, 3B and 8B sizes that pair the Perception Encoder with Llama 3.2 and Llama 3.1 language models. It answers questions about images, documents, charts and video, including fine-grained and spatio-temporally grounded questions about what happens in a clip. Meta released it with its human-labeled video training data and the PLM-VideoBench benchmark.
The GitHub repository, facebookresearch/perception_models, is not declared: it also releases the Perception Encoder vision models, a separate entry under Apache-2.0 whose license does not cover the PLM checkpoints.
Openness
2 high confidence- weights
- open(safetensors for all three sizes on the Hub, behind a manual access request)
- data
- partial(Meta releases its human-labeled PLM-Video-Human set (CC-BY-4.0) and the synthetic PLM-Image-Auto and PLM-Video-Auto sets (Llama 3.2 license)
- code
- open(the training script, with stage 1, 2 and 3 configs for each of the 1B, 3B and 8B models)
- license
- FAIR-Noncommercial-Research(the FAIR Noncommercial Research License on every checkpoint, which bars any commercial use of the models or their outputs)
Meta publishes the weights, its own training data and the training code, but the checkpoints carry a research license that forbids any commercial use of the models or their outputs, and each download needs Meta's approval. The open recipe does not lift that restriction.
- https://huggingface.co/api/datasets/facebook/PLM-Image-Auto recorded 2026-09-28
Hub record for facebook/PLM-Image-Auto: public, not gated, license llama3.2.
- https://huggingface.co/api/datasets/facebook/PLM-Video-Auto recorded 2026-09-28
Hub record for facebook/PLM-Video-Auto: public, not gated, license llama3.2.
- https://huggingface.co/api/datasets/facebook/PLM-Video-Human recorded 2026-09-28
Hub record for facebook/PLM-Video-Human: public, not gated, license cc-by-4.0.
- https://huggingface.co/api/models/facebook/Perception-LM-1B recorded 2026-09-28
Hub record for Perception-LM-1B: gated "manual", license_name fair-noncommercial-research, model.safetensors.
- https://huggingface.co/api/models/facebook/Perception-LM-3B recorded 2026-09-28
Hub record for Perception-LM-3B: gated "manual", license_name fair-noncommercial-research, two safetensors shards.
- https://huggingface.co/api/models/facebook/Perception-LM-8B recorded 2026-09-28
Hub record for Perception-LM-8B: gated "manual", license_name fair-noncommercial-research, four safetensors shards; the gate prompt is the FAIR Noncommercial Research License.
- https://raw.githubusercontent.com/facebookresearch/perception_models/main/apps/plm/configs/datasets.yaml recorded 2026-09-28
Dataset registry: entries only for dummy_image, dummy_multi_image, dummy_video, dummy_text and the other apps/plm/dummy_datasets sets.
- https://raw.githubusercontent.com/facebookresearch/perception_models/main/apps/plm/configs/stage_3/plm_8b.yaml recorded 2026-09-28
Stage-3 config for the 8B model: "datamix: <Please consider using data split listed in Table A2 of our paper https://arxiv.org/pdf/2504.13180. ...>".
- https://raw.githubusercontent.com/facebookresearch/perception_models/main/apps/plm/docs/training.md recorded 2026-09-28
Training guide: "We provide configurations to run warm-up and sft to facilitate reproducibility of PLM training"; torchrun with apps/plm/configs/stage_3/plm_3b.yaml; examples use dummy-datasets.
- https://raw.githubusercontent.com/facebookresearch/perception_models/main/LICENSE.PLM recorded 2026-09-28
FAIR Noncommercial Research License, last updated April 17, 2025: "You will not use the Research Materials or any outputs or results of the Research Materials in connection with any commercial uses or for any uses other than Noncommercial Research Uses".
- https://ungh.cc/repos/facebookresearch/perception_models/files/main recorded 2026-09-28
Repository tree: apps/plm/configs/stage_1, stage_2 and stage_3, each with plm_1b.yaml, plm_3b.yaml and plm_8b.yaml, beside apps/plm/configs/datasets.yaml.
Adoption
1 medium confidenceHugging Face downloads over the trailing 30 days, summed across the 1B, 3B and 8B checkpoints, most of them for the 1B model.
- https://huggingface.co/api/models?author=facebook&search=Perception-LM&limit=100 recorded 2026-09-28
Three facebook repositories matching Perception-LM: 1B with 1,265, 8B with 202 and 3B with 108 downloads in the trailing 30 days, 1,575 in total.
Capability
3 high confidencePerception LM reads documents, charts and video and answers fine-grained and grounded questions about what happens in a clip, level with MiniCPM-V.
- https://arxiv.org/abs/2504.13180 recorded 2026-09-28
PerceptionLM abstract: a model "for transparent research in image and video understanding"; "we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions", and PLM-VideoBench.
- https://raw.githubusercontent.com/facebookresearch/perception_models/main/apps/plm/docs/training.md recorded 2026-09-28
Training guide: "The repo also support text-only, multi-image, image-region, video-region-caption (RCap), video-region-temporal-localization (RTLoc) and video-region-dense-captioning (RDCap) tasks."
- https://raw.githubusercontent.com/facebookresearch/perception_models/main/README.md recorded 2026-09-28
PLM section: 1B, 3B and 8B models on Llama-3.2-1B/3B-Instruct and Llama-3.1-8B-Instruct; image results table with DocVQA 94.6 and ChartQA 85.5 for PLM8B; video results table with VideoMME 58.3 for PLM8B.
Verified 2026-09-28