Florence-2
MicrosoftFlorence-2 is Microsoft's prompt-based vision foundation model, in 0.23B and 0.77B sizes, that performs captioning, object detection, phrase grounding, dense region captioning and OCR from short task tokens. It was trained on FLD-5B, a set of 5.4 billion annotations over 126 million images, and fine-tuned variants cover downstream tasks.
Openness
3 high confidence- weights
- open(safetensors and PyTorch weights on the Hub, ungated)
- data
- described(FLD-5B and the fine-tuning task collection are named
- code
- partial(modeling, processing and an inference notebook in the Hub repository
- license
- MIT(OSI
Florence-2 ships under the MIT license with the code needed to run it. The FLD-5B training set is described but not released, and no training code is published.
- https://huggingface.co/api/datasets?search=FLD-5B&limit=100 recorded 2026-09-27
An empty list: no FLD-5B dataset on the Hub
- https://huggingface.co/api/models/microsoft/Florence-2-large?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for Florence-2-large: gated false, model.safetensors and pytorch_model.bin, modeling and processing code, card license mit
- https://huggingface.co/microsoft/Florence-2-large/raw/main/LICENSE recorded 2026-09-27
"MIT License / Copyright (c) Microsoft Corporation"
- https://huggingface.co/microsoft/Florence-2-large/raw/main/README.md recorded 2026-09-27
"It leverages our FLD-5B dataset, containing 5.4 billion annotations across 126 million images"; the -ft models are "Finetuned model on a colletion of downstream tasks"
Adoption
4 high confidenceHugging Face downloads over the trailing 30 days, summed across Microsoft's four Florence-2 checkpoints. Most of it is the base model, which is widely used as a captioning and detection step in image pipelines.
- https://huggingface.co/api/models?author=microsoft&search=Florence&sort=downloads&direction=-1&limit=100 recorded 2026-09-27
Four Florence-2 checkpoints with 3,790,676 downloads in the trailing 30 days; Florence-2-base 3,046,535
Capability
1 high confidenceFlorence-2 performs a fixed set of prompted tasks on a single image, including captions, detection, grounding and OCR, with strong results for its size. It does not hold a conversation, read documents as a whole or take several images, two steps below MiniCPM-V.
- https://huggingface.co/microsoft/Florence-2-large/raw/main/README.md recorded 2026-09-27
"Florence-2 can interpret simple text prompts to perform tasks like captioning, object detection, and segmentation"; zero-shot COCO Cap. 135.6, RefCOCO 56.3; fine-tuned VQAv2 81.7 against PaLI-X 86.0
Verified 2026-09-27