AI Potluck
Back to Gap Map Model components / Multimodal models

Florence-2

Microsoft
open weights / Overall score: 1.9

Florence-2 is Microsoft's prompt-based vision foundation model, in 0.23B and 0.77B sizes, that performs captioning, object detection, phrase grounding, dense region captioning and OCR from short task tokens. It was trained on FLD-5B, a set of 5.4 billion annotations over 126 million images, and fine-tuned variants cover downstream tasks.

Openness

3 high confidence
3.0
weights
open(safetensors and PyTorch weights on the Hub, ungated)
data
described(FLD-5B and the fine-tuning task collection are named
code
partial(modeling, processing and an inference notebook in the Hub repository
license
MIT(OSI

Florence-2 ships under the MIT license with the code needed to run it. The FLD-5B training set is described but not released, and no training code is published.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days, summed across Microsoft's four Florence-2 checkpoints. Most of it is the base model, which is widely used as a captioning and detection step in image pipelines.

Capability

1 high confidence
1.0

Florence-2 performs a fixed set of prompted tasks on a single image, including captions, detection, grounding and OCR, with strong results for its size. It does not hold a conversation, read documents as a whole or take several images, two steps below MiniCPM-V.

Verified 2026-09-27