Jina Embeddings
Jina AIJina AI's text embedding line, from the v2 English and bilingual encoders through v3's task-LoRA multilingual model to the v5-text family released in early 2026. v5-text-small is a 677M Qwen3-0.6B-Base distillation of Qwen3-Embedding-4B scoring 67.7 on MMTEB across 119+ languages at 32K tokens, with retrieval, text-matching, clustering and classification adapters shipped separately for vLLM and ONNX. The weights are published on the Hub, but the current generation is non-commercial.
Added 2026-09-11 for the embeddings_retrieval roster, as the embedding sibling of jina-reranker. Scope is the TEXT line in the jina-embeddings namespace of the jinaai account: the v1 and v2 encoders, v3 and the v5-text tier. The v4 universal multimodal retriever and the v5-omni tier are excluded on every axis under the category's rule that a vision-language line is a separate product from the text line; that is why the Qwen Research License v4 inherits from Qwen2.5-VL-3B is not part of this record's licence compound. The vendor's own GGUF and vLLM re-publishes are the same weights in another container, so they belong to this line but are not added to the download sum, which would double-count them. jina-reranker covers the rerankers; jina-clip, jina-code-embeddings and jina-colbert-v2 are separate lines and are not counted here. The licence is the thing to read carefully, as with the reranker: v3 and v5-text are cc-by-nc-4.0 while the v1 and v2 checkpoints remain Apache-2.0, and the most restrictive distributed SKU governs. Verified 2026-09-11 via the model cards and the Hub API.
Openness
2 high confidence- weights
- open(ungated safetensors on the Hub across the v2, v3 and v5-text lines)
- data
- closed(no training corpus or construction scripts released
- code
- partial(trust_remote_code inference modules and ONNX/GGUF conversions ship with the repositories
- license
- Apache-2.0(OSI, and only the older v1 and v2 encoders)+CC-BY-NC-4.0(non-commercial
The current text SKUs are non-commercial: v3 and the v5-text line read cc-by-nc-4.0, while Apache-2.0 covers only the superseded v1 and v2 encoders, so the compound resolves on the non-commercial half at commercial_forbidden - whose one question, does the license permit commercial use at all, the cards answer no, directing commercial use to a sales conversation. The compound covers the text line and nothing else: the v4 universal multimodal retriever, which inherits the Qwen Research License from Qwen2.5-VL-3B, is excluded from this record on every axis under the category's vision-language rule, so no unmapped licence is being passed over here.
- https://huggingface.co/jinaai/jina-embeddings-v3/raw/main/README.md recorded 2026-09-11
Front matter reads license: cc-by-nc-4.0, and the License section says the model is licensed under CC BY-NC 4.0 beyond the AWS and Azure listings and directs commercial usage inquiries to contact sales. Training is described against the arXiv report with no corpus released; the usage sections ship inference code only, via trust_remote_code and ONNX.
- https://huggingface.co/api/models/jinaai/jina-embeddings-v3 recorded 2026-09-11
license cc-by-nc-4.0 in cardData; gated: false; private: false; model.safetensors and an onnx/ folder among the files.
- https://huggingface.co/jinaai/jina-embeddings-v5-text-small/raw/main/README.md recorded 2026-09-11
Front matter reads license: cc-by-nc-4.0 and the License section states the model is licensed under CC BY-NC 4.0 with commercial use routed to sales. Training and evaluation detail is deferred to arXiv 2602.15547 with no corpus released, and the usage sections cover inference only.
- https://huggingface.co/api/models/jinaai/jina-embeddings-v5-text-small recorded 2026-09-11
license cc-by-nc-4.0 in cardData; gated: false; private: false; task adapter safetensors present, confirming the v5-text line ships on the same non-commercial terms as v3.
- https://huggingface.co/jinaai/jina-embeddings-v4/raw/main/README.md recorded 2026-09-11
License section reads "This model was initially released under cc-by-nc-4.0 due to an error. The correct license is the Qwen Research License, as this model is derived from Qwen-2.5-VL-3B which is governed by that license."
Adoption
4 high confidence4,687,946 downloads in the trailing 30 days across the jina-embeddings TEXT checkpoints. The v4 universal multimodal retriever and the v5-omni tier are excluded under the category's rule that a vision-language line is a separate product from the text line; counting them gives 5,402,605 and the same band. Vendor GGUF and vLLM re-publishes are excluded, as are the separate jina-clip, jina-code-embeddings and jina-colbert lines.
- https://huggingface.co/api/models/jinaai/jina-embeddings-v3 recorded 2026-09-11
downloads: 2151116; gated: false; private: false for jinaai/jina-embeddings-v3.
- https://huggingface.co/api/models/jinaai/jina-embeddings-v5-text-small recorded 2026-09-11
downloads: 300461 for jinaai/jina-embeddings-v5-text-small.
- https://huggingface.co/api/models/jinaai/jina-embeddings-v4 recorded 2026-09-11
downloads: 324909 for jinaai/jina-embeddings-v4.
Capability
4 high confidenceRung 4, competitive frontier. Current generation, multilingual and long-context, with published numbers on the same instrument as the anchor: 67.7 MMTEB against harrier-oss at 74.3, so one band below the leader rather than level with it. The card's own claim is narrower than the leaderboard - the highest MMTEB among multilingual embedding models under 1B parameters - which is a frontier position in its size class rather than the top of the table.
- https://huggingface.co/jinaai/jina-embeddings-v5-text-small/raw/main/README.md recorded 2026-09-11
Card states jina-embeddings-v5-text-small scores 71.7 average on MTEB English v2 and 67.7 on MMTEB with 677M parameters, the highest among multilingual embedding models under 1B, supporting 119+ languages at up to 32K tokens and built on Qwen3-0.6B-Base by distillation from Qwen3-Embedding-4B.
Verified 2026-09-11