Best Embedding Models for Semantic Search in 2026: Qwen3-VL-Embedding-8B and Qwen3-VL-Embedding-2B

Rankings 2026-09-28 Last updated 2026-09-28 15 min read By Q4KM

Quick Answer

Qwen3-VL-Embedding-8B (Qwen) is the top pick as of September 2026 because its mean main score of 62.2 across 100 published MTEB task results is the highest among the listed candidates, measured as of September 25, 2026 [3]. In order, the ranking is Qwen3-VL-Embedding-8B (Qwen), Qwen3-VL-Embedding-2B (Qwen), Nemotron-3-Embed-1B-BF16 (nvidia), jina-embeddings-v5-text-small (jinaai), harrier-oss-v1-0.6b (microsoft), bge-small-en-v1.5 (BAAI), paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers), and all-MiniLM-L6-v2 (sentence-transformers) [3].

Key Takeaways

How do local embedding models compare on specifications and semantic search benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
Qwen3-VL-Embedding-8B [8] Qwen [8] 8.1B [8] BF16 weights: 16.3 GB [8] 2026-01-07 [8] Apache-2.0 [8] MTEB mean main score: 62.2 across 100 task results (2026-09-25) [3]
Qwen3-VL-Embedding-2B [10] Qwen [10] 2.1B [10] BF16 weights: 4.3 GB [10] 2026-01-07 [10] Apache-2.0 [10] MTEB mean main score: 58.2 across 100 task results (2026-09-25) [3]
Nemotron-3-Embed-1B-BF16 [11] nvidia [11] 1.1B [11] BF16 weights: 2.3 GB [11] 2026-07-14 [11] openmdw-1.1 [11] MTEB mean main score: 54.2 across 100 task results (2026-09-25) [3]
jina-embeddings-v5-text-small [14] jinaai [14] 596M [14] BF16 weights: 1.4 GB [14] 2026-01-22 [14] cc-by-nc-4.0 [14] MTEB mean main score: 45.2 across 100 task results (2026-09-25) [3]
harrier-oss-v1-0.6b [13] microsoft [13] 596M [13] BF16 weights: 1.2 GB [13] 2026-03-30 [13] MIT [13] MTEB mean main score: 33.5 across 100 task results (2026-09-25) [3]
bge-small-en-v1.5 — established pick [4] BAAI [4] 33M [4] F32 weights: 0.1 GB [4] 2023-09-12 [4] MIT [4] MTEB mean main score: 20.5 across 100 task results (2026-09-25) [3]
paraphrase-multilingual-MiniLM-L12-v2 — established pick [6] sentence-transformers [6] 118M [6] F32 weights: 0.5 GB [6] 2022-03-02 [6] Apache-2.0 [6] MTEB mean main score: 8.3 across 100 task results (2026-09-25) [3]
all-MiniLM-L6-v2 — established pick [1] sentence-transformers [1] 23M [1] F32 weights: 0.1 GB [1] 2022-03-02 [1] Apache-2.0 [1] MTEB mean main score: 6.3 across 100 task results (2026-09-25) [3]

Which local embedding models should you choose for semantic search?

1. Qwen3-VL-Embedding-8B

Qwen3-VL-Embedding-8B by Qwen ranks first here because its MTEB mean main score of 62.2 is the highest among the listed candidates across 100 published task results, as of September 25, 2026.[3] Choose it when benchmark performance is your priority for a local semantic-search deployment.

Qwen published the model on January 7, 2026, with 8.1 billion parameters, a 256K-token context length and an Apache-2.0 license.[8] Local hardware planning starts with its 16.3 GB BF16 weight footprint.[8] Quantized weight storage would be approximately 4.05 GB at 4-bit precision—estimated (params x 0.5 bytes), using the cited 8.1 billion parameters.[8] Neither weight figure is a complete RAM or VRAM budget; allow additional memory for execution.

The benchmark caveat matters: the reported score averages 100 MTEB task results, rather than providing a retrieval-only result for your application.[3] Evaluate search relevance on your own queries and documents before committing hardware to deployment.

2. Qwen3-VL-Embedding-2B

Qwen3-VL-Embedding-2B by Qwen [10] ranks second among the admitted candidates on the shared MTEB results: its mean main score is 58.2 across 100 published task results as of 2026-09-25, behind Qwen’s Qwen3-VL-Embedding-8B at 62.2 [3]. Choose it when you want a smaller local deployment than its sibling: their published BF16 weight sizes are 4.3 GB [10] and 16.3 GB [8], respectively.

Published on 2026-01-07, the model has 2.1B parameters, a 256K-token context length and an Apache-2.0 license [10]. For local hardware planning, the BF16 weights alone occupy 4.3 GB [10]. A quantized weight-only budget would be 1.05 GB, estimated (params x 0.5 bytes) from the cited 2.1B parameters [10]. Allow additional RAM or VRAM for execution; neither figure establishes total memory requirements.

Use it as a semantic-search candidate when weight storage constrains your deployment. The caveat is benchmark scope: the reported score averages MTEB task results [3], so evaluate retrieval quality on your own queries and documents before committing.

3. Nemotron-3-Embed-1B-BF16

Nemotron-3-Embed-1B-BF16 by nvidia ranks third here because its mean main score of 54.2 across 100 published MTEB task results places it behind the Qwen candidates and ahead of the remaining candidates, as of September 25, 2026.[3] Its position reflects benchmark performance, rather than download popularity.

Published on July 14, 2026, the model has 1.1B parameters, a 256K-token context length, and 2.3 GB of BF16 weights under the openmdw-1.1 license.[11] For local hardware planning, use that weight payload as a starting point, rather than a complete memory requirement. Quantized weight storage would be approximately 0.55 GB, estimated (params x 0.5 bytes) from the cited 1.1B parameters.[11] Quantization support and a specific GPU fit are not established here.

Consider it for local semantic search when balancing benchmark performance against weight storage.[3][11] The caveat: the reported score averages multiple MTEB tasks; it is not a retrieval-only benchmark or a guarantee of performance on your documents.[3]

4. jina-embeddings-v5-text-small

jina-embeddings-v5-text-small by jinaai ranks fourth: its mean main score of 45.2 across 100 published MTEB task results places it behind both Qwen candidates and NVIDIA’s candidate, but ahead of Microsoft’s candidate, as of September 25, 2026.[3] Treat that aggregate as a comparison point, rather than a retrieval-specific guarantee.

Published on January 22, 2026, the model has 596M parameters, a 32K-token context window and 1.4 GB of BF16 weights.[14] For local hardware planning, that weight size is a starting point, not a total memory requirement. Quantized weight storage would be approximately 0.298 GB, estimated (params x 0.5 bytes) from the cited 596M parameters; runtime memory remains additional.[14]

Consider it for local, noncommercial semantic-search experiments where the 32K-token context window suits your documents.[14] Evaluate retrieval quality on your own queries before deployment. The caveat is its cc-by-nc-4.0 license: commercial use requires separate attention to licensing.[14]

5. harrier-oss-v1-0.6b

harrier-oss-v1-0.6b by microsoft ranks fifth because its MTEB mean main score of 33.5 places it below jinaai’s jina-embeddings-v5-text-small and above BAAI’s bge-small-en-v1.5 on the same comparison.[3] Published on Hugging Face on March 30, 2026, the model has 596M parameters, a 32K-token context window and an MIT license.[13] Its position follows benchmark performance, rather than parameter count or download popularity.

Consider harrier for local semantic-search projects where an MIT license and its context window suit your deployment requirements.[13] For hardware planning, the published BF16 weights occupy 1.2 GB; treat that as weight storage, not a complete RAM or VRAM requirement.[13] A specific GPU fit cannot be established from weight size alone. The benchmark caveat matters: the score averages 100 published MTEB task results as of September 25, 2026, rather than reporting a retrieval-only evaluation.[3] Validate search quality on your own documents before committing.

6. bge-small-en-v1.5

bge-small-en-v1.5 by BAAI ranks sixth [3] as an established pick [4]. Its mean main score is 20.5 across 100 published MTEB task results, dated September 25, 2026 [3]. Benchmark ordering places it below the newer candidates and above the other established picks [3]. Its established-pick status permits inclusion outside the release window; popularity does not determine its position [4].

The model has 33 million parameters, a context length of 512 tokens, and an F32 weight footprint of 0.1 GB [4]. Its license is MIT [4]. For local hardware planning, treat that weight footprint as a starting point, not a complete RAM or VRAM requirement. A specific GPU recommendation would require a separate estimate of runtime memory.

Use it as a compact baseline for semantic search over short passages, with the context limit guiding chunk size [4]. The caveat is benchmark scope: the reported score averages mixed MTEB tasks, rather than measuring retrieval alone [3].

7. paraphrase-multilingual-MiniLM-L12-v2

paraphrase-multilingual-MiniLM-L12-v2 by sentence-transformers ranks seventh on the supplied MTEB comparison, with a mean main score of 8.3 across 100 published task results as of September 25, 2026.[3] Its score places it below BAAI’s bge-small-en-v1.5 and above sentence-transformers’ all-MiniLM-L6-v2.[3] It qualifies as an established pick through the family-download exemption, despite its March 2, 2022 publication date.[6]

The model has 118M parameters, a 512-token context length and 0.5 GB of F32 weights, with an Apache-2.0 license.[6] For local hardware planning, treat that weight footprint as a starting point for memory budgeting. Allow additional memory for execution; the weight size alone does not establish a total RAM requirement or guarantee compatibility with a particular GPU.

Use the model as an established baseline when evaluating multilingual semantic search. The caveat is benchmark scope: its reported mean aggregates MTEB task results, so the ranking should not substitute for an evaluation on your own search queries and documents.[3]

8. all-MiniLM-L6-v2

all-MiniLM-L6-v2 by sentence-transformers ranks eighth because its mean main score of 6.3 is the lowest among the candidates on MTEB results dated September 25, 2026, across 100 published task results.[3] Published on March 2, 2022, it qualifies as an established pick through the family’s download-based exemption from the recency gate.[1]

The model has 23M parameters, a 512-token context length, and 0.1 GB of F32 weights under the Apache-2.0 license.[1] For local hardware planning, treat that weight footprint as a storage baseline, rather than a complete RAM or VRAM requirement. A specific CPU, GPU, or total memory requirement is not established here.

Consider it for a local semantic-search prototype over short passages when compact weights are a priority.[1] The caveat is benchmark scope: its reported score averages 100 MTEB task results; that aggregate should not substitute for evaluating retrieval on your own queries and documents.[3]

How much memory should you budget for local embedding models?

Budget for the published weights first: the listed models range from 0.1 GB to 16.3 GB, but use those figures as a starting point rather than a complete RAM or VRAM budget.[1][4][8] Leave room for execution, and verify memory use with your intended workload before choosing hardware.

Qwen’s Qwen3-VL-Embedding-8B (Qwen) has 16.3 GB of BF16 weights.[8] Qwen’s Qwen3-VL-Embedding-2B (Qwen) has 4.3 GB of BF16 weights, while NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) has 2.3 GB.[10][11] Those published sizes give you a concrete starting point for comparing memory requirements without claiming that a particular GPU will fit the workload.

For a smaller weight footprint, Microsoft’s harrier-oss-v1-0.6b (microsoft) lists 1.2 GB of BF16 weights, and Jina AI’s jina-embeddings-v5-text-small (jinaai) lists 1.4 GB.[13][14] The established sentence-transformers model all-MiniLM-L6-v2 (sentence-transformers) and BAAI’s bge-small-en-v1.5 (BAAI) each list 0.1 GB of F32 weights.[1][4]

Quantization calculations should remain explicitly provisional. For Qwen3-VL-Embedding-8B (Qwen)’s 8.1B parameters, a 4-bit weight payload is 4.05 GB, estimated (params x 0.5 bytes).[8] Treat that calculation as a weight-storage estimate, not a measured runtime requirement or confirmation that a compatible quantized release exists. Choose hardware against the model format you will actually run, then check memory use under your intended batch size and input length.

Which licenses do local embedding models use?

Local embedding models use Apache-2.0, MIT, openmdw-1.1 and cc-by-nc-4.0 licenses across the candidates covered here.[1][4][11][14] Record the license alongside the exact repository name when documenting a deployment, and review its terms against the intended use.

Qwen’s Qwen3-VL-Embedding-8B (Qwen) and Qwen3-VL-Embedding-2B (Qwen) both list Apache-2.0.[8][10] The sentence-transformers organization also lists Apache-2.0 for all-MiniLM-L6-v2 (sentence-transformers) and paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers).[1][6] For a shortlist organized by license, those models belong in the same group; evaluate their search performance separately.

BAAI’s bge-small-en-v1.5 (BAAI) and Microsoft’s harrier-oss-v1-0.6b (microsoft) list MIT.[4][13] Alongside the Apache-2.0 candidates, they give engineers alternatives to evaluate under either license.[1][4][8][10][13] Choose the deployment requirements before narrowing the model shortlist.

NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) lists openmdw-1.1, while jinaai’s jina-embeddings-v5-text-small (jinaai) lists cc-by-nc-4.0.[11][14] Keep those exact identifiers in the deployment review rather than treating every downloadable model as having interchangeable terms.

For deployment, check the selected license for obligations concerning use, modification and redistribution. Include the planned application and distribution method in that review, and resolve license suitability before spending time on integration or benchmarking. Treat license approval and semantic-search evaluation as separate decisions.

How are recent releases and established picks ranked?

Recent releases pass the 12-month eligibility gate, then rank by comparable public benchmark results; established picks bypass that gate but follow the same scoring rule and cannot take first place through the exemption.[1][4][6][8][10][11][13][14]

Eligibility is limited to embedding models from labs with a published paper or leaderboard record, or other publishers exceeding 10,000 Hugging Face downloads in the last 30 days under the sentence-similarity or feature-extraction pipeline tags; rerankers are excluded.[3][13][14]

Qwen’s Qwen3-VL-Embedding-8B (Qwen) ranks first because its mean main score of 62.2 is the highest among the eligible candidates on the shared MTEB results record.[3][8] Qwen’s Qwen3-VL-Embedding-2B (Qwen) follows at 58.2, then NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) at 54.2.[3][10][11]

Next come jinaai’s jina-embeddings-v5-text-small (jinaai) at 45.2 and Microsoft’s harrier-oss-v1-0.6b (microsoft) at 33.5.[3][13][14] The established picks follow: BAAI’s bge-small-en-v1.5 (BAAI) at 20.5, sentence-transformers’ paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers) at 8.3, and sentence-transformers’ all-MiniLM-L6-v2 (sentence-transformers) at 6.3.[1][3][4][6]

Each score is the mean across the 100 published MTEB task results for that model, dated September 25, 2026.[3] Those aggregates establish this ordering; they should not be read as scores for your particular search workload.

A leaderboard comparison applies only to models evaluated together. A newer unscored release receives “no public benchmark yet” and cannot be placed below an older scored model on benchmark grounds. Model-card benchmarks must be labeled “self-reported.” Downloads, likes and stars break ties only; parameter count does not determine placement.

Frequently Asked Questions

Which embedding model should I try first for semantic search?

Qwen3-VL-Embedding-8B (Qwen) takes first place because its mean main score of 62.2 leads the eligible candidates across their shared 100 published MTEB task results, dated September 25, 2026 [3]. Its BF16 weights occupy 16.3 GB, its context length is 256K tokens, and its license is Apache-2.0 [8]. Treat the aggregate score as a starting point for evaluation against your own search queries.

How are models admitted and ordered?

Ranking note: Scope admits embedding models under sentence-similarity or feature-extraction from labs with a published paper or leaderboard record, or other publishers exceeding 10,000 downloads in the last 30 days, a threshold Microsoft’s entry clears [13]; rerankers are excluded. Recent releases qualify; established picks bypass recency but cannot lead. Shared MTEB scores determine ordering [3]; popularity only breaks ties. Older scored models cannot outrank newer unscored releases, which receive “no public benchmark yet.”

Which alternatives offer smaller weight downloads?

Qwen3-VL-Embedding-2B (Qwen) ranks next with a mean main score of 58.2, followed by Nemotron-3-Embed-1B-BF16 (nvidia) at 54.2 in MTEB results dated September 25, 2026 [3]. Their BF16 weights occupy 4.3 GB and 2.3 GB respectively [10][11]. Qwen uses Apache-2.0 [10]; NVIDIA uses openmdw-1.1 [11]. Compare their retrieval results on your documents before choosing solely by weight size.

How much local memory should I budget?

Published weight sizes provide a budgeting baseline rather than a complete runtime memory requirement. harrier-oss-v1-0.6b (microsoft) lists 596M parameters, 1.2 GB of BF16 weights, and a 32K-token context [13]. Its mean main score is 33.5 in MTEB results dated September 25, 2026 [3]. Use the weight figure when planning an evaluation, then measure memory with your intended runtime, input lengths, and batching before committing hardware.

Can I use these models in a commercial search product?

Check each license before selecting a model. jina-embeddings-v5-text-small (jinaai) carries cc-by-nc-4.0, so treat it as a noncommercial candidate unless you arrange separate permission [14]. Its mean main score of 45.2 places it below NVIDIA’s candidate and above Microsoft’s in MTEB results dated September 25, 2026 [3]. The Qwen candidates use Apache-2.0 [8][10], while Microsoft’s candidate uses MIT [13]; review the applicable terms for your deployment.

Are older embedding models still worth evaluating?

Established picks remain evaluation candidates, ordered here by their shared MTEB results dated September 25, 2026: bge-small-en-v1.5 (BAAI) scores 20.5; paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers) scores 8.3; all-MiniLM-L6-v2 (sentence-transformers) scores 6.3 [3]. Each is an established pick that bypasses the recency gate [4][6][1]. Their F32 weights occupy 0.1 GB, 0.5 GB, and 0.1 GB respectively [4][6][1]. Download popularity does not establish retrieval quality.

Sources

  1. sentence-transformers/all-MiniLM-L6-v2 model card (Hugging Face) — 2026-09-25
  2. A Repository of Conversational Datasets — 2019-04-13
  3. MTEB results (mteb/results, revision 2026-09-25) — 2026-09-25
  4. BAAI/bge-small-en-v1.5 model card (Hugging Face) — 2026-09-25
  5. Soaring from 4K to 400K: Extending LLM's Context with Activation Beacon — 2024-01-07
  6. sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 model card (Hugging Face) — 2026-09-25
  7. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks — 2019-08-27
  8. Qwen/Qwen3-VL-Embedding-8B model card (Hugging Face) — 2026-09-25
  9. Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — 2026-01-08
  10. Qwen/Qwen3-VL-Embedding-2B model card (Hugging Face) — 2026-09-25
  11. nvidia/Nemotron-3-Embed-1B-BF16 model card (Hugging Face) — 2026-09-25
  12. Compact Language Models via Pruning and Knowledge Distillation — 2024-07-19
  13. microsoft/harrier-oss-v1-0.6b model card (Hugging Face) — 2026-09-25
  14. jinaai/jina-embeddings-v5-text-small model card (Hugging Face) — 2026-09-25
  15. jina-embeddings-v5-text: Task-Targeted Embedding Distillation — 2026-02-17

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog