Quick Answer
Qwen3-VL-Embedding-8B (Qwen) is the top pick to evaluate for code search as of September 2026 because its aggregate MTEB mean score of 62.2 leads the eligible candidates [3][8]. In order, the ranking is Qwen3-VL-Embedding-8B (Qwen), Qwen3-VL-Embedding-2B (Qwen), Nemotron-3-Embed-1B-BF16 (nvidia), jina-embeddings-v5-text-small (jinaai), harrier-oss-v1-0.6b (microsoft), bge-small-en-v1.5 (BAAI), paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers), and all-MiniLM-L6-v2 (sentence-transformers) [3][8].
Key Takeaways
- Qwen’s Qwen3-VL-Embedding-8B (Qwen) leads because its mean MTEB main score of 62.2 is the highest among the ranked candidates as of September 25, 2026; the aggregate covers 100 published task results and does not establish a code-search-specific winner.[3] Validate retrieval quality on your own repositories.
- Qwen’s Qwen3-VL-Embedding-2B (Qwen) follows with 58.2, ahead of nvidia’s Nemotron-3-Embed-1B-BF16 (nvidia) at 54.2.[3] Their published BF16 weights are 4.3 GB and 2.3 GB, respectively.[10][11] Those weight sizes should not be treated as complete runtime memory requirements.
- jinaai’s jina-embeddings-v5-text-small (jinaai) scores 45.2, followed by microsoft’s harrier-oss-v1-0.6b (microsoft) at 33.5.[3] License choice matters: jina-embeddings-v5-text-small lists cc-by-nc-4.0, while harrier-oss-v1-0.6b lists MIT.[14][13]
- BAAI’s bge-small-en-v1.5 (BAAI), sentence-transformers’ paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers), and sentence-transformers’ all-MiniLM-L6-v2 (sentence-transformers) follow in that order, scoring 20.5, 8.3, and 6.3.[3] Each is an established pick admitted through the family popularity exemption.[4][6][1]
- For local hardware planning, Qwen3-VL-Embedding-8B lists 16.3 GB of BF16 weights and Apache-2.0 licensing.[8] bge-small-en-v1.5 lists 0.1 GB of F32 weights and MIT licensing.[4] Use published weight sizes as a starting point, then measure runtime memory with your intended input lengths and batches.
- Ranking note: scope admits only embedding models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days under Hugging Face’s sentence-similarity or feature-extraction tags; rerankers are excluded. Ordering follows the shared MTEB aggregate.[3] Current candidates pass the 12-month release gate; established picks qualify among their families’ 3 most-downloaded models, bypass that gate, and cannot take first place through the exemption.[8][10][11][14][13][4][6][1] Popularity serves only as a tiebreaker.
How do local embedding models compare on parameters, weight memory, context and MTEB scores?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Qwen/Qwen3-VL-Embedding-8B [8] | Qwen [8] | 8.1B [8] | BF16 weights: 16.3 GB [8] | 2026-01-07 [8] | Apache-2.0 [8] | MTEB mean main score: 62.2 across 100 published task results (2026-09-25) [3] |
| Qwen/Qwen3-VL-Embedding-2B [10] | Qwen [10] | 2.1B [10] | BF16 weights: 4.3 GB [10] | 2026-01-07 [10] | Apache-2.0 [10] | MTEB mean main score: 58.2 across 100 published task results (2026-09-25) [3] |
| nvidia/Nemotron-3-Embed-1B-BF16 [11] | nvidia [11] | 1.1B [11] | BF16 weights: 2.3 GB [11] | 2026-07-14 [11] | openmdw-1.1 [11] | MTEB mean main score: 54.2 across 100 published task results (2026-09-25) [3] |
| jinaai/jina-embeddings-v5-text-small [14] | jinaai [14] | 596M [14] | BF16 weights: 1.4 GB [14] | 2026-01-22 [14] | cc-by-nc-4.0 [14] | MTEB mean main score: 45.2 across 100 published task results (2026-09-25) [3] |
| microsoft/harrier-oss-v1-0.6b [13] | microsoft [13] | 596M [13] | BF16 weights: 1.2 GB [13] | 2026-03-30 [13] | MIT [13] | MTEB mean main score: 33.5 across 100 published task results (2026-09-25) [3] |
| BAAI/bge-small-en-v1.5 — established pick [4] | BAAI [4] | 33M [4] | F32 weights: 0.1 GB [4] | 2023-09-12 [4] | MIT [4] | MTEB mean main score: 20.5 across 100 published task results (2026-09-25) [3] |
| sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 — established pick [6] | sentence-transformers [6] | 118M [6] | F32 weights: 0.5 GB [6] | 2022-03-02 [6] | Apache-2.0 [6] | MTEB mean main score: 8.3 across 100 published task results (2026-09-25) [3] |
| sentence-transformers/all-MiniLM-L6-v2 — established pick [1] | sentence-transformers [1] | 23M [1] | F32 weights: 0.1 GB [1] | 2022-03-02 [1] | Apache-2.0 [1] | MTEB mean main score: 6.3 across 100 published task results (2026-09-25) [3] |
Which embedding model should you choose for local code search?
1. Qwen3-VL-Embedding-8B
Qwen3-VL-Embedding-8B by Qwen ranks first because its mean MTEB main score of 62.2 is the highest among the listed candidates on the same benchmark record, dated 2026-09-25.[3] That aggregate covers 100 published task results per model; it does not establish a code-search-specific lead.[3]
Qwen3-VL-Embedding-8B has 8.1B parameters, a 256K-token context window and an Apache-2.0 license, with first publication on Hugging Face on 2026-01-07.[8] Local hardware planning starts with its BF16 weight footprint of 16.3 GB.[8] Four-bit weight storage would be 4.05 GB, estimated (params x 0.5 bytes) from its 8.1B parameters.[8] Weight storage alone does not establish total runtime memory or confirm that a particular GPU can run it.
Consider the model for local search across code and accompanying documentation when long input capacity matters. Evaluate retrieval on representative repository queries before deployment: the aggregate benchmark supports its ranking, but leaves code-specific retrieval quality unresolved.[3]
2. Qwen3-VL-Embedding-2B
Qwen3-VL-Embedding-2B by Qwen ranks second here with a mean main score of 58.2 across 100 published MTEB task results, behind Qwen’s Qwen3-VL-Embedding-8B at 62.2, as of September 25, 2026 [3]. The ordering reflects aggregate MTEB performance; the score does not establish code-search accuracy [3].
Published on January 7, 2026, the model has 2.1 billion parameters, a 256K-token context window, and an Apache-2.0 license [10]. Its published BF16 weights occupy 4.3 GB [10]. For local hardware planning, treat that figure as the weight payload: a complete RAM or VRAM requirement is not established, so a specific GPU fit cannot be promised.
Consider it for local code-search evaluation when you want a smaller weight footprint than Qwen’s Qwen3-VL-Embedding-8B, whose BF16 weights occupy 16.3 GB [8]. Validate retrieval quality on your repository before choosing it; the aggregate ranking leaves code-specific performance unresolved [3].
3. Nemotron-3-Embed-1B-BF16
Nemotron-3-Embed-1B-BF16 by nvidia ranks third on the shared MTEB results, with a mean main score of 54.2 across 100 published task results as of September 25, 2026.[3] Qwen’s Qwen3-VL-Embedding-8B and Qwen3-VL-Embedding-2B score 62.2 and 58.2 respectively on that comparison.[3] The ordering reflects benchmark performance, not parameter count or download popularity.
Published on July 14, 2026, Nemotron-3-Embed-1B-BF16 has 1.1B parameters, a 256K-token context window and an openmdw-1.1 license.[11] For local hardware planning, its BF16 weights occupy 2.3 GB.[11] Treat that footprint as weight storage, not a verified total RAM or VRAM requirement; a specific GPU fit remains unverified.
Consider it for local code-search evaluation when the weight footprint and context allowance suit your deployment.[11] The caveat is benchmark relevance: the reported aggregate does not establish code-search accuracy.[3] Validate retrieval against representative queries and relevant code from your repository before choosing it for production.
4. jina-embeddings-v5-text-small
jina-embeddings-v5-text-small by jinaai ranks fourth on the shared MTEB comparison, with a mean main score of 45.2 across 100 published task results as of 2026-09-25 [3]. Its position follows benchmark performance; download counts and model size do not determine the order. The aggregate result supports its shortlist position but does not establish code-search accuracy.
Published on 2026-01-22, the model has 596M parameters, a 32K-token context window and BF16 weights totaling 1.4 GB [14]. For local hardware planning, use that 1.4 GB weight payload as a starting point [14]. Total runtime RAM and VRAM requirements remain unspecified, so a particular GPU fit cannot be confirmed.
Consider it for local code-search evaluation when its context window suits your code chunks, then measure retrieval quality on your own repositories. The licensing caveat is cc-by-nc-4.0 [14]; check whether those terms suit your intended deployment before adopting it.
5. harrier-oss-v1-0.6b
harrier-oss-v1-0.6b by microsoft ranks fifth on the supplied MTEB comparison, with a mean main score of 33.5 across 100 published task results as of September 25, 2026 [3]. Its position follows benchmark scores rather than parameter count or download popularity.
Published on March 30, 2026, the model has 596M parameters, a 32K-token context window, and an MIT license [13]. BF16 weights occupy 1.2 GB [13]. For local hardware planning, use that weight footprint as a starting point; it does not establish a verified GPU or system-RAM requirement. Validate runtime memory on your target machine before deployment.
Consider it for a local code-search prototype where MIT licensing matters [13], then evaluate retrieval against queries from your own repository. The caveat is benchmark relevance: the reported MTEB mean aggregates published task results and does not establish code-search accuracy [3].
6. bge-small-en-v1.5
BAAI’s bge-small-en-v1.5 ranks sixth [3] as an established pick that bypasses the recency gate [4]. Its mean main score is 20.5 across 100 published MTEB task results as of 2026-09-25 [3], placing it below the current-generation candidates and above the other established picks in this comparison [3]. Those aggregate results do not establish code-search performance.
The model has 33M parameters, a 512-token context window, and an MIT license [4]. Local hardware planning starts with its published F32 weight footprint of 0.1 GB [4]. Total runtime RAM and a validated CPU or GPU configuration are not specified, so the weight footprint alone should not determine hardware capacity.
Consider it for a local search baseline over short code snippets or individual functions that fit within its 512-token window [4]. Evaluate retrieval on your own repository before adopting it; the practical caveat is the absence of a code-specific benchmark here.
7. paraphrase-multilingual-MiniLM-L12-v2
paraphrase-multilingual-MiniLM-L12-v2 by sentence-transformers is an established pick with an MTEB mean main score of 8.3, placing it below bge-small-en-v1.5 (BAAI) at 20.5 and above all-MiniLM-L6-v2 (sentence-transformers) at 6.3.[3][6] Those averages cover 100 published task results per model as of September 25, 2026; they do not establish a code-search-specific advantage.[3]
The model has 118M parameters, a 512-token context window, published F32 weights of 0.5 GB, and an Apache-2.0 license.[6] For local hardware planning, treat that weight footprint as a starting point, not a total RAM requirement. Allow additional memory for execution; a specific CPU, GPU or measured runtime memory requirement is not established.
Use the model as an established baseline when evaluating search over short code snippets. Keep inputs within its 512-token context limit.[6] The practical caveat is benchmark relevance: evaluate retrieval on your own code queries before choosing it on the strength of its aggregate MTEB score.[3]
8. all-MiniLM-L6-v2
all-MiniLM-L6-v2 by sentence-transformers ranks eighth because its mean main score of 6.3 is the lowest among the candidates across the 100 published MTEB task results, dated 2026-09-25.[3] Published on 2022-03-02, it qualifies as an established pick through its place among the family’s three most-downloaded models, rather than release recency.[1] Its practical role is a compact baseline for evaluating whether embeddings help retrieve relevant code in your repository.
The model has 23M parameters, a 512-token context length, and 0.1 GB of F32 weights under the Apache-2.0 license.[1] For local hardware planning, use that weight footprint as a starting point; total RAM requirements and GPU fit remain unquantified. Try it with short code snippets and accompanying descriptions. The caveat is benchmark relevance: the reported MTEB aggregate does not establish code-search accuracy, so validate retrieval against representative queries before adopting it.[3]
What do MTEB scores tell you about code search quality?
MTEB scores provide a broad benchmark comparison, but the reported averages do not establish how well a model retrieves code. The published figures summarize mean main scores across 100 task results per model; they are not identified as code-search scores.[3] Treat them as a starting point for evaluation, rather than a measured answer about your repository.
Qwen’s Qwen3-VL-Embedding-8B (Qwen) records a mean main score of 62.2, while Qwen’s Qwen3-VL-Embedding-2B (Qwen) records 58.2 in the MTEB results dated 2026-09-25.[3] Those figures support an ordering by reported aggregate score. They do not establish the same ordering for finding a function from a natural-language description, locating an implementation example, or retrieving code that uses a particular API.
An aggregate also leaves practical questions unanswered. The reported means do not show which code-search tasks contributed, how individual programming languages performed, or whether a model handled your query style well.[3] Avoid translating a difference in aggregate score into an expected improvement in code retrieval.
For a local deployment, evaluate candidate models against representative queries and relevant code from your repository. Keep the indexed content and retrieval setup consistent, then inspect whether useful matches appear early in the results. Use that evaluation to make the code-search decision; use MTEB to inform the shortlist.
How much memory should you budget to run an embedding model locally?
Budget for model weights plus runtime memory, using the published weight size as your starting point. Weight sizes alone do not establish a complete RAM or VRAM requirement.
For a compact baseline, sentence-transformers’ all-MiniLM-L6-v2 (sentence-transformers) and BAAI’s bge-small-en-v1.5 (BAAI) each list F32 weights of 0.1 GB.[1][4] sentence-transformers’ paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers) lists F32 weights of 0.5 GB.[6]
For newer candidates, Microsoft’s harrier-oss-v1-0.6b (microsoft) lists BF16 weights of 1.2 GB,[13] while Jina AI’s jina-embeddings-v5-text-small (jinaai) lists BF16 weights of 1.4 GB.[14] NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) lists BF16 weights of 2.3 GB.[11]
Qwen’s Qwen3-VL-Embedding-2B (Qwen) lists BF16 weights of 4.3 GB,[10] and Qwen’s Qwen3-VL-Embedding-8B (Qwen) lists BF16 weights of 16.3 GB.[8] Use those footprints when deciding which models to evaluate on your hardware, but avoid treating them as verified GPU capacity requirements.
Quantization calculations also need a clear boundary. For Qwen3-VL-Embedding-8B (Qwen), with 8.1B parameters,[8] a weight-only budget is approximately 4.05 GB at 4-bit precision—estimated (params x 0.5 bytes).[8] Actual runtime memory remains a separate measurement.
Before committing hardware, measure peak memory with your intended code chunks and batch settings. Record the runtime, precision and workload alongside that measurement so your deployment budget is reproducible.
Which licenses apply to these embedding models?
The models use Apache-2.0 [1][6][8][10], MIT [4][13], openmdw-1.1 [11], or cc-by-nc-4.0 [14], depending on the checkpoint.
The Apache-2.0 group includes sentence-transformers’ all-MiniLM-L6-v2 (sentence-transformers) [1] and paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers) [6]. Qwen’s Qwen3-VL-Embedding-8B (Qwen) [8] and Qwen3-VL-Embedding-2B (Qwen) [10] also use Apache-2.0. For a shortlist organized by license, keep those checkpoints together, then evaluate their suitability for your code-search workload separately.
The MIT group contains BAAI’s bge-small-en-v1.5 (BAAI) [4] and Microsoft’s harrier-oss-v1-0.6b (microsoft) [13]. Record the license alongside the exact repository identifier when documenting your deployment choice.
NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) uses openmdw-1.1 [11]. Review that license directly when assessing a deployment; do not substitute an Apache or MIT review for the checkpoint’s own terms.
Jina AI’s jina-embeddings-v5-text-small (jinaai) uses cc-by-nc-4.0 [14]. Treat that license as a separate review item, particularly when evaluating the model for a commercial code-search service or an internal company tool.
For implementation, make license review part of checkpoint selection. Keep the repository identifier, license text, and intended use in your deployment record. Assess permission for your intended use separately from retrieval quality and hardware requirements; a benchmark ranking should not make the licensing decision for you.
Frequently Asked Questions
Which embedding model should I try first for local code search?
Qwen’s Qwen3-VL-Embedding-8B (Qwen) takes first place among these candidates because its MTEB mean main score of 62.2 leads the shared comparison dated September 25, 2026.[3] Qwen’s Qwen3-VL-Embedding-2B (Qwen) follows at 58.2.[3] Both means cover 100 published task results per model, rather than a dedicated code-search evaluation.[3] Use that ordering to build a shortlist, then evaluate retrieval against questions and relevant functions from your own repository.
How are the embedding models ranked?
Ranking note: eligibility admits embedding models from labs with a paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days under sentence-similarity or feature-extraction; rerankers are excluded.[3][13] Apply the 12-month gate.[8][11] Established picks bypass it but cannot win through exemption.[1][4][6] Comparable MTEB scores determine order; popularity only breaks ties, never size.[3] Never place an older scored model above a newer unscored candidate; label unscored candidates “no public benchmark yet.”
How much GPU memory do I need to run these models locally?
Budget from the published weights: Qwen3-VL-Embedding-8B (Qwen) has 16.3 GB of BF16 weights, while Qwen3-VL-Embedding-2B (Qwen) has 4.3 GB.[8][10] NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) has 2.3 GB of BF16 weights.[11] Treat those figures as weight storage, not verified GPU-memory requirements. Measure peak memory with your intended input length and batch size before committing hardware; a specific GPU-fit claim would require a separate estimate or measurement.
Which alternatives have smaller published weight files?
NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) scores 54.2, jinaai’s jina-embeddings-v5-text-small (jinaai) scores 45.2, and Microsoft’s harrier-oss-v1-0.6b (microsoft) scores 33.5 in the MTEB comparison dated September 25, 2026.[3] Their BF16 weights occupy 2.3 GB, 1.4 GB, and 1.2 GB, respectively.[11][14][13] That benchmark order gives you alternatives to evaluate as your weight-storage budget changes. Treat the aggregate scores as comparison points, without assuming the same ordering on your codebase.
Which licenses should I check before using an embedding model at work?
Qwen3-VL-Embedding-8B (Qwen) and Qwen3-VL-Embedding-2B (Qwen) list Apache-2.0, while harrier-oss-v1-0.6b (microsoft) lists MIT.[8][10][13] NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) lists openmdw-1.1, and jina-embeddings-v5-text-small (jinaai) lists cc-by-nc-4.0.[11][14] Make license review part of your selection process before integrating a model into an internal service or distributed product. Evaluate the named license against your intended use rather than treating every downloadable model as having interchangeable terms.
Are older embedding models still useful comparison baselines?
The established picks rank as BAAI’s bge-small-en-v1.5 (BAAI) at 20.5, sentence-transformers’ paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers) at 8.3, and sentence-transformers’ all-MiniLM-L6-v2 (sentence-transformers) at 6.3 in MTEB results dated September 25, 2026.[3] Each qualifies for the established-pick exemption through family download standing.[4][6][1] Their F32 weight files are 0.1 GB, 0.5 GB, and 0.1 GB, respectively.[4][6][1] Include them as local comparison baselines, while keeping their exemption separate from benchmark merit.
Sources
- sentence-transformers/all-MiniLM-L6-v2 model card (Hugging Face) — 2026-09-25
- A Repository of Conversational Datasets — 2019-04-13
- MTEB results (mteb/results, revision 2026-09-25) — 2026-09-25
- BAAI/bge-small-en-v1.5 model card (Hugging Face) — 2026-09-25
- Soaring from 4K to 400K: Extending LLM's Context with Activation Beacon — 2024-01-07
- sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 model card (Hugging Face) — 2026-09-25
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks — 2019-08-27
- Qwen/Qwen3-VL-Embedding-8B model card (Hugging Face) — 2026-09-25
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — 2026-01-08
- Qwen/Qwen3-VL-Embedding-2B model card (Hugging Face) — 2026-09-25
- nvidia/Nemotron-3-Embed-1B-BF16 model card (Hugging Face) — 2026-09-25
- Compact Language Models via Pruning and Knowledge Distillation — 2024-07-19
- microsoft/harrier-oss-v1-0.6b model card (Hugging Face) — 2026-09-25
- jinaai/jina-embeddings-v5-text-small model card (Hugging Face) — 2026-09-25
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation — 2026-02-17