Best Feature Extraction Models to Run Locally in 2026: Qwen3-VL-Embedding-8B

Rankings 2026-09-28 Last updated 2026-09-28 15 min read By Q4KM

Quick Answer

As of September 2026, Qwen3-VL-Embedding-8B (Qwen) is the top pick for local feature extraction, with the highest mean main MTEB score among these candidates: 62.2 across 100 published task results [3][13]. In order, the ranking is Qwen3-VL-Embedding-8B (Qwen), Qwen3-VL-Embedding-2B (Qwen), jina-embeddings-v5-text-small (jinaai), harrier-oss-v1-0.6b (microsoft), bge-small-en-v1.5 (BAAI), jina-embeddings-v5-omni-small (jinaai), paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers), and all-MiniLM-L6-v2 (sentence-transformers) [3][13].

Key Takeaways

How do local feature extraction models compare on specifications and benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
Qwen/Qwen3-VL-Embedding-8B [13] Qwen [13] 8.1B [13] 4-bit weights only: 4.05 GB estimated (params x 0.5 bytes) [13] 2026-01-07 [13] Apache-2.0 [13] MTEB mean main score: 62.2 across 100 task results (2026-09-25) [3]
Qwen/Qwen3-VL-Embedding-2B [15] Qwen [15] 2.1B [15] 4-bit weights only: 1.05 GB estimated (params x 0.5 bytes) [15] 2026-01-07 [15] Apache-2.0 [15] MTEB mean main score: 58.2 across 100 task results (2026-09-25) [3]
jinaai/jina-embeddings-v5-text-small [9] jinaai [9] 596M [9] 4-bit weights only: 298 MB estimated (params x 0.5 bytes) [9] 2026-01-22 [9] cc-by-nc-4.0 [9] MTEB mean main score: 45.2 across 100 task results (2026-09-25) [3]
microsoft/harrier-oss-v1-0.6b [8] Microsoft [8] 596M [8] 4-bit weights only: 298 MB estimated (params x 0.5 bytes) [8] 2026-03-30 [8] MIT [8] MTEB mean main score: 33.5 across 100 task results (2026-09-25) [3]
BAAI/bge-small-en-v1.5 — established pick [4] BAAI [4] 33M [4] 4-bit weights only: 16.5 MB estimated (params x 0.5 bytes) [4] 2023-09-12 [4] MIT [4] MTEB mean main score: 20.5 across 100 task results (2026-09-25) [3]
jinaai/jina-embeddings-v5-omni-small [11] jinaai [11] 1.6B [11] 4-bit weights only: 800 MB estimated (params x 0.5 bytes) [11] 2026-04-01 [11] cc-by-nc-4.0 [11] MTEB mean main score: 16.6 across 100 task results (2026-09-25) [3]
sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 — established pick [6] sentence-transformers [6] 118M [6] 4-bit weights only: 59 MB estimated (params x 0.5 bytes) [6] 2022-03-02 [6] Apache-2.0 [6] MTEB mean main score: 8.3 across 100 task results (2026-09-25) [3]
sentence-transformers/all-MiniLM-L6-v2 — established pick [1] sentence-transformers [1] 23M [1] 4-bit weights only: 11.5 MB estimated (params x 0.5 bytes) [1] 2022-03-02 [1] Apache-2.0 [1] MTEB mean main score: 6.3 across 100 task results (2026-09-25) [3]

Which feature extraction models should you run locally?

1. Qwen3-VL-Embedding-8B

Qwen3-VL-Embedding-8B by Qwen ranks first because its MTEB mean main score of 62.2 is the highest among the eligible candidates in the shared benchmark snapshot dated 2026-09-25.[3] The score averages 100 published task results, making it a broad comparison rather than a guarantee for your workload.[3]

Published on Hugging Face on 2026-01-07, Qwen3-VL-Embedding-8B has 8.1B parameters, a 256K-token context window and an Apache-2.0 license.[13] Its intended application is multimodal retrieval, making local retrieval across text and visual content a relevant use case.[14]

For local hardware planning, the BF16 weights occupy 16.3 GB.[13] At the cited 8.1B parameters, 4-bit weight storage is 4.05 GB, estimated (params x 0.5 bytes).[13] The practical caveat is memory sizing: that calculation covers weights alone. Allow additional memory for runtime overhead, and verify your chosen runtime and workload before committing to a GPU or system RAM configuration.

2. Qwen3-VL-Embedding-2B

Qwen3-VL-Embedding-2B by Qwen ranks second on the shared benchmark, with a mean main score of 58.2 across its 100 published MTEB task results as of 2026-09-25, behind Qwen’s Qwen3-VL-Embedding-8B at 62.2.[3] Its position follows benchmark performance rather than parameter count or download volume.

The model has 2.1B parameters, a 256K-token context length, an Apache-2.0 license, and 4.3 GB of BF16 weights.[15] For local hardware planning, its 4-bit weight footprint is 1.05 GB, estimated (params x 0.5 bytes) from the cited 2.1B parameters.[15] Treat that estimate as a weight budget, not a complete RAM or VRAM requirement; allow additional memory for execution.

Use it for local multimodal retrieval when that capability matches your application.[14] The caveat is benchmark interpretation: the reported 58.2 is an aggregate across 100 published task results, so it does not establish performance on your particular retrieval workload.[3] Evaluate representative queries and documents before committing to deployment.

3. jina-embeddings-v5-text-small

jina-embeddings-v5-text-small by jinaai ranks third here: its mean main score of 45.2 across 100 published MTEB task results places it below the two higher-scoring candidates and above the remaining candidates, as of 2026-09-25.[3] Published on 2026-01-22, the model has 596M parameters, a 32K-token context window and BF16 weights listed at 1.4 GB.[9]

For local hardware planning, its 4-bit weight footprint is estimated at 298 MB (params x 0.5 bytes), based on the cited 596M parameters.[9] Treat that calculation as a weight-storage estimate, not a complete RAM or VRAM requirement. The listed BF16 weight size provides another starting point for planning a local deployment.[9]

Use it as a candidate for local text-embedding workloads that benefit from task-targeted embedding distillation.[10] The practical caveat is licensing: the model uses CC-BY-NC-4.0, so commercial deployment requires attention to its noncommercial restriction.[9]

4. harrier-oss-v1-0.6b

harrier-oss-v1-0.6b from microsoft ranks fourth because its MTEB mean main score of 33.5 falls below jina-embeddings-v5-text-small from jinaai at 45.2 and above bge-small-en-v1.5 from BAAI at 20.5.[3] Those means cover the 100 task results published for each model as of 2026-09-25.[3]

Published on Hugging Face on 2026-03-30, harrier-oss-v1-0.6b has 596M parameters, a 32K-token context window, BF16 weights listed at 1.2 GB, and an MIT license.[8] Consider it for local feature extraction when the MIT license and stated context window match your application’s requirements.

For hardware planning, its 596M parameters imply 298 MB of 4-bit weights, estimated (params x 0.5 bytes).[8] Treat that figure as a weights-only budget: a specific GPU or system-RAM requirement cannot be established from it alone. The practical caveat is that the listed BF16 weight size and the calculated quantized size do not establish total runtime memory needs.

5. bge-small-en-v1.5

bge-small-en-v1.5 by BAAI ranks fifth on the shared MTEB comparison, with a mean main score of 20.5 across 100 published task results as of September 25, 2026.[3] The model qualifies as an established pick: its position among the family’s three most-downloaded models permits an exception to the 12-month release window.[4] Benchmark performance determines its position here; download volume only explains its eligibility. The model has 33M parameters, a 512-token context window, and an MIT license.[4]

For local hardware planning, its 4-bit weight memory is 16.5 MB, estimated (params x 0.5 bytes) from the cited 33M parameters.[4] Treat that figure as a weight-storage estimate, not a complete RAM or GPU-memory requirement. Short-text embedding workloads are a practical use within its 512-token context limit.[4] The caveat is input length: check document lengths against that limit before choosing it for your pipeline.[4]

6. jina-embeddings-v5-omni-small

jina-embeddings-v5-omni-small by jinaai ranks sixth because its mean main score of 16.6 trails five candidates in the shared comparison across 100 published MTEB task results, dated 2026-09-25.[3] Its position follows benchmark scores rather than parameter count or download popularity. Hugging Face first published the model on 2026-04-01.[11]

The model has 1.6B parameters, a 32K-token context window and 3.4 GB of BF16 weights.[11] Quantized weight storage is approximately 0.8 GB, estimated (params x 0.5 bytes) from the cited 1.6B parameters.[11] For local hardware planning, treat that estimate as weight storage alone; total RAM or VRAM requirements remain unspecified, so it does not establish a particular GPU fit.

Choose it for local multimodal embedding work: its architecture uses frozen-tower composition to preserve text geometry while adding multimodal embeddings.[12] The practical caveat is its cc-by-nc-4.0 license, which makes commercial deployment a licensing consideration.[11]

7. paraphrase-multilingual-MiniLM-L12-v2

paraphrase-multilingual-MiniLM-L12-v2 by sentence-transformers ranks seventh here: its mean main score of 8.3 across 100 published MTEB task results places it below bge-small-en-v1.5 (BAAI) and above all-MiniLM-L6-v2 (sentence-transformers), as of 2026-09-25.[3] Its inclusion is an established-pick exemption from the release window, supported by its position among the family’s three most-downloaded models.[6]

The model has 118M parameters, a 512-token context limit and an Apache-2.0 license.[6] Published F32 weights occupy 0.5 GB.[6] For local hardware planning, 4-bit weight storage is 59 MB, estimated (params x 0.5 bytes) from the cited parameter count.[6] Treat that estimate as a weight-storage budget; total runtime RAM and a specific GPU fit are not established by that calculation.

Use it as an established multilingual embedding baseline when evaluating a local deployment. The practical caveat is its 512-token context limit: check input lengths before choosing it for document workloads.[6] Its original publication date was 2022-03-02, so its place here reflects the established-pick exception.[6]

8. all-MiniLM-L6-v2

all-MiniLM-L6-v2 by sentence-transformers ranks eighth because its mean main score of 6.3 across 100 published MTEB task results is below every other candidate’s score as of September 25, 2026.[3] First published on March 2, 2022, it qualifies as an established pick through the family’s download-based exemption from the recency gate.[1] Popularity explains its inclusion, while benchmark performance determines its position.

The model has 23M parameters, a context length of 512 tokens, and an Apache-2.0 license.[1] For local hardware planning, its listed F32 weights occupy 0.1 GB.[1] At 4-bit precision, weight storage would be 11.5 MB, estimated (params x 0.5 bytes) from the cited 23M parameters.[1] Allow additional memory for execution; that estimate covers weights alone.

Use it as a compact baseline for embedding short text when keeping parameter storage low matters more than its benchmark position.[1][3] The practical caveat is context: keep inputs within the stated 512-token limit.[1]

How do you estimate memory requirements from an embedding model's parameter count?

Estimate embedding-model weight memory by multiplying the parameter count by the storage required per parameter: for Qwen’s Qwen3-VL-Embedding-8B, 8.1B parameters imply 4.05 GB of weight storage at 4-bit, estimated (params x 0.5 bytes).[13] Treat that result as a weight-storage budget, rather than a complete runtime memory requirement.

Apply the same calculation to smaller candidates. Microsoft’s harrier-oss-v1-0.6b has 596M parameters, giving 298 MB at 4-bit, estimated (params x 0.5 bytes).[8] Jina AI’s jina-embeddings-v5-omni-small has 1.6B parameters, giving 800 MB at 4-bit, estimated (params x 0.5 bytes).[11] Keep the estimate beside the parameter count so readers can distinguish calculated storage from published weight sizes.

Check the checkpoint’s actual format before planning a deployment. Qwen3-VL-Embedding-8B lists BF16 weights of 16.3 GB; the quantized estimate does not describe that downloadable checkpoint’s size.[13] A parameter-based calculation also does not establish that a compatible quantized checkpoint is available.

Use the estimate to compare weight budgets, then measure total memory with your chosen runtime and workload. Avoid turning a weight-only calculation into a GPU-fit claim: a deployment budget must account for more than the model’s stored parameters.

Which licenses do local embedding models use?

Local embedding models in this selection use Apache-2.0, MIT, or CC-BY-NC-4.0; the license depends on the model you download.[1][4][9]

Apache-2.0 covers Qwen’s Qwen3-VL-Embedding-8B (Qwen) and Qwen3-VL-Embedding-2B (Qwen).[13][15] The sentence-transformers organization also uses Apache-2.0 for all-MiniLM-L6-v2 (sentence-transformers) and paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers).[1][6] Those models share a license label, but evaluate their suitability for your application separately from their licensing.

MIT covers Microsoft’s harrier-oss-v1-0.6b (microsoft) and BAAI’s bge-small-en-v1.5 (BAAI).[8][4] For a deployment inventory, record the exact repository identifier alongside its license. Keeping that mapping explicit makes a later model replacement easier to review.

CC-BY-NC-4.0 covers jinaai’s jina-embeddings-v5-text-small (jinaai) and jina-embeddings-v5-omni-small (jinaai).[9][11] Treat that license as a separate review item rather than assuming all downloadable embedding weights carry interchangeable terms. For a planned commercial deployment, check the applicable license text before committing to either model.

Choose a model against both your technical requirements and your intended use. Read the license attached to the exact model revision, retain the relevant notices, and document your licensing decision alongside the deployment configuration. Running inference on your own hardware should be part of that review, not a substitute for checking the terms.

How do established embedding models compare with recent releases?

Established embedding models offer smaller parameter counts, but several recent releases score higher on the reported MTEB aggregate.[1][4][6][8][9][13][15] Qwen’s Qwen3-VL-Embedding-8B (Qwen) takes first place because its mean main score of 62.2 leads the eligible candidates on the same benchmark.[3]

The established picks are sentence-transformers’ all-MiniLM-L6-v2 (sentence-transformers), BAAI’s bge-small-en-v1.5 (BAAI) and sentence-transformers’ paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers).[1][4][6] Their respective scores are 6.3, 20.5 and 8.3, measured as mean main scores across each model’s 100 published MTEB task results as of 2026-09-25.[3]

Recent releases include Qwen’s Qwen3-VL-Embedding-2B (Qwen) at 58.2, jinaai’s jina-embeddings-v5-text-small (jinaai) at 45.2 and microsoft’s harrier-oss-v1-0.6b (microsoft) at 33.5 on that aggregate.[3] Newer does not guarantee a higher score: jinaai’s jina-embeddings-v5-omni-small (jinaai), published on 2026-04-01,[11] scores 16.6, below the established BAAI model.[3]

Parameter counts expose the local-memory tradeoff. The BAAI model has 33M parameters,[4] giving estimated weight memory of 16.5 MB at 4-bit (estimated (params x 0.5 bytes)).[4] The winning Qwen model has 8.1B parameters,[13] giving 4.05 GB by the same calculation (estimated (params x 0.5 bytes)).[13]

Ranking note: scope admits only embedding models from labs with a published paper or leaderboard record, using Hugging Face’s feature-extraction or sentence-similarity tags; rerankers are excluded. Eligible releases are ordered by comparable MTEB scores, with popularity used only for ties.[3] Established picks bypass the recency gate through their family’s three-most-downloaded exemption; that exemption cannot award first place.[1][4][6]

Frequently Asked Questions

Which model should I start with for local feature extraction?

Qwen’s Qwen3-VL-Embedding-8B (Qwen) takes first place because its mean MTEB main score of 62.2 is the highest among the eligible candidates as of 2026-09-25 [3]. The model has 8.1B parameters, a 256K-token context window and an Apache-2.0 license [13]. Qwen’s Qwen3-VL-Embedding-2B (Qwen) offers a smaller alternative with 2.1B parameters and a mean score of 58.2 on the same reporting basis [15][3].

How does the ranking handle recent releases and established models?

The ranking admits only embedding models from labs with a published paper or leaderboard record, under Hugging Face’s feature-extraction or sentence-similarity tags; rerankers are excluded. Recency eligibility comes first, followed by comparable MTEB scores [3]; popularity only breaks ties. Established picks bypass the recency gate but cannot win through that exemption [1][4][6]. An older scored model cannot outrank a newer unscored model, which receives “no public benchmark yet”; model-card benchmarks are labeled self-reported.

How much memory should I budget for quantized weights?

Microsoft’s harrier-oss-v1-0.6b (microsoft) has 596M parameters [8]: 4-bit weight memory is 0.298 GB, estimated (params x 0.5 bytes) [8]. Jina AI’s jina-embeddings-v5-text-small (jinaai) has 596M parameters [9]: 4-bit weight memory is also 0.298 GB, estimated (params x 0.5 bytes) [9]. Sentence-transformers’ all-MiniLM-L6-v2 (sentence-transformers) has 23M parameters [1]: 4-bit weight memory is 0.0115 GB, estimated (params x 0.5 bytes) [1]. Treat these as weight estimates, not complete runtime memory budgets.

Which licenses should I check before choosing a model?

The Qwen embedding candidates use Apache-2.0 [13][15], while Microsoft’s harrier model uses MIT [8]. Jina AI’s text and omni candidates use cc-by-nc-4.0, so account for the noncommercial restriction when selecting a deployment model [9][11]. Among established picks, BAAI’s bge-small-en-v1.5 (BAAI) uses MIT [4]. Sentence-transformers’ paraphrase-multilingual-MiniLM-L12-v2 (sentence-transformers) and its all-MiniLM model use Apache-2.0 [6][1]. License suitability belongs alongside benchmark performance in your selection.

Which candidates support long input documents?

The Qwen embedding candidates list context windows of 256K tokens [13][15]. Microsoft’s harrier model and Jina AI’s text model list 32K tokens [8][9]. The established BAAI and sentence-transformers candidates each list 512 tokens [4][1][6]. Compare your intended input length with those limits before selecting a model. The context figures describe supported input length; the MTEB means provide a separate benchmark comparison [3].

Where does the Jina omni model belong in the comparison?

Jina AI’s jina-embeddings-v5-omni-small (jinaai) was first published on Hugging Face on 2026-04-01 [11]. The model has 1.6B parameters, a 32K-token context window and a cc-by-nc-4.0 license [11]. Its mean MTEB main score is 16.6 across 100 published task results as of 2026-09-25 [3]. On that benchmark basis, it ranks below Jina AI’s text model at 45.2 and the established BAAI model at 20.5 [3].

Sources

  1. sentence-transformers/all-MiniLM-L6-v2 model card (Hugging Face) — 2026-09-25
  2. A Repository of Conversational Datasets — 2019-04-13
  3. MTEB results (mteb/results, revision 2026-09-25) — 2026-09-25
  4. BAAI/bge-small-en-v1.5 model card (Hugging Face) — 2026-09-25
  5. Soaring from 4K to 400K: Extending LLM's Context with Activation Beacon — 2024-01-07
  6. sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 model card (Hugging Face) — 2026-09-25
  7. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks — 2019-08-27
  8. microsoft/harrier-oss-v1-0.6b model card (Hugging Face) — 2026-09-25
  9. jinaai/jina-embeddings-v5-text-small model card (Hugging Face) — 2026-09-25
  10. jina-embeddings-v5-text: Task-Targeted Embedding Distillation — 2026-02-17
  11. jinaai/jina-embeddings-v5-omni-small model card (Hugging Face) — 2026-09-25
  12. jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition — 2026-05-08
  13. Qwen/Qwen3-VL-Embedding-8B model card (Hugging Face) — 2026-09-25
  14. Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — 2026-01-08
  15. Qwen/Qwen3-VL-Embedding-2B model card (Hugging Face) — 2026-09-25

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog