Quick Answer
Qwen3-VL-Embedding-8B (Qwen) is the top pick for Apple Silicon Macs as of September 2026 because its mean main score of 62.2 across 100 published MTEB task results is the highest among the scored candidates [1][3]. In order, the ranking is Qwen3-VL-Embedding-8B (Qwen), Qwen3-VL-Embedding-2B (Qwen), Nemotron-3-Embed-1B-BF16 (nvidia), jina-embeddings-v5-text-small (jinaai), harrier-oss-v1-0.6b (microsoft), harrier-oss-v1-270m (microsoft), jina-embeddings-v5-omni-small (jinaai), and llama-nemotron-embed-1b-v2 (nvidia) [1][3].
Key Takeaways
- Qwen’s Qwen3-VL-Embedding-8B ranks first because its published MTEB mean main score of 62.2 leads the scored candidates across 100 task results.[3] Its 8.1B parameters imply 4.05 GB of weight storage at 4-bit, estimated (params x 0.5 bytes).[1]
- Qwen’s Qwen3-VL-Embedding-2B offers a smaller weight footprint: 2.1B parameters imply 1.05 GB at 4-bit, estimated (params x 0.5 bytes), with a 256K-token context and Apache-2.0 license.[4] Its published MTEB mean main score is 58.2.[3]
- NVIDIA’s Nemotron-3-Embed-1B-BF16 combines a published MTEB mean main score of 54.2[3] with 256K-token context and an openmdw-1.1 license; its 1.1B parameters imply 0.55 GB at 4-bit, estimated (params x 0.5 bytes).[5]
- Microsoft’s harrier-oss-v1-270m is a compact MIT-licensed option with 32K-token context: 268M parameters imply 0.134 GB at 4-bit, estimated (params x 0.5 bytes).[14] Its published MTEB mean main score is 29.6.[3] Weight-storage estimates alone do not establish Apple Silicon runtime memory requirements or speed.
- NVIDIA’s llama-nemotron-embed-1b-v2 has no public benchmark yet (MTEB results, archived date unknown).[3] Its final placement reflects missing benchmark coverage, rather than a demonstrated quality deficit; its license is nvidia-open-model-license.[12]
- Ranking note: admit only embedding models tagged sentence-similarity or feature-extraction; exclude rerankers. Apply release eligibility first, then order comparable scored candidates by published MTEB mean main score; never place an older scored model above a newer unscored model.[3] Model size does not determine ordering, and popularity serves only as a tiebreak.
How do these embedding models compare on memory, context, licensing and benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Qwen3-VL-Embedding-8B [1] | Qwen [1] | 8.1B [1] | 4-bit weights: 4.05 GB estimated (params x 0.5 bytes); not total runtime memory [1] | 2026-01-07 [1] | Apache-2.0 [1] | MTEB mean main score: 62.2 across 100 task results (2026-09-25) [3] |
| Qwen3-VL-Embedding-2B [4] | Qwen [4] | 2.1B [4] | 4-bit weights: 1.05 GB estimated (params x 0.5 bytes); not total runtime memory [4] | 2026-01-07 [4] | Apache-2.0 [4] | MTEB mean main score: 58.2 across 100 task results (2026-09-25) [3] |
| Nemotron-3-Embed-1B-BF16 [5] | nvidia [5] | 1.1B [5] | 4-bit weights: 0.55 GB estimated (params x 0.5 bytes); not total runtime memory [5] | 2026-07-14 [5] | openmdw-1.1 [5] | MTEB mean main score: 54.2 across 100 task results (2026-09-25) [3] |
| jina-embeddings-v5-text-small [8] | jinaai [8] | 596M [8] | 4-bit weights: 0.298 GB estimated (params x 0.5 bytes); not total runtime memory [8] | 2026-01-22 [8] | cc-by-nc-4.0 [8] | MTEB mean main score: 45.2 across 100 task results (2026-09-25) [3] |
| harrier-oss-v1-0.6b [7] | microsoft [7] | 596M [7] | 4-bit weights: 0.298 GB estimated (params x 0.5 bytes); not total runtime memory [7] | 2026-03-30 [7] | MIT [7] | MTEB mean main score: 33.5 across 100 task results (2026-09-25) [3] |
| harrier-oss-v1-270m [14] | microsoft [14] | 268M [14] | 4-bit weights: 0.134 GB estimated (params x 0.5 bytes); not total runtime memory [14] | 2026-03-30 [14] | MIT [14] | MTEB mean main score: 29.6 across 100 task results (2026-09-25) [3] |
| jina-embeddings-v5-omni-small [10] | jinaai [10] | 1.6B [10] | 4-bit weights: 0.80 GB estimated (params x 0.5 bytes); not total runtime memory [10] | 2026-04-01 [10] | cc-by-nc-4.0 [10] | MTEB mean main score: 16.6 across 100 task results (2026-09-25) [3] |
| llama-nemotron-embed-1b-v2 [12] | nvidia [12] | 1.2B [12] | 4-bit weights: 0.60 GB estimated (params x 0.5 bytes); not total runtime memory [12] | 2025-10-16 [12] | nvidia-open-model-license [12] | no public benchmark yet (MTEB results, archived date unknown) [3] |
Which embedding models should you consider for an Apple Silicon Mac?
1. Qwen3-VL-Embedding-8B
Qwen3-VL-Embedding-8B by Qwen ranks first because its mean main score of 62.2 across 100 published Massive Text Embedding Benchmark (MTEB) task results is the highest among the scored candidates here.[3] The model has 8.1B parameters, a 256K-token context window and an Apache-2.0 license.[1] Use it for retrieval projects where published embedding quality takes priority over a smaller weight footprint.
For local Apple Silicon planning, weight storage is approximately 4.05 GB at 4-bit, estimated (params x 0.5 bytes) from the cited 8.1B parameters.[1] Treat that as a weight-storage estimate, not a complete Mac RAM requirement. Published BF16 weights occupy 16.3 GB.[1] The caveat: Mac runtime compatibility, total memory requirements and throughput remain unverified here.
Ranking note: eligibility is limited to recent embedding releases tagged sentence-similarity or feature-extraction; rerankers are excluded. Ordering uses comparable public benchmark scores, with popularity only breaking ties; an older scored model cannot outrank a newer model with no public benchmark yet. Weight size does not determine rank.
2. Qwen3-VL-Embedding-2B
Qwen3-VL-Embedding-2B by Qwen ranks second with a mean main score of 58.2 across 100 published MTEB task results, behind Qwen’s Qwen3-VL-Embedding-8B at 62.2.[3] The ordering follows published benchmark scores among eligible embedding models, with rerankers excluded; model size does not determine placement.
The model has 2.1B parameters, a 256K-token context window and an Apache-2.0 license.[4] For an Apple Silicon Mac, its raw 4-bit weight memory is approximately 1.05 GB, estimated (params x 0.5 bytes) from the cited 2.1B parameters; published BF16 weights occupy 4.3 GB.[4] Treat the weight estimate as a starting point for a unified-memory budget, rather than a complete RAM requirement.
Consider it for local retrieval when you want the Qwen family’s context allowance with a modest estimated weight footprint.[4] The caveat is hardware sizing: the weight calculation does not establish total runtime memory, full-context feasibility or measured speed on a particular Mac.
3. Nemotron-3-Embed-1B-BF16
Nemotron-3-Embed-1B-BF16 by nvidia ranks third in the eligible lineup, with a published MTEB mean main score of 54.2 across 100 task results.[3] Its position follows benchmark scores, rather than parameter count or popularity. Consider it for local document retrieval when you want a compact model with a long input window.
The model supports a context length of 256K tokens and uses the openmdw-1.1 license.[5] With 1.1B parameters, its 4-bit weight footprint is approximately 0.55 GB, estimated (params x 0.5 bytes).[5] The published BF16 weights occupy 2.3 GB.[5] For an Apple Silicon Mac, treat the quantized figure as a starting point for memory planning. Budget additional unified memory for the runtime and input processing before deciding whether your machine has enough capacity.
The caveat is that the published MTEB score does not establish Apple Silicon throughput or quality after quantization.[3] Validate retrieval quality and memory use with your intended runtime and documents.
4. jina-embeddings-v5-text-small
jina-embeddings-v5-text-small by jinaai ranks fourth on the supplied MTEB comparison, with a mean main score of 45.2 across 100 published task results.[3] Published on January 22, 2026, the model pairs a compact parameter count with a text-focused embedding design.[8][9]
The model has 596M parameters, a 32K-token context window and 1.4 GB of BF16 weights.[8] For local Mac memory planning, its 4-bit weight footprint is approximately 298 MB, estimated (params x 0.5 bytes) from the cited 596M parameters.[8] Treat that figure as a weight budget: a specific Mac RAM requirement also needs runtime measurements. The estimate does not establish Apple Silicon throughput or memory use at the full context limit.
Consider the model for noncommercial local text-retrieval experiments where weight storage matters.[8][9] The practical caveat is its cc-by-nc-4.0 license: check the noncommercial restriction before choosing it for a commercial application.[8]
5. harrier-oss-v1-0.6b
harrier-oss-v1-0.6b by microsoft ranks fifth because its mean main score of 33.5 across 100 published MTEB task results places it behind the higher-scoring candidates in this comparison.[3] Its 596M parameters, 32K-token context window and MIT license make it worth considering for local retrieval projects with limited memory budgets.[7]
For Apple Silicon Macs, weight storage at 4-bit is approximately 0.298 GB, estimated (params x 0.5 bytes) from the cited 596M parameters.[7] Published BF16 weights occupy 1.2 GB.[7] Budget unified memory beyond the weights for the runtime and processing workload. The weight estimate alone cannot establish which Mac configuration will run a given workload comfortably.
Consider it for a local document-search index when permissive licensing and a compact model matter to your deployment.[7] The caveat is retrieval quality: its published MTEB score supports its position here, but does not establish accuracy on your documents or execution speed on your Mac.[3]
6. harrier-oss-v1-270m
harrier-oss-v1-270m by microsoft ranks sixth on the published MTEB comparison, with a mean main score of 29.6 across 100 task results.[3] Its appeal is a compact weight footprint and MIT license.[14] The model has 268M parameters, supports a 32K-token context, and provides 0.5 GB of BF16 weights.[14] Consider it for local retrieval projects where memory budget and licensing matter more than benchmark position.
For an Apple Silicon Mac, weight memory at 4-bit is 134 MB, estimated (params x 0.5 bytes) from the 268M parameter count.[14] Treat that figure as a weight-storage budget when choosing hardware; runtime memory requirements remain unestablished. A specific minimum Mac configuration cannot be justified here. The practical caveat is that its published MTEB result does not establish Mac throughput or confirm compatibility with your chosen local runtime.[3] Verify runtime support before committing to an implementation.
7. jina-embeddings-v5-omni-small
jina-embeddings-v5-omni-small by jinaai ranks here because its published MTEB mean main score of 16.6 across 100 task results falls below the other scored candidates in this comparison.[3] Its appeal is multimodal embedding: the architecture uses frozen-tower composition to preserve text geometry while adding multimodal representations.[11]
The model has 1.6B parameters and a 32K-token context window, with BF16 weights listed at 3.4 GB.[10] Quantized weight storage would be approximately 0.8 GB, estimated (params x 0.5 bytes) from the cited 1.6B parameters.[10] For an Apple Silicon Mac, treat that weight-only estimate as a starting point for memory planning. A local deployment still needs runtime memory beyond the weights; the estimate does not establish a sufficient unified-memory capacity.
Consider it for multimodal retrieval experiments where preserving text representations matters.[11] The licensing caveat is cc-by-nc-4.0: check whether its terms suit your intended deployment before committing to the model.[10]
8. llama-nemotron-embed-1b-v2
llama-nemotron-embed-1b-v2 by nvidia ranks eighth as an older, unscored candidate: its October 16, 2025 release predates nvidia’s Nemotron-3-Embed-1B-BF16, published July 14, 2026.[12][5] Its benchmark status is “no public benchmark yet” (MTEB results, archived date unknown), so its position does not establish weaker retrieval quality.[3] The model has 1.2B parameters, a 128K-token context window, and 2.5 GB of BF16 weights under the nvidia-open-model-license.[12]
For local use on an Apple Silicon Mac, the weight budget at 4-bit is approximately 0.6 GB—estimated (params x 0.5 bytes) from the 1.2B parameter count.[12] Treat that estimate as a weight-storage allowance, not a total unified-memory requirement. Consider the model for embedding long documents when its 128K-token context window matches your workload.[12] The practical caveat is that a weight estimate alone cannot establish which Mac configuration will run your intended workload well.
How can you estimate embedding weight memory from parameter counts?
Estimate embedding weight memory by multiplying the published parameter count by 0.5 bytes per parameter for 4-bit weights, labeling the result “estimated (params x 0.5 bytes).”[1] Use the parameter count rather than the rounded size embedded in a model’s name.
Qwen’s Qwen3-VL-Embedding-8B (Qwen) has 8.1B parameters: its weight memory is approximately 4.05 GB, estimated (params x 0.5 bytes).[1] Qwen’s Qwen3-VL-Embedding-2B (Qwen) has 2.1B parameters, giving approximately 1.05 GB, estimated (params x 0.5 bytes).[4] Both calculations express memory in decimal gigabytes.
NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) has 1.1B parameters, giving approximately 0.55 GB, estimated (params x 0.5 bytes).[5] Microsoft’s harrier-oss-v1-270m (microsoft) has 268M parameters, giving approximately 0.134 GB, estimated (params x 0.5 bytes).[14] Applying the same calculation makes weight budgets directly comparable across model families.
For comparison, the published BF16 weights for Qwen3-VL-Embedding-8B (Qwen) occupy 16.3 GB; the calculated 4-bit weight budget is approximately one quarter of that size.[1] Keep the estimate labeled as weight memory, rather than total application memory or required Mac unified memory. Use it as a starting budget when selecting a model, without treating the calculation as proof of hardware fit or local performance.
Which licenses should you check before choosing an embedding model?
Check the exact model’s license before choosing: the candidates use Apache-2.0,[1][4] MIT,[7][14] cc-by-nc-4.0,[8][10] openmdw-1.1,[5] or nvidia-open-model-license.[12] Make license review part of model selection, alongside memory requirements and retrieval quality.
Qwen’s Qwen3-VL-Embedding-8B (Qwen) and Qwen3-VL-Embedding-2B (Qwen) use Apache-2.0.[1][4] Microsoft’s harrier-oss-v1-0.6b (microsoft) and harrier-oss-v1-270m (microsoft) use MIT.[7][14] Record the license attached to the specific repository you plan to deploy, and check its terms against your intended use.
Jina AI’s jina-embeddings-v5-text-small (jinaai) and jina-embeddings-v5-omni-small (jinaai) both use cc-by-nc-4.0.[8][10] For a commercial application, put those terms through review before investing in integration. Ask whether your planned use, modifications and distribution are permitted; do not treat downloadable weights as permission for every deployment.
NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) uses openmdw-1.1,[5] while NVIDIA’s llama-nemotron-embed-1b-v2 (nvidia) uses nvidia-open-model-license.[12] Review each separately rather than carrying a licensing decision from one NVIDIA model to another.
For an Apple Silicon deployment, keep the repository identifier and license text with your model configuration. Before shipping an application or distributing a converted checkpoint, check the applicable terms for that activity, including any notice or attribution requirements. Resolve licensing questions before committing to a model, then evaluate its practical fit on your Mac.
What do published embedding benchmarks establish about Apple Silicon performance?
Published embedding benchmarks establish embedding-quality comparisons, but the reported scores do not establish Apple Silicon speed, runtime memory use or performance after quantization.[3]
Qwen’s Qwen3-VL-Embedding-8B records a mean main score of 62.2, compared with 58.2 for Qwen’s Qwen3-VL-Embedding-2B and 54.2 for NVIDIA’s Nemotron-3-Embed-1B-BF16, each across 100 published MTEB task results.[3] Those aggregate scores support a quality comparison on that benchmark collection. A higher aggregate score does not establish faster document indexing on a Mac or better retrieval for your own documents.
Weight storage offers a separate planning estimate. Qwen3-VL-Embedding-8B has 8.1 billion parameters:[1] its 4-bit weights would occupy approximately 4.05 GB, estimated (params x 0.5 bytes).[1] Nemotron-3-Embed-1B-BF16 has 1.1 billion parameters:[5] its corresponding weight footprint would be approximately 0.55 GB, estimated (params x 0.5 bytes).[5] Neither calculation establishes total application memory or proves that a particular Mac can run the model.
NVIDIA’s llama-nemotron-embed-1b-v2[12] has no public benchmark yet in the supplied MTEB results snapshot; the archive date is unknown.[3] Missing results should remain an evidence gap, rather than become an assumed quality score.
For a Mac deployment, use the published scores to choose candidates for evaluation. Measure indexing throughput, query latency, peak memory and retrieval quality with your intended runtime, quantization and documents before making a hardware-specific recommendation.
Frequently Asked Questions
Which embedding model should I start with on an Apple Silicon Mac?
Qwen’s Qwen3-VL-Embedding-8B (Qwen) ranks first because its MTEB mean main score of 62.2 is the highest among the scored candidates in the supplied comparison.[3] Its 8.1B parameters imply 4.05 GB of quantized weights, estimated (params x 0.5 bytes), and its license is Apache-2.0.[1] Use that ranking to select a retrieval candidate, then validate execution and runtime memory on your Mac.
How are the embedding models ranked?
Ranking note: only embedding models with Hugging Face pipeline tags sentence-similarity or feature-extraction qualify; rerankers are excluded. Apply the recency gate, then order comparable public MTEB scores; never place an older scored model above a newer unscored model.[3] Popularity only breaks ties, and size does not determine order. NVIDIA’s llama-nemotron-embed-1b-v2 (nvidia) has no public benchmark yet (MTEB results, archived date unknown).[3][12]
How much memory should I budget for quantized embedding weights?
Qwen’s Qwen3-VL-Embedding-2B (Qwen) has 2.1B parameters: 1.05 GB at 4-bit, estimated (params x 0.5 bytes).[4] NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) has 1.1B parameters: 0.55 GB estimated (params x 0.5 bytes).[5] Microsoft’s harrier-oss-v1-270m (microsoft) has 268M parameters: 0.134 GB estimated (params x 0.5 bytes).[14] Use these weight estimates as a starting point for memory planning, and measure total runtime memory before deciding whether a model fits your Mac.
How do the Jina and Microsoft alternatives compare on retrieval scores?
jinaai’s jina-embeddings-v5-text-small (jinaai) scores 45.2, Microsoft’s harrier-oss-v1-0.6b (microsoft) scores 33.5, Microsoft’s harrier-oss-v1-270m (microsoft) scores 29.6, and jinaai’s jina-embeddings-v5-omni-small (jinaai) scores 16.6.[3] Each figure is the mean main score across the model’s 100 published MTEB task results.[3] Their benchmark ordering follows those scores. For a Mac deployment, evaluate that ordering alongside weight memory and license requirements, then check retrieval quality on your own documents.
Which models support long document inputs?
Qwen3-VL-Embedding-8B (Qwen) and Qwen3-VL-Embedding-2B (Qwen) support 256K-token contexts.[1][4] NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) also supports 256K tokens, while llama-nemotron-embed-1b-v2 (nvidia) supports 128K tokens.[5][12] The listed Microsoft Harrier and jinaai models support 32K tokens.[7][8][10][14] Choose a context allowance around your document workflow, and validate memory use with representative inputs before committing to a deployment. Advertised context limits alone do not establish practical Mac capacity.
What licenses should I check before using these models in a product?
The listed Qwen models use Apache-2.0, and the listed Microsoft Harrier models use MIT.[1][4][7][14] jina-embeddings-v5-text-small (jinaai) and jina-embeddings-v5-omni-small (jinaai) use cc-by-nc-4.0.[8][10] NVIDIA’s Nemotron-3-Embed-1B-BF16 (nvidia) uses openmdw-1.1, while llama-nemotron-embed-1b-v2 (nvidia) uses nvidia-open-model-license.[5][12] Review the applicable license against your intended product use before selecting a model. Keep license suitability separate from retrieval score when making the deployment decision.
Sources
- Qwen/Qwen3-VL-Embedding-8B model card (Hugging Face) — 2026-09-25
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — 2026-01-08
- MTEB results (mteb/results, revision 2026-09-25) — 2026-09-25
- Qwen/Qwen3-VL-Embedding-2B model card (Hugging Face) — 2026-09-25
- nvidia/Nemotron-3-Embed-1B-BF16 model card (Hugging Face) — 2026-09-25
- Compact Language Models via Pruning and Knowledge Distillation — 2024-07-19
- microsoft/harrier-oss-v1-0.6b model card (Hugging Face) — 2026-09-25
- jinaai/jina-embeddings-v5-text-small model card (Hugging Face) — 2026-09-25
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation — 2026-02-17
- jinaai/jina-embeddings-v5-omni-small model card (Hugging Face) — 2026-09-25
- jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition — 2026-05-08
- nvidia/llama-nemotron-embed-1b-v2 model card (Hugging Face) — 2026-09-25
- NV-Retriever: Improving text embedding models with effective hard-negative mining — 2024-07-22
- microsoft/harrier-oss-v1-270m model card (Hugging Face) — 2026-09-25