Best Image Embedding Models for Apple Silicon Macs in 2026: Mage-ViT and VTP-Large-f16d64

Rankings 2026-09-29 Last updated 2026-09-29 14 min read By Q4KM

Quick Answer

Mage-ViT (Microsoft) is the top pick as of September 2026 because its release is newest; this ranking admits only image feature extraction models from labs with a published paper or leaderboard record and uses newest release first, then downloads, with no public benchmark yet [6][7][9]. In order, the ranking is Mage-ViT (Microsoft), VTP-Large-f16d64 (MiniMaxAI), VTP-Base-f16d64 (MiniMaxAI), DINOv2 Small (facebook), DINOv2 Base (facebook), and Vision Transformer (ViT) Base, patch16-224, ImageNet-21k (Google) [6][7].

Key Takeaways

How do image embedding models compare on parameters, estimated weight memory, licenses and public benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
Mage-ViT [6] Microsoft 316M [6] 4-bit weights: 158 MB, estimated (316M params × 0.5 bytes) [6] 2026-07-26 [6] MIT [6] no public benchmark yet
VTP-Large-f16d64 [9] MiniMaxAI 732M [9] 4-bit weights: 366 MB, estimated (732M params × 0.5 bytes) [9] 2025-12-16 [9] modified-mit [9] no public benchmark yet
VTP-Base-f16d64 [7] MiniMaxAI 296M [7] 4-bit weights: 148 MB, estimated (296M params × 0.5 bytes) [7] 2025-12-16 [7] modified-mit [7] no public benchmark yet
DINOv2 Small — established pick [1] facebook 22M [1] 4-bit weights: 11 MB, estimated (22M params × 0.5 bytes) [1] 2023-07-31 [1] Apache-2.0 [1] no public benchmark yet
DINOv2 Base — established pick [3] facebook 87M [3] 4-bit weights: 43.5 MB, estimated (87M params × 0.5 bytes) [3] 2023-07-17 [3] Apache-2.0 [3] no public benchmark yet
Vision Transformer (ViT) Base, patch16-224, ImageNet-21k — established pick [4] Google 86M [4] 4-bit weights: 43 MB, estimated (86M params × 0.5 bytes) [4] 2022-03-02 [4] Apache-2.0 [4] no public benchmark yet

Which image embedding models should you consider for an Apple Silicon Mac?

1. Mage-ViT

Mage-ViT by Microsoft ranks first because its Hugging Face publication date is July 26, 2026, later than the other recent candidates’ publication dates.[6][7][9] Mage-ViT has no public benchmark yet, so its position does not establish superior embedding quality or Mac performance. The ordering basis is newest release first, then downloads. The ranking admits only image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s image-feature-extraction pipeline tag.

Mage-ViT has 316M parameters, a listed BF16 weight size of 0.6 GB, and an MIT license.[6] A hypothetical 4-bit weight payload is 158 MB, estimated (params x 0.5 bytes) from the cited 316M parameters.[6] That calculation covers weights alone. Before choosing an Apple Silicon Mac configuration, verify runtime support and measure total unified-memory use; the estimate does not establish how much RAM a working application needs.

Use Mage-ViT as a candidate for evaluating recent image feature extraction weights in a local workflow. The practical caveat is that a calculated weight footprint does not demonstrate an available quantized implementation or successful Apple Silicon execution. Validate loading, image preprocessing, and embedding usefulness on your intended workload before committing to deployment. No supported image-input limit or verified minimum Mac configuration is specified.

2. VTP-Large-f16d64

VTP-Large-f16d64 by MiniMaxAI ranks second behind Microsoft’s Mage-ViT under the ordering rule: newest release first, then downloads.[6][9] Eligibility is limited to image feature extraction models from labs with a published paper or leaderboard record. VTP-Large-f16d64 shares its publication date of December 16, 2025, with MiniMaxAI’s VTP-Base-f16d64 and takes precedence through higher download counts.[7][9] The benchmark status is no public benchmark yet; its position does not establish superior embedding quality.

VTP-Large-f16d64 has 732M parameters and published F32 weights of 2.9 GB, under the modified-mit license.[9] For Apple Silicon memory planning, its 4-bit weight footprint is 366 MB, estimated (params x 0.5 bytes) from the cited 732M parameters.[9] Treat that estimate as a weight-storage budget; verify total memory consumption with the intended local runtime before choosing a Mac configuration. A supported quantized execution path and minimum Mac memory requirement are not established.

Consider the model for visual-tokenizer experiments connected to image generation, matching the focus of MiniMaxAI’s paper, “Towards Scalable Pre-training of Visual Tokenizers for Generation.”[8] For a local image-search application, evaluate retrieval quality on representative images before adopting it. The practical caveat is that Apple Silicon execution remains unverified: the parameter count alone cannot establish compatibility, speed or end-to-end memory use.

3. VTP-Base-f16d64

VTP-Base-f16d64 by MiniMaxAI ranks third under the ordering rule “newest release first, then downloads”: its Hugging Face publication date matches MiniMaxAI’s VTP-Large-f16d64, but its recorded download count is lower.[7][9] Microsoft’s Mage-ViT has a later publication date.[6] Ranking note: eligibility is limited to image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s image-feature-extraction tag. VTP-Base-f16d64 has no public benchmark yet, so its position does not establish an embedding-quality advantage.[7]

The model has 296M parameters, distributed as F32 weights totaling 1.2 GB, under the modified-mit license.[7] At 4-bit precision, weight storage would be 148 MB, estimated (params x 0.5 bytes) from the cited 296M parameters.[7] An Apple Silicon Mac needs additional unified memory for execution beyond the weights. Treat that calculation as a weight-storage budget; a supported quantized implementation and measured runtime memory remain necessary before choosing hardware.

Evaluate VTP-Base-f16d64 for image feature extraction in projects involving visual tokenizers for generation, the focus of MiniMaxAI’s accompanying paper, “Towards Scalable Pre-training of Visual Tokenizers for Generation.”[8] The practical caveat is deployment uncertainty: no validated Apple Silicon runtime, total memory requirement, or context specification is established here. Check runtime support and input requirements before committing to a local pipeline, and measure feature usefulness on your own images.

4. DINOv2 Small

DINOv2 Small by facebook ranks fourth as an established pick, with newer releases placed ahead of it.[1][6][7][9] The ordering basis is newest release first, then downloads, with established picks admitted through a recency exemption. Its Hugging Face publication date is July 31, 2023, and it recorded 2,985,175 downloads during the reporting window ending September 26, 2026.[1] Its benchmark status is “no public benchmark yet”; download volume does not demonstrate embedding quality.

DINOv2 Small has 22M parameters, uses the Apache-2.0 license, and lists F32 weights at 0.1 GB.[1] Quantized weight storage would be approximately 11 MB, estimated (params x 0.5 bytes) from the 22M parameter count.[1] That calculation covers weights alone. A Mac’s required unified memory also depends on the runtime and its additional allocations, so the estimate cannot establish a hardware requirement or confirm that a compatible quantized implementation is available.

Consider DINOv2 Small an established baseline for a local image-feature-extraction evaluation. The model belongs to the family described in facebook’s “DINOv2: Learning Robust Visual Features without Supervision.”[2] The practical caveat is the gap between compact weight storage and demonstrated Apple Silicon execution: validate runtime compatibility, input handling, and peak memory before choosing hardware or committing it to an application.

5. DINOv2 Base

DINOv2 Base by facebook ranks fifth as an established pick, retaining a place through the family’s download-based exemption from the recency gate.[3] The ordering basis is newest release first, then downloads. Its Hugging Face publication date of July 17, 2023 places it behind facebook’s DINOv2 Small, published July 31, 2023.[3][1] DINOv2 Base has no public benchmark yet for this ranking; its position does not establish an embedding-quality advantage.

DINOv2 Base contains 87 million parameters and ships with 0.3 GB of F32 weights under the Apache-2.0 license.[3] For those 87 million parameters, 4-bit weight storage is 43.5 MB, estimated (params x 0.5 bytes).[3] Treat that calculation as a weight-storage budget: a usable Apple Silicon configuration still requires verified runtime support and measured total memory use. The estimate alone cannot establish which Mac or RAM configuration will run the model successfully.

Choose DINOv2 Base for local image feature extraction when its Apache-2.0 license suits your project; the family’s published work focuses on learning robust visual features without supervision.[3][2] The ranking admits only image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s image-feature-extraction pipeline tag. The practical caveat is that input limits, quantized execution support and measured Apple Silicon performance are unspecified, so hardware suitability remains unverified.

6. Vision Transformer (ViT) Base, patch16-224, ImageNet-21k

Vision Transformer (ViT) Base, patch16-224, ImageNet-21k by Google is an established pick in this ranking.[4] Its Hugging Face publication date is March 2, 2022,[4] placing it after the more recently published candidates under the ordering rule: newest release first, then downloads. Its established-pick status permits inclusion outside the recency window.[4] Benchmark status: no public benchmark yet; its position does not establish comparative embedding quality.

Google lists 86M parameters, F32 weights of 0.3 GB and an Apache-2.0 license.[4] For those 86M parameters, a 4-bit weight payload would be approximately 43 MB, estimated (params x 0.5 bytes).[4] Treat that calculation as a weight-storage budget when planning an Apple Silicon deployment. Total memory requirements and a suitable Mac configuration remain unverified; the estimate does not establish that a compatible quantized implementation is available.

Consider the model as an established image feature extraction baseline for a local evaluation, particularly when Apache-2.0 licensing suits the project.[4] The architecture is documented in Google’s paper, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.”[5] The practical caveat is deployment uncertainty: Apple Silicon throughput, peak runtime memory and supported input limits are not established here. Validate those requirements in the intended runtime before choosing hardware.

How can you estimate image embedding model weight memory from parameter counts?

Estimate image embedding model weight memory by multiplying the parameter count by the assumed storage per parameter: for 4-bit weights, report the result as estimated (params x 0.5 bytes) alongside the cited parameter count.[6] The calculation estimates weight storage, not total memory required to run the model.

Microsoft’s Mage-ViT (Mage-ViT (Microsoft)) has 316M parameters.[6] Its quantized weight memory is approximately 158 MB, estimated (params x 0.5 bytes).[6] Keep that calculation separate from the published BF16 weight size of 0.6 GB; the published size describes the available weights, while the calculation assumes a different storage precision.[6]

MiniMaxAI’s VTP-Base-f16d64 has 296M parameters, giving approximately 148 MB, estimated (params x 0.5 bytes).[7] MiniMaxAI’s VTP-Large-f16d64 has 732M parameters, giving approximately 366 MB, estimated (params x 0.5 bytes).[9] Use the same assumption across models so the comparison stays consistent. Label the results as weight estimates wherever they appear, including comparison tables.

For an Apple Silicon Mac, use those estimates as a starting point for a memory budget. A parameter calculation alone cannot establish whether a model fits in available unified memory, supports the intended quantization method, or runs well in a particular runtime. Before choosing a model, check its runtime support and measure total memory with your intended image inputs and batch settings.

Which licenses do the image embedding models use?

The image embedding models use MIT for Microsoft’s Mage-ViT, modified-mit for MiniMaxAI’s VTP-Base-f16d64 and VTP-Large-f16d64, and Apache-2.0 for Facebook’s DINOv2 Small and DINOv2 Base and Google’s ViT Base model, google/vit-base-patch16-224-in21k.[6][7][9][1][3][4]

Mage-ViT is the MIT-licensed option in this comparison.[6] For an engineering evaluation, record that license alongside the exact repository and revision you intend to deploy. Keep license review separate from decisions about memory use or embedding quality.

MiniMaxAI lists both VTP variants under modified-mit.[7][9] Preserve that exact label in your dependency inventory rather than shortening it to MIT. Review the actual license text before deciding whether either model suits your intended deployment; do not assume the modification has no effect on your use case.

Facebook’s DINOv2 Small and DINOv2 Base share the Apache-2.0 designation, as does Google’s listed ViT model.[1][3][4] Teams evaluating these candidates can group them by that declared license, while keeping a separate record for each model.

For a local Mac application, use these labels as an initial license filter. Before distributing model weights with your application, review the selected repository’s terms for your planned packaging and distribution.

How should release dates and downloads determine the order when there is no public benchmark yet?

When no public benchmark yet scores the candidates against each other, order them by newest release first, then downloads as a tiebreaker—not by parameter count or assumed performance on Apple Silicon.[1][3][4][6][7][9]

Microsoft’s Mage-ViT (Microsoft) takes first place because its Hugging Face publication date, July 26, 2026, is the newest among the eligible candidates.[1][3][4][6][7][9] Its position reflects release timing; no public benchmark yet establishes a performance advantage.

MiniMaxAI’s VTP-Large-f16d64 (MiniMaxAI) comes next, followed by MiniMaxAI’s VTP-Base-f16d64 (MiniMaxAI). Both were first published on December 16, 2025.[7][9] Downloads break that tie: Large recorded 962 downloads over the preceding 30 days as of September 25, 2026; Base recorded 181 as of September 26, 2026.[7][9] The snapshots differ by a day, so treat popularity as a dated tiebreaker.

Facebook’s facebook/dinov2-small, Facebook’s facebook/dinov2-base and Google’s google/vit-base-patch16-224-in21k follow as established picks, in that order. Their publication dates are July 31, 2023; July 17, 2023; and March 2, 2022, respectively.[1][3][4] Each has no public benchmark yet for this comparison.

Ranking note: newest release first, then downloads. Scope admits only image feature extraction models tagged image-feature-extraction from labs with a published paper or leaderboard record. Established picks qualify through their family’s three most-downloaded models and bypass the twelve-month recency gate, but that exemption cannot give them first place.[1][3][4] Download totals do not establish embedding quality, memory fit or local execution speed.

Frequently Asked Questions

Which image embedding model should I try first on an Apple Silicon Mac?

Microsoft’s Mage-ViT takes first place because its July 2026 publication is newer than the December 2025 MiniMaxAI releases under the release-date ordering rule.[6][7][9] Mage-ViT has 316M parameters and an MIT license.[6] Its benchmark status is “no public benchmark yet.” Treat the ranking as a starting point for evaluation, not proof of superior retrieval quality or Apple Silicon performance.

How are the models ranked?

Ranking note: newest release first, then downloads.[6][9][7][1][3][4] The order is Microsoft’s Mage-ViT,[6] MiniMaxAI’s VTP-Large-f16d64,[9] MiniMaxAI’s VTP-Base-f16d64,[7] facebook’s dinov2-small (established pick),[1] facebook’s dinov2-base (established pick),[3] and Google’s vit-base-patch16-224-in21k (established pick).[4] Scope: this ranking admits only image feature extraction models from labs with a published paper or leaderboard record (Hugging Face pipeline tag: image-feature-extraction). Established picks bypass the recency gate; the exemption cannot award first place. Model size does not determine placement.

How much memory would quantized weights need?

Mage-ViT has 316M parameters:[6] its 4-bit weight payload is 158 MB, estimated (params x 0.5 bytes).[6] VTP-Large-f16d64 has 732M parameters:[9] its corresponding payload is 366 MB, estimated (params x 0.5 bytes).[9] Treat these calculations as weight-storage estimates only. Neither calculation establishes a working quantized implementation, total runtime memory, or a Mac RAM requirement.

Which licenses do these models use?

Mage-ViT uses MIT.[6] VTP-Base-f16d64 and VTP-Large-f16d64 use modified-mit.[7][9] The established picks—dinov2-small, dinov2-base and vit-base-patch16-224-in21k—use Apache-2.0.[1][3][4] For deployment, inspect the exact license text attached to your chosen repository. In particular, do not assume modified-mit grants identical permissions or imposes identical conditions to MIT merely because the names overlap.

Why are older DINOv2 and Vision Transformer models included?

The established-pick exception admits models among their family’s three most-downloaded Hugging Face entries despite the twelve-month recency gate.[1][3][4] facebook’s dinov2-small and dinov2-base were published in July 2023; Google’s vit-base-patch16-224-in21k was published in March 2022.[1][3][4] Their established status explains inclusion, not superior quality. The exemption cannot place an older model first, and download counts serve only as a tiebreak.

Can I choose confidently from benchmark scores and context limits?

No public benchmark scores these candidates against one another, so each carries the status “no public benchmark yet.” Context limits and measured Apple Silicon performance remain unverified in this comparison. Before committing, evaluate your intended image collection and runtime. Check retrieval relevance, preprocessing requirements, supported image dimensions and actual memory use; parameter counts alone cannot establish those results.

Sources

  1. facebook/dinov2-small model card (Hugging Face) — 2026-09-26
  2. DINOv2: Learning Robust Visual Features without Supervision — 2023-04-14
  3. facebook/dinov2-base model card (Hugging Face) — 2026-09-26
  4. google/vit-base-patch16-224-in21k model card (Hugging Face) — 2026-09-26
  5. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — 2020-10-22
  6. microsoft/Mage-ViT model card (Hugging Face) — 2026-09-26
  7. MiniMaxAI/VTP-Base-f16d64 model card (Hugging Face) — 2026-09-26
  8. Towards Scalable Pre-training of Visual Tokenizers for Generation — 2025-12-15
  9. MiniMaxAI/VTP-Large-f16d64 model card (Hugging Face) — 2026-09-26

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog