Quick Answer
Mage-ViT (microsoft) is the top pick for CPU-only PCs as of September 2026 because it is the newest eligible release, with no public benchmark yet [6][7][9]. In order, the ranking is Mage-ViT (microsoft), VTP-Large-f16d64 (MiniMaxAI), VTP-Base-f16d64 (MiniMaxAI), DINOv2 Small (facebook), DINOv2 Base (facebook), and Vision Transformer Base (vit-base-patch16-224-in21k) (google) [6][7].
Key Takeaways
- Microsoft’s Mage-ViT takes first place because its July release is the newest among the eligible candidates, not because of a demonstrated CPU performance advantage. Its 316M parameters imply 158 MB of 4-bit weights, estimated (params x 0.5 bytes); its license is MIT. [6]
- MiniMaxAI’s VTP-Large-f16d64 precedes MiniMaxAI’s VTP-Base-f16d64 because their release dates match and Large has more downloads. Their respective 732M and 296M parameters imply 366 MB and 148 MB of 4-bit weights, estimated (params x 0.5 bytes); both use modified-mit licenses. [9][7]
- Facebook’s DINOv2 Small and Facebook’s DINOv2 Base are established picks under Apache-2.0. Small’s 22M parameters imply 11 MB of 4-bit weights, estimated (params x 0.5 bytes); Base’s 87M parameters imply 43.5 MB, estimated (params x 0.5 bytes). [1][3]
- Google’s Vision Transformer Base (vit-base-patch16-224-in21k) is an established pick under Apache-2.0. Its 86M parameters imply 43 MB of 4-bit weights, estimated (params x 0.5 bytes). [4]
- Benchmark status across the ranking is no public benchmark yet; the order does not establish comparative CPU speed or embedding quality. Weight estimates describe parameter storage, not measured runtime RAM requirements. [6][9][7][1][3][4]
- Ranking note: newest release first, then downloads. Scope admits only image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s
image-feature-extractionpipeline tag. DINOv2 Small, DINOv2 Base and Vision Transformer Base are established picks that bypass the recency gate; the exemption does not grant first place, and model size does not determine order. [6][9][7][1][3][4]
How do image embedding models compare on parameters, estimated weight memory, licenses and public benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Mage-ViT [6] | microsoft | 316M [6] | 4-bit weights only: 158 MB estimated (params x 0.5 bytes) [6] | 2026-07-26 [6] | MIT [6] | no public benchmark yet |
| VTP-Large-f16d64 [9] | MiniMaxAI | 732M [9] | 4-bit weights only: 366 MB estimated (params x 0.5 bytes) [9] | 2025-12-16 [9] | modified-mit [9] | no public benchmark yet |
| VTP-Base-f16d64 [7] | MiniMaxAI | 296M [7] | 4-bit weights only: 148 MB estimated (params x 0.5 bytes) [7] | 2025-12-16 [7] | modified-mit [7] | no public benchmark yet |
| DINOv2 Small — established pick [1] | 22M [1] | 4-bit weights only: 11 MB estimated (params x 0.5 bytes) [1] | 2023-07-31 [1] | Apache-2.0 [1] | no public benchmark yet | |
| DINOv2 Base — established pick [3] | 87M [3] | 4-bit weights only: 43.5 MB estimated (params x 0.5 bytes) [3] | 2023-07-17 [3] | Apache-2.0 [3] | no public benchmark yet | |
| Vision Transformer Base (vit-base-patch16-224-in21k) — established pick [4] | 86M [4] | 4-bit weights only: 43 MB estimated (params x 0.5 bytes) [4] | 2022-03-02 [4] | Apache-2.0 [4] | no public benchmark yet |
Which image embedding models should you consider for a CPU-only PC?
1. Mage-ViT
Mage-ViT by microsoft ranks first because its Hugging Face publication date makes it the newest release in this selection: July 26, 2026.[6] The ordering basis is newest release first, then downloads; the position does not establish superior embedding quality or CPU performance. The ranking admits only image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s image-feature-extraction pipeline tag.
Mage-ViT has 316M parameters, lists 0.6 GB of BF16 weights, and uses the MIT license.[6] Its estimated 4-bit weight memory is 158 MB—estimated (params x 0.5 bytes), using the cited 316M parameter count.[6] Treat that calculation as a weight-storage estimate, not a system RAM requirement or confirmation that a compatible quantized checkpoint exists. A specific CPU model or minimum RAM capacity cannot be recommended from the available specifications.
Consider Mage-ViT for evaluating local image feature extraction when its MIT license suits your project and you can validate deployment yourself.[6] The practical caveat is “no public benchmark yet”: its ranking provides no measured assurance about embedding quality or CPU speed. Before committing to a CPU-only installation, confirm runtime compatibility and measure memory consumption and latency on your own machine. Those checks are necessary before calling the model a practical fit.
2. VTP-Large-f16d64
VTP-Large-f16d64 by MiniMaxAI ranks second under the ordering rule “newest release first, then downloads”: Microsoft’s Mage-ViT has a later publication date, while MiniMaxAI’s VTP-Base-f16d64 shares its publication date but has fewer downloads.[6][7][9] VTP-Large-f16d64 was first published on Hugging Face on December 16, 2025, and recorded 962 downloads over the preceding 30 days as of September 25, 2026.[9] The ranking admits only image feature extraction models from labs with a published paper or leaderboard record, using the Hugging Face image-feature-extraction pipeline tag.
The model has 732M parameters and published F32 weights of 2.9 GB, with a modified-mit license.[9] Weight storage at 4-bit is approximately 366 MB, estimated (params x 0.5 bytes) from the cited 732M parameters.[9] Treat that calculation as a weight-storage estimate, not a complete RAM requirement or confirmation of a working quantized CPU implementation. A specific CPU and minimum system RAM configuration cannot be recommended from those figures alone.
Consider VTP-Large-f16d64 for local experiments involving visual tokenizers for generation, the subject of MiniMaxAI’s accompanying paper, “Towards Scalable Pre-training of Visual Tokenizers for Generation.”[8] The caveat for a CPU-only purchase decision is straightforward: no public benchmark yet. Its ranking therefore establishes neither CPU throughput nor embedding-quality superiority. Verify runtime compatibility, input limits and measured memory use before choosing hardware around it.
3. VTP-Base-f16d64
VTP-Base-f16d64 by MiniMaxAI ranks third under the ordering of newest release first, then downloads.[6][7][9] Its Hugging Face publication date is December 16, 2025, behind Microsoft’s Mage-ViT.[6][7] MiniMaxAI’s VTP-Large-f16d64 shares that release date but has 962 downloads in the reported monthly snapshot, compared with 181 for the Base model.[7][9] VTP-Base-f16d64 has no public benchmark yet; its position does not establish an advantage in embedding quality or CPU speed.
The model has 296 million parameters and 1.2 GB of F32 weights.[7] Its theoretical 4-bit weight storage is 148 MB, estimated (params x 0.5 bytes) from the cited 296 million parameters.[7] Treat that calculation as a weight-storage budget, not a total RAM requirement or confirmation that a compatible quantized build exists. For local CPU deployment, verify runtime support and measure total memory use before choosing hardware; the weight figures alone cannot establish a minimum PC configuration.
Consider VTP-Base-f16d64 for image-feature extraction experiments involving visual tokenizers for generation, the focus of MiniMaxAI’s accompanying paper.[8] Evaluate its features on your own images before adopting it for a retrieval workflow. One caveat is the modified-mit license: review its terms for your intended deployment rather than assuming the permissions of an unmodified MIT license.[7]
4. DINOv2 Small
DINOv2 Small by facebook ranks here as an established pick: it is among its family’s three most-downloaded Hugging Face models and bypasses the twelve-month recency gate.[1] Its July 31, 2023 publication places it behind the recent releases under the ordering basis: newest release first, then downloads.[1][6][7][9] The ranking admits only image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s image-feature-extraction pipeline tag.
DINOv2 Small has 22M parameters, published F32 weights of 0.1 GB, and an Apache-2.0 license.[1] Its weight storage at four-bit precision would be approximately 11 MB, estimated (params x 0.5 bytes) from the cited 22M parameters.[1] Treat that estimate as a weight-storage budget when planning a local CPU deployment. Total system RAM requirements, quantized runtime compatibility, and CPU latency are not specified, so the weight figure alone cannot establish which PC will run it comfortably.
Consider DINOv2 Small for local image feature extraction when a compact parameter count and Apache-2.0 licensing suit your project.[1] The model belongs to the family introduced in “DINOv2: Learning Robust Visual Features without Supervision.”[2] The caveat is performance evidence: no public benchmark yet establishes its standing in this comparison. Its position reflects the stated release ordering and established-pick exemption, rather than a demonstrated CPU speed or embedding-quality advantage.
5. DINOv2 Base
DINOv2 Base by facebook ranks fifth as an established pick, with a Hugging Face publication date of July 17, 2023.[3] The ordering basis is newest release first, then downloads; facebook’s DINOv2 Small precedes it by publication date.[1][3] Established picks bypass the recency gate because they are among their family’s three most-downloaded models.[3] The ranking admits only image feature extraction models from labs with a published paper or leaderboard record, using the Hugging Face image-feature-extraction pipeline tag.
DINOv2 Base has 87M parameters, published F32 weights of 0.3 GB, and an Apache-2.0 license.[3] Quantized weight storage would be 43.5 MB, estimated (params x 0.5 bytes) from the cited 87M parameters.[3] Treat that estimate as a weight-storage calculation when planning a CPU-only PC. Actual RAM must also accommodate the runtime and intermediate data; the estimate does not establish a complete hardware requirement or confirm that a compatible quantized implementation is available.
Choose DINOv2 Base as an established starting point for local image feature extraction when the Apache-2.0 license suits your project.[3] The model belongs to the family described in “DINOv2: Learning Robust Visual Features without Supervision.”[2] The benchmark status is “no public benchmark yet.” Its position therefore provides no measured CPU performance advantage; validate runtime compatibility and latency on your own PC before committing to it.
6. Vision Transformer Base (vit-base-patch16-224-in21k)
Vision Transformer Base (vit-base-patch16-224-in21k) by google ranks sixth as an established pick, with a Hugging Face publication date of March 2, 2022.[4] Its established-pick status permits inclusion outside the twelve-month recency gate.[4] The ordering is newest release first, then downloads; popularity serves only as a tiebreak. The ranking admits only image feature extraction models from labs with a published paper or leaderboard record.
The model has 86 million parameters, published F32 weights of 0.3 GB, and an Apache-2.0 license.[4] Weight memory at 4-bit is approximately 43 MB, estimated (params x 0.5 bytes) from the cited 86 million parameters.[4] Treat that calculation as a weight-storage estimate, not a total system RAM requirement or confirmation that a compatible quantized runtime is available. A specific CPU and RAM configuration cannot be recommended from weight size alone.
Consider this established model when you want an Apache-2.0 image-feature baseline for a local evaluation.[4] The associated paper is “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.”[5] The main caveat is performance uncertainty: no public benchmark yet scores this candidate within the ranking. Its position therefore does not establish embedding quality or CPU speed, and the weight estimate does not establish end-to-end memory use.
How can you estimate image embedding model weight memory from parameter counts?
Estimate image embedding model weight memory by multiplying its parameter count by the assumed bytes per parameter, and label the result as an estimate rather than measured RAM use.
For a four-bit weight estimate, use the calculation “estimated (params x 0.5 bytes).” Microsoft’s Mage-ViT (microsoft) has 316M parameters: estimated (params x 0.5 bytes), its weight payload is 158 MB in decimal units.[6] The calculation describes the assumed quantized weights; it does not establish that a compatible quantized release is available.
The same calculation applies across the candidates. Facebook’s facebook/dinov2-small has 22M parameters: estimated (params x 0.5 bytes), its weight payload is 11 MB.[1] MiniMaxAI’s VTP-Large-f16d64 (MiniMaxAI) has 732M parameters: estimated (params x 0.5 bytes), its weight payload is 366 MB.[9] Keep the parameter count beside each estimate so readers can check the arithmetic.
Compare estimates with published weight sizes carefully. Microsoft lists Mage-ViT (microsoft) with BF16 weights of 0.6 GB.[6] At four bits, the assumed weight payload is approximately one quarter of BF16, using 0.5 rather than 2 bytes per parameter; that ratio is an estimate, not a measured file-size reduction.[6]
Use the result to compare theoretical weight storage. A weight-only estimate cannot establish total application RAM requirements or CPU inference speed, so avoid turning it into a claim that a model fits or runs well on a particular PC.
Which image embedding models use Apache-2.0, MIT or modified-mit licenses?
facebook’s facebook/dinov2-small and facebook/dinov2-base, plus google’s google/vit-base-patch16-224-in21k, use Apache-2.0.[1][3][4] microsoft’s Mage-ViT (microsoft) uses MIT.[6] MiniMaxAI’s VTP-Base-f16d64 (MiniMaxAI) and VTP-Large-f16d64 (MiniMaxAI) use modified-mit.[7][9]
The Apache-2.0 group contains the established picks.[1][3][4] facebook/dinov2-small has 22M parameters and a listed F32 weight size of 0.1 GB.[1] facebook/dinov2-base has 87M parameters and 0.3 GB of F32 weights.[3] google/vit-base-patch16-224-in21k has 86M parameters and 0.3 GB of F32 weights.[4] Those weight sizes describe the listed F32 weights, rather than estimated quantized memory requirements.[1][3][4]
The MIT entry, Mage-ViT (microsoft), was first published on Hugging Face on 2026-07-26.[6] Its listed configuration has 316M parameters and 0.6 GB of BF16 weights.[6] Keep that precision label alongside the size when comparing its weight footprint with the Apache-2.0 entries.[1][3][4][6]
The modified-mit entries share a first-publication date of 2025-12-16 on Hugging Face.[7][9] VTP-Base-f16d64 (MiniMaxAI) has 296M parameters and 1.2 GB of F32 weights; VTP-Large-f16d64 (MiniMaxAI) has 732M parameters and 2.9 GB of F32 weights.[7][9] Record their license as modified-mit when documenting a deployment, preserving the distinction from Mage-ViT (microsoft)’s MIT designation.[6][7][9]
Which established image embedding models belong on a CPU-only shortlist?
Facebook’s DINOv2 Small (facebook/dinov2-small), Facebook’s DINOv2 Base (facebook/dinov2-base) and Google’s Vision Transformer Base (google/vit-base-patch16-224-in21k) belong on the CPU-only shortlist as established picks. [1][3][4] All carry the Apache-2.0 license. [1][3][4]
DINOv2 Small has 22M parameters; its quantized weight storage would be 11 MB, estimated (params x 0.5 bytes). [1] Published F32 weights occupy 0.1 GB. [1] Its benchmark status is no public benchmark yet; its position does not establish CPU speed or embedding quality.
DINOv2 Base has 87M parameters; its quantized weight storage would be 43.5 MB, estimated (params x 0.5 bytes). [3] Published F32 weights occupy 0.3 GB. [3] Its benchmark status is no public benchmark yet. The DINOv2 family describes learning visual features without supervision. [2]
Google’s Vision Transformer Base has 86M parameters; its quantized weight storage would be 43 MB, estimated (params x 0.5 bytes). [4] Published F32 weights occupy 0.3 GB. [4] Its benchmark status is no public benchmark yet. Treat the calculated storage figures as weight estimates; CPU latency and total runtime memory still need validation.
Ranking note: this ranking admits only image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s image-feature-extraction pipeline tag. With no public benchmark scoring the candidates, the ordering is newest release first, then downloads. Established picks bypass the recency gate and retain that label. [1][3][4] Microsoft’s Mage-ViT (Mage-ViT (microsoft)) takes the first position because its release is newest among the admitted candidates. [6][7][9][1][3][4]
Frequently Asked Questions
Which image embedding model should I evaluate first on a CPU-only PC?
Microsoft’s Mage-ViT (Mage-ViT (microsoft)) leads because its Hugging Face publication date is newer than those of the other eligible recent releases [6][7][9]. Mage-ViT has 316M parameters and an MIT license [6]. Its benchmark status is no public benchmark yet [6]. Treat its placement as a reason to evaluate it, rather than proof of CPU speed or embedding quality.
How are the image embedding models ranked?
Ranking note: newest release first, then downloads. Scope: only image feature extraction models from labs with a published paper or leaderboard record, using Hugging Face’s image-feature-extraction pipeline tag. Order: Microsoft’s Mage-ViT (microsoft) [6], MiniMaxAI’s VTP-Large-f16d64 (MiniMaxAI) [9], VTP-Base-f16d64 (MiniMaxAI) [7], Facebook’s facebook/dinov2-small [1], facebook/dinov2-base [3], Google’s google/vit-base-patch16-224-in21k [4]. Established picks bypass the recency gate through their family download standing, but cannot lead by exemption [1][3][4]. Parameter count does not determine placement.
How much memory would quantized weights need?
For 4-bit weights, Mage-ViT’s 316M parameters imply 158 MB estimated (params x 0.5 bytes) [6]. VTP-Large-f16d64’s 732M parameters imply 366 MB estimated (params x 0.5 bytes) [9]. DINOv2 Small’s 22M parameters imply 11 MB estimated (params x 0.5 bytes) [1]. Treat these calculations as weight-storage estimates; verify actual process RAM and quantization support before choosing a deployment configuration.
Do published benchmarks establish CPU speed or supported image context?
Every listed candidate has the status no public benchmark yet [1][3][4][6][7][9]. No CPU latency, throughput or supported image-context figure is established for this comparison. Before deployment, check accepted image dimensions, preprocessing requirements and embedding output shape. Measure latency and memory with your intended runtime and images; neither release date nor parameter count establishes practical CPU performance.
Are older DINOv2 and Google Vision Transformer models still worth considering?
Facebook’s DINOv2 Small (facebook/dinov2-small), Facebook’s DINOv2 Base (facebook/dinov2-base), and Google’s Vision Transformer (google/vit-base-patch16-224-in21k) are established picks [1][3][4]. Each qualifies through membership among its family’s three most-downloaded models, bypassing the recency gate [1][3][4]. Their parameter counts are 22M, 87M and 86M, respectively [1][3][4]. Consider them for evaluation, while keeping download popularity separate from embedding quality and measured CPU performance.
What licenses apply to these local image embedding models?
Microsoft’s Mage-ViT uses MIT [6]. MiniMaxAI’s VTP-Base-f16d64 and VTP-Large-f16d64 list modified-mit [7][9]. Facebook’s DINOv2 Small and DINOv2 Base, plus Google’s vit-base-patch16-224-in21k, list Apache-2.0 [1][3][4]. Review the actual license text before redistribution or product integration, particularly the modified-mit terms. A license label alone should not substitute for checking the obligations applicable to your deployment.
Sources
- facebook/dinov2-small model card (Hugging Face) — 2026-09-26
- DINOv2: Learning Robust Visual Features without Supervision — 2023-04-14
- facebook/dinov2-base model card (Hugging Face) — 2026-09-26
- google/vit-base-patch16-224-in21k model card (Hugging Face) — 2026-09-26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — 2020-10-22
- microsoft/Mage-ViT model card (Hugging Face) — 2026-09-26
- MiniMaxAI/VTP-Base-f16d64 model card (Hugging Face) — 2026-09-26
- Towards Scalable Pre-training of Visual Tokenizers for Generation — 2025-12-15
- MiniMaxAI/VTP-Large-f16d64 model card (Hugging Face) — 2026-09-26