Quick Answer
Microsoft’s colipri is the top pick as of September 2026 because it is the newest listed release; the ranking admits only zero-shot image classification models from labs with a published paper or leaderboard record, includes older releases because the family has too few recent candidates, and follows newest release first, then downloads, with every candidate marked “no public benchmark yet” [12] [8] [4] [6] [10] [7] [1] [3]. In order, the ranking is colipri (Microsoft), PE-Core-L14-336 (Facebook), SigLIP 2 Base Patch16-256 (Google), SigLIP 2 Giant Opt Patch16-384 (Google), MobileCLIP-S1-OpenCLIP (Apple), BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 (Microsoft), CLIP ViT-Base-Patch32 (OpenAI), and CLIP ViT-Large-Patch14 (OpenAI) [12][8].
Key Takeaways
- Microsoft colipri takes first place because its publication date, January 14, 2026, is the newest among the listed candidates; no public benchmark yet establishes an accuracy advantage. [12]
- Ranking note: no public benchmark scores any candidate, so the ordering is newest release first, then downloads. The ranking admits only zero-shot image classification models from labs with a published paper or leaderboard record, using Hugging Face’s
zero-shot-image-classificationpipeline tag. Too few candidates were released within the last 12 months, so the newest available releases are included. [1][3][4][6][7][8][10][12] - Established picks are OpenAI CLIP ViT-Base-Patch32, OpenAI CLIP ViT-Large-Patch14 and Google SigLIP 2 Base Patch16-256—the three most-downloaded candidates. Their recency exemption keeps them eligible under the same ordering rule; popularity serves only as a tiebreak and does not earn first place. [1][3][4]
- Weight memory estimates: Microsoft colipri has 258M parameters, giving 129 MB at 4-bit, estimated (params x 0.5 bytes). [12] Google SigLIP 2 Giant Opt Patch16-384 has 1.9B parameters, giving 950 MB at 4-bit, estimated (params x 0.5 bytes). [6] Both estimates cover weights alone.
- License choices differ: Microsoft colipri and Microsoft BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 use MIT; Facebook PE-Core-L14-336 and both Google SigLIP 2 variants use Apache-2.0. [12][7][8][4][6] Apple MobileCLIP-S1-OpenCLIP uses apple-amlr. [10]
- OpenAI CLIP ViT-Base-Patch32 and OpenAI CLIP ViT-Large-Patch14 each have a text context limit of 77 tokens; account for that limit when composing classification prompts. [1][3]
How do local zero-shot image classifiers compare on specs and benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| colipri [12] | Microsoft | 258M [12] | 4-bit weights: 129 MB estimated (params x 0.5 bytes); VRAM not published [12] | 2026-01-14 [12] | MIT [12] | no public benchmark yet |
| PE-Core-L14-336 [8] | — | — | 2025-04-11 [8] | Apache-2.0 [8] | no public benchmark yet | |
| SigLIP 2 Base Patch16-256 [4] | 375M [4] | 4-bit weights: 187.5 MB estimated (params x 0.5 bytes); VRAM not published [4] | 2025-02-17 [4] | Apache-2.0 [4] | no public benchmark yet | |
| SigLIP 2 Giant Opt Patch16-384 [6] | 1.9B [6] | 4-bit weights: 950 MB estimated (params x 0.5 bytes); VRAM not published [6] | 2025-02-17 [6] | Apache-2.0 [6] | no public benchmark yet | |
| MobileCLIP-S1-OpenCLIP [10] | Apple | — | Full-precision weights: 0.3 GB; VRAM not published [10] | 2024-06-07 [10] | apple-amlr [10] | no public benchmark yet |
| BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 [7] | Microsoft | — | — | 2023-04-05 [7] | MIT [7] | no public benchmark yet |
| CLIP ViT-Base-Patch32 [1] | OpenAI | — | — | 2022-03-02 [1] | — | no public benchmark yet |
| CLIP ViT-Large-Patch14 [3] | OpenAI | 428M [3] | 4-bit weights: 214 MB estimated (params x 0.5 bytes); VRAM not published [3] | 2022-03-02 [3] | — | no public benchmark yet |
Which zero-shot image classifiers should you run locally?
1. colipri
colipri by Microsoft ranks first because its Hugging Face publication date, January 14, 2026, puts it ahead on recency.[12][8] Microsoft lists 258M parameters, F32 weights of 1.0 GB, and an MIT license.[12] Consider it for local zero-shot image classification evaluation when you want to assess this release under that license.[12]
For its 258M parameters, 4-bit weight storage is approximately 129 MB, estimated (params x 0.5 bytes).[12] A local machine also needs memory for runtime overhead; that estimate is not a measured RAM or VRAM requirement. Hardware selection therefore needs a runtime memory check. The caveat: no public benchmark yet, so the ranking does not establish an accuracy advantage. Context length is unspecified.
Ranking note: newest release first, then downloads. Too few candidates meet the recency window, so the newest available releases are included.[1][3][4][6][7][8][10][12] Scope admits only zero-shot image classification models from labs with a published paper or leaderboard record, using Hugging Face’s zero-shot-image-classification pipeline tag.
2. PE-Core-L14-336
PE-Core-L14-336 by Facebook ranks second under the ordering rule: newest release first, then downloads. Its first Hugging Face publication was April 11, 2025, behind Microsoft’s colipri on January 14, 2026.[8][12] Placement reflects publication order rather than demonstrated accuracy: no public benchmark yet.
PE-Core-L14-336 carries the Apache-2.0 license and accompanies Facebook’s paper “Perception Encoder,” published April 17, 2025.[8][9] Parameter count, text context length and quantized weight memory remain unspecified, so a concrete GPU or RAM recommendation would require additional verification. For a local deployment, measure peak memory with your intended image inputs and batch settings before choosing hardware.
Use PE-Core-L14-336 as a candidate for local zero-shot image classification when Apache-2.0 licensing suits your project.[8] The practical caveat is performance uncertainty: evaluate your own labels and images before treating its ranking as a reason to deploy it.
3. SigLIP 2 Base Patch16-256
SigLIP 2 Base Patch16-256 by Google ranks third under the “newest release first, then downloads” rule.[4][8][12] Its Hugging Face publication date is February 17, 2025, so it appears as an older available option outside the twelve-month release window.[4] It precedes Google’s SigLIP 2 Giant Opt Patch16-384 because both share that publication date, while Base has more downloads: 3,894,401 versus 2,161,515 in the recorded thirty-day window.[4][6]
The model has 375 million parameters, a listed F32 weight size of 1.5 GB, and an Apache-2.0 license.[4] Four-bit weight storage is approximately 187.5 MB, estimated (params x 0.5 bytes) from that parameter count.[4] Treat this as a weight-storage estimate when planning local hardware; a total RAM or VRAM requirement and a compatible quantized runtime are not established.
Consider it for local zero-shot image classification when Apache-2.0 licensing and its parameter budget suit your project.[4] The caveat is performance uncertainty: no public benchmark yet, so its placement does not establish an accuracy advantage.
4. SigLIP 2 Giant Opt Patch16-384
SigLIP 2 Giant Opt Patch16-384 by Google ranks here under the ordering rule: newest release first, then downloads. Its first Hugging Face publication was February 17, 2025, the same date as Google’s SigLIP 2 Base Patch16-256.[4][6] The base variant wins the tiebreak with 3,894,401 downloads versus 2,161,515 over the preceding 30 days, as of September 25, 2026.[4][6] The giant variant has no public benchmark yet.
The model has 1.9 billion parameters and a listed F32 weight size of 7.5 GB.[6] Four-bit weight storage is approximately 950 MB, estimated (params x 0.5 bytes) from that parameter count.[6] Treat that calculation as a weight-storage budget; a specific GPU or total RAM requirement cannot be established from it.
Consider this Apache-2.0-licensed model for local zero-shot image classification experiments where license terms matter.[6] The practical caveat is that the calculated footprint does not establish a working quantized configuration or measured runtime memory requirement.
5. MobileCLIP-S1-OpenCLIP
MobileCLIP-S1-OpenCLIP by Apple occupies fifth place under the ordering rule: newest release first, then downloads. Its Hugging Face publication date is June 7, 2024, placing it among the older alternatives admitted because too few recent releases are available.[10] Benchmark status: no public benchmark yet, so its position does not establish an accuracy advantage.
Apple lists full-precision weights of 0.3 GB and the apple-amlr license.[10] Local zero-shot image-classification experiments where checkpoint storage matters are a practical use to consider. Plan disk space around that published weight size, but do not treat the download size as a runtime memory requirement. A minimum GPU specification or system RAM target cannot be established from the available specifications.
The caveat is incomplete deployment sizing: parameter count and text context length are unspecified here. Without a parameter count, a defensible quantized-weight estimate is unavailable. Check the license terms and measure memory use with your intended runtime before committing hardware.
6. BiomedCLIP-PubMedBERT_256-vit_base_patch16_224
BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 by Microsoft ranks here under the ordering rule: newest release first, then downloads. Its Hugging Face publication date of April 5, 2023[7] places it after Apple’s MobileCLIP-S1-OpenCLIP, published June 7, 2024,[10] and before OpenAI’s CLIP ViT models, published March 2, 2022.[1][3] The position reflects publication timing, not demonstrated accuracy.
The model carries the MIT license.[7] Consider it for a local evaluation when that license meets your project’s requirements. Parameter count, text context length and weight size are unspecified, so neither a quantized weight-memory estimate nor a concrete GPU or system RAM recommendation is justified. Check those requirements before allocating hardware.
The accuracy caveat is straightforward: no public benchmark yet. No dated benchmark result supports a workload-specific recommendation, so validate performance on your intended images and labels before adopting it.
7. CLIP ViT-Base-Patch32
CLIP ViT-Base-Patch32 by OpenAI is an established pick whose position follows release date, with downloads breaking ties. The ranking uses newest release first, then downloads because no public benchmark scores the candidates against one another. Its Hugging Face publication date is March 2, 2022, shared with OpenAI’s CLIP ViT-Large-Patch14; its higher download count places it ahead of that checkpoint.[1][3]
The text context is 77 tokens, so keep candidate-label prompts within that limit.[1] Use the model as an established baseline for zero-shot image classification when comparing a local workflow with newer checkpoints. OpenAI describes the underlying approach in “Learning Transferable Visual Models From Natural Language Supervision.”[2]
For local hardware planning, verify memory consumption in your intended runtime before choosing a GPU or committing system RAM. A parameter-based quantized-weight estimate cannot be supplied without a verified parameter count. The caveat is comparative accuracy: no public benchmark yet establishes its standing against this candidate set, so its placement should not imply an accuracy advantage.
8. CLIP ViT-Large-Patch14
CLIP ViT-Large-Patch14 by OpenAI ranks eighth as an established pick: its Hugging Face publication date is 2022-03-02, and its 8,651,752 downloads trail the same-date CLIP ViT-Base-Patch32 by OpenAI.[3][1] The ordering is newest release first, then downloads. The established-pick exemption admits older releases; popularity serves only as a tiebreak. Benchmark status: no public benchmark yet.
The model has 428M parameters and a text context of 77 tokens.[3] Published F32 weights occupy 1.7 GB.[3] A 4-bit weight budget is 214 MB, estimated (params x 0.5 bytes) from the cited 428M parameters.[3] Local hardware needs memory beyond that weight budget for execution; the estimate does not establish a complete RAM or GPU requirement.
Use the model as an established CLIP reference for local zero-shot image classification. The practical caveat is prompt length: keep candidate-label descriptions within the 77-token text context.[3] Choose it for continuity with CLIP workflows, without treating its placement as evidence of comparative accuracy.
How much memory do local zero-shot image classifiers need?
Local zero-shot image classifiers can have compact weight storage: Microsoft’s colipri (Microsoft) has 258M parameters, giving a 4-bit weight footprint of 129 MB, estimated (params x 0.5 bytes); its listed F32 weights occupy 1.0 GB.[12] Treat the calculated footprint as a weight-storage estimate, not a measured requirement for running the model.
Google’s google/siglip2-base-patch16-256 has 375M parameters: 187.5 MB at 4-bit, estimated (params x 0.5 bytes), versus listed F32 weights of 1.5 GB.[4] Google’s google/siglip2-giant-opt-patch16-384 has 1.9B parameters: 950 MB at 4-bit, estimated (params x 0.5 bytes), versus listed F32 weights of 7.5 GB.[6]
OpenAI’s openai/clip-vit-large-patch14 has 428M parameters: 214 MB at 4-bit, estimated (params x 0.5 bytes), versus listed F32 weights of 1.7 GB.[3] Apple’s MobileCLIP-S1-OpenCLIP (Apple) lists full-precision weights of 0.3 GB.[10] A full-precision file size alone does not establish a quantized runtime footprint.
For hardware planning, distinguish stored weights from total memory during inference. The arithmetic above does not establish whether a particular implementation supports the assumed quantization or how much additional memory execution needs. Before committing to hardware, verify quantization support and measure peak memory with your intended runtime, images and candidate labels. Avoid treating a calculated weight footprint as a GPU or system-RAM fit guarantee.
Which licenses apply to local zero-shot image classifiers?
The listed local zero-shot image classifiers use MIT, Apache-2.0, and apple-amlr licenses, so license review should follow the exact model repository you plan to deploy.[4][6][7][8][10][12] Treat the license as a separate selection criterion from classification quality or hardware requirements.
Microsoft’s colipri (Microsoft) and Microsoft’s BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 (Microsoft) list the MIT license.[12][7] For either model, include the repository’s license text in your deployment review. Keep the model identifier alongside that record so the review remains tied to the intended artifact.
Google’s google/siglip2-base-patch16-256 and google/siglip2-giant-opt-patch16-384, along with Facebook’s PE-Core-L14-336 (Facebook), list Apache-2.0.[4][6][8] Those entries share a license label; evaluate their technical suitability separately. A shared license does not establish which model suits your classification task.
Apple’s MobileCLIP-S1-OpenCLIP (Apple) lists apple-amlr.[10] Review that named license directly when considering the model for your application. Avoid treating every downloadable model as having interchangeable terms.
For OpenAI’s openai/clip-vit-base-patch32 and openai/clip-vit-large-patch14, leave the license decision pending a repository license check. Do not infer permission from download availability. Before deployment, check the applicable terms against your intended use, modification, and redistribution plans.
How should you choose a zero-shot image classifier without public benchmarks?
Choose by license, estimated weight memory, and performance on your own labeled images; use release recency to prioritize evaluation when public benchmarks are unavailable. Compare candidates using the same images, label descriptions, and acceptance criteria.
Ranking note: this ranking admits only zero-shot image classification models from labs with a published paper or leaderboard record, using Hugging Face’s zero-shot-image-classification pipeline tag. The recent-release pool is too small, so the newest available candidates include older releases.[4][8][10] No public benchmark scores any candidate, so the ordering is “newest release first, then downloads.” Microsoft’s colipri (Microsoft) leads because its publication date, January 14, 2026, is the newest among the listed candidates.[12]
Established picks bypass the recency gate but cannot take the lead through that exemption. Those picks are OpenAI’s openai/clip-vit-base-patch32, OpenAI’s openai/clip-vit-large-patch14, and Google’s google/siglip2-base-patch16-256, based on their download counts.[1][3][4] Downloads break ties; they do not establish classification quality.
Check deployment constraints before evaluating accuracy. Microsoft’s colipri has an MIT license and 258M parameters; its 4-bit weight memory is 129 MB, estimated (params x 0.5 bytes).[12] Google’s SigLIP candidate has an Apache-2.0 license and 375M parameters; its 4-bit weight memory is 187.5 MB, estimated (params x 0.5 bytes).[4] Measure runtime memory separately.
Record benchmark status as “no public benchmark yet.” Keep your evaluation results separate from public scores, and label any model-card results “self-reported.” Check text limits when designing labels: both listed OpenAI models have a context length of 77 tokens.[1][3]
Frequently Asked Questions
How are the local zero-shot image classifiers ranked?
Ranking note: scope admits only zero-shot image classification models from labs with a published paper or leaderboard record, tagged zero-shot-image-classification on Hugging Face. Ordering is newest release first, then downloads: Microsoft colipri (Microsoft) [12]; Meta PE-Core-L14-336 (Facebook) [8]; Google google/siglip2-base-patch16-256 [4]; Google google/siglip2-giant-opt-patch16-384 [6]; Apple MobileCLIP-S1-OpenCLIP (Apple) [10]; Microsoft BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 (Microsoft) [7]; OpenAI openai/clip-vit-base-patch32 [1]; OpenAI openai/clip-vit-large-patch14 [3]. Established picks—SigLIP base and both OpenAI models—bypass the recency gate but cannot take first place through that exemption. [1][3][4]
Which model should I try first?
Microsoft colipri ranks first because its Hugging Face publication date, January 14, 2026, is the newest among the listed candidates. [12] The category has too few releases within the last twelve months, so the newest available candidates are included alongside older established picks. [1][3][4][6][7][8][10][12] Colipri has no public benchmark yet; its position reflects release recency rather than a demonstrated accuracy advantage.
How much memory would quantized weights need?
For colipri's 258M parameters, weight memory is 129 MB estimated (params x 0.5 bytes). [12] SigLIP base has 375M parameters: 187.5 MB estimated (params x 0.5 bytes). [4] SigLIP giant has 1.9B parameters: 950 MB estimated (params x 0.5 bytes). [6] CLIP large has 428M parameters: 214 MB estimated (params x 0.5 bytes). [3] Treat these as weight-storage calculations, not measured runtime memory or confirmation that a compatible quantized implementation exists.
What text context limits should I plan around?
OpenAI clip-vit-base-patch32 and clip-vit-large-patch14 both have a text context length of 77 tokens. [1][3] Context limits for the other candidates remain unverified in this comparison. Check the chosen checkpoint's tokenizer configuration before designing long classification prompts, and avoid inferring a text context limit from numbers embedded in a repository name.
Which licenses apply to these local models?
Microsoft colipri and BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 use the MIT license. [12][7] Google's listed SigLIP variants and Meta's PE-Core-L14-336 use Apache-2.0. [4][6][8] Apple's MobileCLIP-S1-OpenCLIP uses apple-amlr. [10] License details for the OpenAI checkpoints remain unverified in this comparison. Read the applicable license before distributing weights or incorporating a checkpoint into a product.
Does the ranking show which classifier is most accurate?
No public leaderboard scores these candidates against each other, so every entry carries the status “no public benchmark yet.” No comparable benchmark score or benchmark date is available for this ranking. Model-card results, when included, must be labeled self-reported. Downloads only break release-date ties: Google's base and giant SigLIP checkpoints share a February 17, 2025 publication date, and the base checkpoint has more downloads. [4][6]
Sources
- openai/clip-vit-base-patch32 model card (Hugging Face) — 2026-09-25
- Learning Transferable Visual Models From Natural Language Supervision — 2021-02-26
- openai/clip-vit-large-patch14 model card (Hugging Face) — 2026-09-25
- google/siglip2-base-patch16-256 model card (Hugging Face) — 2026-09-25
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features — 2025-02-20
- google/siglip2-giant-opt-patch16-384 model card (Hugging Face) — 2026-09-25
- microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 model card (Hugging Face) — 2026-09-25
- facebook/PE-Core-L14-336 model card (Hugging Face) — 2026-09-25
- Perception Encoder: The best visual embeddings are not at the output of the network — 2025-04-17
- apple/MobileCLIP-S1-OpenCLIP model card (Hugging Face) — 2026-09-25
- MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training — 2023-11-28
- microsoft/colipri model card (Hugging Face) — 2026-09-25