Quick Answer
Google TIPSv2-so400m14 from Google is the top pick as of September 2026 because it shares the newest release date and wins the downloads tiebreak, with no public benchmark yet. [1][3][4]
Ranking note: newest release first, then downloads; eligibility is limited to zero-shot image classification models from labs with a published paper or leaderboard record: 1. In order, the ranking is Google TIPSv2-so400m14 (Google), Google TIPSv2-b14 (Google), and Microsoft CoLIPri (Microsoft) [1][3].
Key Takeaways
- Ranking note: This ranking admits only zero-shot image classification models from labs with a published paper or leaderboard record (Hugging Face pipeline tag:
zero-shot-image-classification). No public benchmark scores any candidate, so the ordering is newest release first, then downloads; popularity serves only as a tiebreak. [1][2][3][4] - Google TIPSv2-so400m14 (Google) ranks first because it shares the newest release date, 2026-04-09, with Google TIPSv2-b14 (Google), but has 272,428 downloads versus 210,563 over the last 30 days. [3][1] Weight memory at 4-bit: 431 MB, estimated (params x 0.5 bytes) from 862M parameters; text context: 64 tokens; license: Apache-2.0. [3] Benchmark status: no public benchmark yet. [3]
- Google TIPSv2-b14 (Google) ranks second under the release-date and downloads rule. [1][3] Weight memory at 4-bit: 98 MB, estimated (params x 0.5 bytes) from 196M parameters; text context: 64 tokens; license: Apache-2.0. [1] Benchmark status: no public benchmark yet. [1]
- Microsoft CoLIPri (Microsoft) ranks third because its publication date, 2026-01-14, precedes the Google releases. [4][1][3] Weight memory at 4-bit: 129 MB, estimated (params x 0.5 bytes) from 258M parameters; license: MIT. [4] Benchmark status: no public benchmark yet. [4]
How do these CLIP-style models compare on memory, context, licenses and published benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Google TIPSv2-so400m14 [3] | Google [3] | 862M [3] | 4-bit weights: 431 MB estimated (params x 0.5 bytes) [3]; runtime VRAM: not published | 2026-04-09 [3] | Apache-2.0 [3] | no public benchmark yet [3] |
| Google TIPSv2-b14 [1] | Google [1] | 196M [1] | 4-bit weights: 98 MB estimated (params x 0.5 bytes) [1]; runtime VRAM: not published | 2026-04-09 [1] | Apache-2.0 [1] | no public benchmark yet [1] |
| Microsoft CoLIPri [4] | Microsoft [4] | 258M [4] | 4-bit weights: 129 MB estimated (params x 0.5 bytes) [4]; runtime VRAM: not published | 2026-01-14 [4] | MIT [4] | no public benchmark yet [4] |
Which CLIP-style models should you consider for an Apple Silicon Mac?
1. Google TIPSv2-so400m14
Google TIPSv2-so400m14 by Google ranks first because it shares the newest publication date among the eligible candidates and wins the downloads tiebreak.[1][3][4] Published on Hugging Face on April 9, 2026, the model has 862M parameters, a 64-token text context, and an Apache-2.0 license.[3] Its published F32 weights occupy 3.4 GB.[3] Google describes the family in “TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment.”[2]
Use it as a candidate for local zero-shot image classification when its text context and license suit your application.[3] For Apple Silicon hardware planning, a hypothetical 4-bit conversion would occupy 431 MB for weights alone, estimated (params x 0.5 bytes) from the published 862M parameters.[3] That calculation does not establish the total RAM needed to run it locally. A specific Mac recommendation would require a verified runtime and memory measurement; the weight estimate alone cannot establish compatibility or performance.
The ranking admits only zero-shot image classification models from labs with a published paper or leaderboard record, using the Hugging Face pipeline tag zero-shot-image-classification.[1][2][3][4] The ordering basis is newest release first, then downloads; downloads serve only as a tiebreak.[1][3][4] The caveat is no public benchmark yet: the placement does not establish superior classification accuracy or measured Apple Silicon speed.[3]
2. Google TIPSv2-b14
Google TIPSv2-b14 by Google ranks second because the ordering is newest release first, then downloads: its publication date matches Google TIPSv2-so400m14 by Google, but its download count is lower.[1][3] Both appeared on Hugging Face on April 9, 2026; their respective download counts were 210,563 and 272,428 for the preceding 30 days as of September 25, 2026.[1][3] Google TIPSv2-b14 has no public benchmark yet, so this position does not establish a classification-accuracy advantage.
Google TIPSv2-b14 has 196M parameters, a 64-token text context, and 0.8 GB of published F32 weights under the Apache-2.0 license.[1] Its accompanying paper is “TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment.”[2] A practical use is zero-shot image classification with short candidate labels or prompts that stay within its text context.[1] The ranking admits only zero-shot image classification models from labs with a published paper or leaderboard record.
For Apple Silicon hardware planning, its 196M parameters imply 98 MB of weight storage at 4-bit precision, estimated (params x 0.5 bytes).[1] Treat that calculation as a weight-storage budget, not a complete application memory requirement. The caveat is deployment uncertainty: a minimum Mac memory configuration, working Apple Silicon quantization path, and measured local performance are not established here. A specific hardware recommendation would therefore go beyond what can be substantiated.
3. Microsoft CoLIPri
Microsoft CoLIPri by Microsoft ranks third because its January release precedes the April releases of Google’s TIPSv2 models.[4][1][3] The ordering basis is newest release first, then downloads; CoLIPri has no public benchmark yet, so its position does not establish a classification-quality gap. Eligibility is limited to zero-shot image classification models from labs with a published paper or leaderboard record, using Hugging Face’s zero-shot-image-classification pipeline tag.
Microsoft CoLIPri contains 258M parameters, ships with 1.0 GB of F32 weights, and uses the MIT license.[4] Its weight-only memory at 4-bit is approximately 129 MB, estimated (params x 0.5 bytes) from the cited 258M parameters.[4] Treat that calculation as a storage estimate, not evidence of an available quantized checkpoint. No context-length specification is available here, so prompt-length planning remains unresolved.
For an Apple Silicon Mac, the practical use case is evaluating CoLIPri for local zero-shot image classification when its MIT license suits your project.[4] The hardware caveat is that neither the published weight size nor the quantized estimate establishes total runtime memory or Mac compatibility. A specific Mac configuration cannot be recommended from those figures alone; confirm a compatible implementation and its memory requirements before choosing hardware.
How can you estimate quantized weight memory from parameter counts?
Estimate quantized weight memory by multiplying the parameter count by the storage per parameter: for a 4-bit weight payload, use estimated (params x 0.5 bytes).[1][3][4] Treat the result as a weight-storage estimate, rather than a measurement of memory consumed while classifying images.
Google’s google/tipsv2-so400m14 has 862M parameters: 431 MB for quantized weights, estimated (params x 0.5 bytes) at 4-bit.[3] The published F32 weights occupy 3.4 GB; that figure describes the supplied weights, rather than a measured quantized allocation.[3]
Google’s google/tipsv2-b14 has 196M parameters: 98 MB for quantized weights, estimated (params x 0.5 bytes) at 4-bit.[1] The published F32 weights occupy 0.8 GB.[1] Keep the published weight size and the calculated payload separate when comparing storage requirements.
Microsoft’s microsoft/colipri has 258M parameters: 129 MB for quantized weights, estimated (params x 0.5 bytes) at 4-bit.[4] The published F32 weights occupy 1.0 GB.[4] Apply the same calculation consistently across candidates.
Use those estimates to budget the weight payload before evaluating a local runtime. A parameter-based calculation does not establish total application memory, confirm that a compatible quantized implementation exists, or demonstrate that a model fits a particular Apple Silicon Mac. Keep hardware-fit and performance conclusions separate until the intended implementation provides measurements.
What context lengths are listed for these models?
Google’s TIPSv2 models, google/tipsv2-so400m14 and google/tipsv2-b14, both list a context length of 64 tokens.[3][1] Microsoft’s COLIPRI, microsoft/colipri, has no context length listed in the provided specifications.[4]
The shared context limit means context length does not distinguish the Google models. The google/tipsv2-so400m14 model has 862 million parameters, while google/tipsv2-b14 has 196 million parameters.[3][1] Choosing the larger model therefore does not buy a longer listed text context: the stated limit remains 64 tokens for either option.[3][1] For a local classification workflow, treat that shared limit as a constraint when preparing text prompts.
The Google models accompany the paper “TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment.”[2] Their matching context specifications give you a common text budget when comparing them, although that match alone does not establish equivalent classification quality.
For Microsoft’s COLIPRI, leave context length marked “not listed” rather than assigning the Google models’ limit to it.[4] An omitted specification should remain an unknown when choosing a model for prompts that require a verified context budget.
Which licenses cover these models?
Google’s google/tipsv2-so400m14 and google/tipsv2-b14 use the Apache-2.0 license,[1][3] while Microsoft’s microsoft/colipri uses the MIT license.[4]
The Google models share a license despite their different parameter counts: google/tipsv2-b14 has 196M parameters,[1] and google/tipsv2-so400m14 has 862M parameters.[3] Choosing between those variants therefore does not change the listed license.[1][3] Keep the licensing decision separate from the memory and runtime decisions when evaluating them for a local application.
Microsoft’s microsoft/colipri provides the MIT-licensed option in this selection.[4] Treat that license as a separate review item if your application supports switching between Microsoft’s model and either Google model. Record the exact repository identifier alongside the license so a later model replacement does not silently inherit the previous model’s documentation.
For deployment, review the applicable license text against your intended use, including any redistribution of weights. Keep a copy with your project’s dependency records. Avoid treating a license label as confirmation of hardware compatibility, quantization support or classification quality; evaluate those questions separately before choosing a model for an Apple Silicon Mac.
Frequently Asked Questions
Which CLIP-style model ranks first for Apple Silicon Macs?
Google’s google/tipsv2-so400m14 ranks first because it ties Google’s google/tipsv2-b14 on the newest release date and wins the downloads tiebreak.[3][1] Microsoft’s microsoft/colipri ranks third.[4] Ranking note: newest release first, then downloads; no public benchmark scores any candidate.[1][3][4] Scope: this ranking admits only zero shot image classification models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: zero-shot-image-classification).[1][2][3][4]
How much memory would quantized weights need?
For google/tipsv2-so400m14, 862M parameters imply 431 MB of 4-bit weights, estimated (params x 0.5 bytes).[3] For google/tipsv2-b14, 196M parameters imply 98 MB, estimated (params x 0.5 bytes).[1] For microsoft/colipri, 258M parameters imply 129 MB, estimated (params x 0.5 bytes).[4] Treat these as weight-storage calculations, not measured application memory requirements or confirmation that a compatible quantized implementation exists.
How long can my classification prompts be?
Both google/tipsv2-so400m14 and google/tipsv2-b14 have a context length of 64 tokens.[3][1] Keep candidate-label prompts within that limit and check tokenization before using longer descriptions. For microsoft/colipri, confirm the text-input limit before choosing a prompt format; do not assume that its interface matches either Google model.
What licenses do these models use?
Google’s google/tipsv2-so400m14 and google/tipsv2-b14 use Apache-2.0, while Microsoft’s microsoft/colipri uses MIT.[3][1][4] If your project requires a particular license, use that requirement to narrow the shortlist before evaluating classification behavior. Review the applicable license text before redistribution, and keep licensing decisions separate from assumptions about accuracy or Mac compatibility.
Which model has the strongest published classification score?
No candidate has a public benchmark score that establishes a winner: google/tipsv2-so400m14 has no public benchmark yet; google/tipsv2-b14 has no public benchmark yet; microsoft/colipri has no public benchmark yet.[3][1][4] The ordering therefore reflects release recency and a downloads tiebreak. Treat the ranking as a shortlist for evaluation, not a demonstrated accuracy comparison.
Do the listed weight sizes prove a model will run well on my Mac?
No. The listed F32 weight sizes are 3.4 GB for google/tipsv2-so400m14, 0.8 GB for google/tipsv2-b14, and 1.0 GB for microsoft/colipri.[3][1][4] Those figures describe weights; they do not establish Apple Silicon throughput or application memory use. Before committing to a model, verify compatibility with your intended runtime and measure its behavior on your Mac.
Sources
- google/tipsv2-b14 model card (Hugging Face) — 2026-09-25
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment — 2026-04-13
- google/tipsv2-so400m14 model card (Hugging Face) — 2026-09-25
- microsoft/colipri model card (Hugging Face) — 2026-09-25