Quick Answer
As of September 2026, ConvNeXt V2 Base (22k, 384) (Facebook) is the top pick for its newest listed release; eligibility covers image classification models from labs with a published paper or leaderboard record, recent releases are too few so the newest available are listed, and established picks Vision Transformer Base (patch16, 224) (Google), ResNet-50 (Microsoft), and Swin Transformer V2 Tiny (patch4, window16, 256) (Microsoft) receive recency exemptions without winning through that exemption [1][3][5][7][8]. In order, the ranking is ConvNeXt V2 Base (22k, 384) (Facebook), MobileNetV2 1.0 224 (Google), Swin Transformer V2 Tiny (patch4, window16, 256) (Microsoft), ResNet-50 (Microsoft), Vision Transformer Base (patch16, 224) (Google), Data-efficient Image Transformer Tiny (patch16, 224) (Facebook), Mix Transformer B2 (NVIDIA), and Mix Transformer B0 (NVIDIA) [1][3].
Key Takeaways
- Ranking note: scope admits only image classification models from labs with a published paper or leaderboard record, using Hugging Face’s image-classification pipeline tag. Too few recent releases qualify, so the newest available candidates are included despite falling outside the requested release window. Ordering is “newest release first, then downloads”; no public benchmark scores any candidate. [1][3][5][7][9][11][13][15]
- Facebook’s ConvNeXt V2 Base (22k, 384) ranks first because its Hugging Face publication date is the newest among the candidates: 2023-02-19. Its 89M parameters imply 44.5 MB of 4-bit weight storage, estimated (params x 0.5 bytes); the license is Apache-2.0. [7]
- Google’s MobileNetV2 1.0 224 offers a compact weight footprint: 4M parameters imply 2 MB at 4-bit, estimated (params x 0.5 bytes). Its ranking follows publication date, not size; a license is not specified in the supplied model details. [13]
- Established picks are Google’s Vision Transformer Base (patch16, 224), Microsoft’s ResNet-50, and Microsoft’s Swin Transformer V2 Tiny (patch4, window16, 256), identified by downloads. Their recency exemption does not grant first place; the same ordering rule applies. [1][3][5]
- ResNet-50’s 26M parameters imply 13 MB of 4-bit weights, estimated (params x 0.5 bytes); Vision Transformer Base’s 87M parameters imply 43.5 MB, estimated (params x 0.5 bytes). Both carry Apache-2.0 licenses. [3][1]
- Every candidate has “no public benchmark yet” for this comparison. Raspberry Pi latency, runtime RAM, context limits, and working quantized deployment are unspecified, so weight estimates alone cannot establish which model runs well on the device. [1][3][5][7][9][11][13][15]
How do image classification models for Raspberry Pi compare on specifications and published benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| ConvNeXt V2 Base (22k, 384) [7] | 89M [7] | 4-bit weights: 44.5 MB, estimated (params x 0.5 bytes); runtime VRAM not published [7] | 2023-02-19 (first published on Hugging Face) [7] | Apache-2.0 [7] | no public benchmark yet | |
| MobileNetV2 1.0 224 [13] | 4M [13] | 4-bit weights: 2 MB, estimated (params x 0.5 bytes); runtime VRAM not published [13] | 2022-11-10 (first published on Hugging Face) [13] | — | no public benchmark yet | |
| Swin Transformer V2 Tiny (patch4, window16, 256) [5] | Microsoft | — | — | 2022-06-14 (first published on Hugging Face) [5] | Apache-2.0 [5] | no public benchmark yet |
| ResNet-50 [3] | Microsoft | 26M [3] | 4-bit weights: 13 MB, estimated (params x 0.5 bytes); runtime VRAM not published [3] | 2022-03-16 (first published on Hugging Face) [3] | Apache-2.0 [3] | no public benchmark yet |
| Vision Transformer Base (patch16, 224) [1] | 87M [1] | 4-bit weights: 43.5 MB, estimated (params x 0.5 bytes); runtime VRAM not published [1] | 2022-03-02 (first published on Hugging Face) [1] | Apache-2.0 [1] | no public benchmark yet | |
| Data-efficient Image Transformer Tiny (patch16, 224) [9] | — | — | 2022-03-02 (first published on Hugging Face) [9] | Apache-2.0 [9] | no public benchmark yet | |
| Mix Transformer B2 [15] | NVIDIA | — | — | 2022-03-02 (first published on Hugging Face) [15] | — | no public benchmark yet |
| Mix Transformer B0 [11] | NVIDIA | — | — | 2022-03-02 (first published on Hugging Face) [11] | — | no public benchmark yet |
Which image classification models should you consider for Raspberry Pi?
1. ConvNeXt V2 Base (22k, 384)
ConvNeXt V2 Base (22k, 384) by Facebook ranks first because its Hugging Face release is the newest among the listed candidates.[7][13][5] The ordering is “newest release first, then downloads”; no public benchmark yet. Too few recent releases qualify, so the newest available candidates are included. The ranking admits only image classification models from labs with a published paper or leaderboard record.
The model has 89M parameters,[7] with 4-bit weight storage of 44.5 MB, estimated (params x 0.5 bytes).[7] Published F32 weights occupy 0.4 GB, and the license is Apache-2.0.[7] Total RAM needed on Raspberry Pi is unverified; the weight estimate does not establish runtime memory requirements or a working quantized deployment.
Consider this model for a local classification evaluation where you can measure memory use and latency on your target Raspberry Pi before committing. The practical caveat is that its ranking reflects release order: Raspberry Pi speed, accuracy and hardware fit remain unverified.
2. MobileNetV2 1.0 224
MobileNetV2 1.0 224 by Google ranks second under the ordering rule: newest release first, then downloads.[7][13] Its Hugging Face publication date falls between Facebook’s ConvNeXt V2 checkpoint and Microsoft’s Swin Transformer V2 checkpoint.[7][13][5] The ranking uses the newest available candidates because too few recent releases qualify; MobileNetV2 has no public benchmark yet.
Google lists 4M parameters, giving 4-bit weight storage of approximately 2 MB, estimated (params x 0.5 bytes).[13] That calculation covers weights alone. A Raspberry Pi deployment still needs memory for the runtime, intermediate tensors and image processing; a supported quantized runtime and total RAM requirement are not established here. Choose hardware only after checking the intended runtime’s compatibility and measuring total memory use.
Consider MobileNetV2 for a local image-classification prototype where keeping parameter storage small matters.[13] The caveat is that its ranking does not establish Raspberry Pi inference speed or classification accuracy. Confirm the checkpoint’s license and runtime support before deployment.
3. Swin Transformer V2 Tiny (patch4, window16, 256)
Swin Transformer V2 Tiny (patch4, window16, 256) by Microsoft ranks third under the ordering rule: newest release first, then downloads.[5][7][13] Its first publication on Hugging Face was June 14, 2022, behind the listed ConvNeXt V2 and MobileNetV2 checkpoints.[5][7][13] The benchmark status is “no public benchmark yet”; the position does not establish Raspberry Pi performance.
The checkpoint uses the Apache-2.0 license and belongs to the architecture family described in Microsoft’s “Swin Transformer V2: Scaling Up Capacity and Resolution” paper.[5][6] Consider it for a local image-classification evaluation where that license suits your project. Parameter count and weight size are unspecified, so a quantized weight-memory estimate cannot be calculated here.
The hardware caveat is unresolved deployment requirements: required Raspberry Pi RAM, compatible runtime and inference latency remain unverified. Before choosing a board, validate checkpoint loading, peak memory consumption and classification speed in your intended runtime. Treat this entry as an evaluation candidate, with hardware suitability still to establish.
4. ResNet-50
ResNet-50 by Microsoft ranks fourth under the ordering rule: newest release first, then downloads.[3][5][7][13] Its Hugging Face publication date is March 16, 2022.[3] ResNet-50 is an established pick; its recency exemption does not establish a performance advantage. Benchmark status: no public benchmark yet.
Microsoft lists 26M parameters, F32 weights occupying 0.1 GB, and an Apache-2.0 license.[3] Weight storage at 4-bit is approximately 13 MB, estimated (params x 0.5 bytes) from the cited parameter count.[3] Treat that calculation as a weight-storage budget, not a total RAM requirement or confirmation that a compatible quantized implementation exists.
An established image-classification baseline is a practical use for ResNet-50, particularly when Apache-2.0 licensing suits the project.[3] For local Raspberry Pi deployment, validate runtime compatibility, peak RAM and latency on the intended board before choosing hardware. The caveat is deployment uncertainty: a supported Raspberry Pi configuration, total memory requirement and measured on-device speed are not specified.
5. Vision Transformer Base (patch16, 224)
Vision Transformer Base (patch16, 224) by Google [1] ranks fifth under newest release first, then downloads. Its Hugging Face publication date is March 2, 2022 [1], behind the preceding checkpoints [3][5][7][13]. An established pick, it recorded 8,229,454 downloads over the reported 30-day period [1]; popularity breaks ties without demonstrating Raspberry Pi performance.
The checkpoint has 87 million parameters, with 43.5 MB of weight storage at 4-bit precision—estimated (params x 0.5 bytes) [1]. Published F32 weights occupy 0.3 GB, and the license is Apache-2.0 [1]. Treat the quantized figure as a storage estimate, not a measured runtime memory requirement.
Use it as an evaluation candidate when Apache-2.0 licensing matters [1]. Hardware needed remains unverified: no Raspberry Pi RAM requirement or latency measurement is established here. Plan to measure total memory and inference time on your target board before deployment. The benchmark caveat is no public benchmark yet for this ranking, so its position does not establish accuracy or speed.
6. Data-efficient Image Transformer Tiny (patch16, 224)
Data-efficient Image Transformer Tiny (patch16, 224) by Facebook occupies sixth place under the ordering rule “newest release first, then downloads.”[9] The checkpoint was first published on Hugging Face on March 2, 2022, and recorded 145,686 downloads over the preceding 30 days as of September 25, 2026.[9] Its position reflects publication date and download tie-breaking, rather than demonstrated Raspberry Pi performance.
The checkpoint carries an Apache-2.0 license.[9] Benchmark status: no public benchmark yet. Treat the model as a candidate for local image-classification experiments where that license suits your project. A quantized-weight memory estimate is unavailable for this entry, so choosing a Raspberry Pi memory configuration requires deployment measurements.
Before committing hardware, verify that your inference runtime supports the checkpoint, then measure peak RAM use and classification latency on your intended board. The practical caveat is that the “Tiny” name alone does not establish memory fit or acceptable speed on Raspberry Pi.
7. Mix Transformer B2
Mix Transformer B2 by NVIDIA ranks seventh under the ordering rule: newest release first, then downloads.[15][11] Its Hugging Face publication date is March 2, 2022, and its 98,086 downloads over the recorded 30-day period place it ahead of NVIDIA’s Mix Transformer B0 among equally dated candidates.[15][11] The model has no public benchmark yet.
Mix Transformer B2 is associated with NVIDIA’s “SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers” paper.[12] Parameter count, weight size, input dimensions and license terms remain unconfirmed for this comparison. Treat it as an evaluation candidate for an image classification workflow, with accuracy and deployment compatibility still to establish.
A specific Raspberry Pi RAM configuration cannot be recommended without verified weight and runtime memory requirements. Before committing hardware, check runtime compatibility, measure peak memory and latency, and verify licensing. The practical caveat is that its ranking does not establish Raspberry Pi suitability or an accuracy advantage.
8. Mix Transformer B0
Mix Transformer B0 by NVIDIA ranks eighth under the ordering rule: newest release first, then downloads.[11][15] Its Hugging Face publication date is March 2, 2022, with 95,766 downloads over the reporting window ending September 25, 2026.[11] NVIDIA’s Mix Transformer B2 shares that publication date but records 98,086 downloads, placing it ahead on the popularity tiebreak.[15] Mix Transformer B0 has no public benchmark yet.
Use the model as an evaluation candidate for local image classification. Parameter count, quantized weight memory, license terms and Raspberry Pi hardware requirements remain unverified for this entry. Before deployment, confirm the license and runtime compatibility, then measure peak RAM use and inference latency on your intended board. A hardware purchase recommendation would require those measurements.
The associated SegFormer paper addresses semantic segmentation.[12] Verify that the checkpoint’s classification head, preprocessing and output labels match your application before integrating it into a local image classification pipeline.
How can you estimate classification model weight memory from parameter counts?
Estimate classification model weight memory by multiplying the parameter count by storage per parameter: for 4-bit weights, use estimated (params x 0.5 bytes) [13]. Treat the result as a weight-storage budget, rather than a measurement of memory used during inference.
Google’s MobileNetV2 (google/mobilenet_v2_1.0_224) has 4 million parameters [13], giving approximately 2 MB, estimated (params x 0.5 bytes) [13]. Microsoft’s ResNet-50 (ResNet-50 (Microsoft)) has 26 million parameters [3], giving approximately 13 MB, estimated (params x 0.5 bytes) [3].
Google’s Vision Transformer Base (google/vit-base-patch16-224) has 87 million parameters [1], giving approximately 43.5 MB, estimated (params x 0.5 bytes) [1]. Facebook’s ConvNeXt V2 Base (facebook/convnextv2-base-22k-384) has 89 million parameters [7], giving approximately 44.5 MB, estimated (params x 0.5 bytes) [7]. Each calculation assumes the stated precision applies across the parameter count.
For Raspberry Pi planning, keep weight storage separate from the full application memory budget. Account separately for quantization metadata, intermediate activations, input buffers and runtime overhead. Confirm the checkpoint format and execution support before treating a calculated weight budget as a deployment option.
Use parameter-based estimates to screen candidates, then measure memory use with the intended runtime and workload. Avoid turning a weight-only estimate into a claim that a model fits or runs well on a particular Raspberry Pi.
Which image classification models have a documented license?
Google’s Vision Transformer Base (google/vit-base-patch16-224) [1], Microsoft’s ResNet-50 (ResNet-50 (Microsoft)) [3], Microsoft’s Swin Transformer V2 Tiny (microsoft/swinv2-tiny-patch4-window16-256) [5], Facebook’s ConvNeXt V2 Base (facebook/convnextv2-base-22k-384) [7], and Facebook’s Data-efficient Image Transformer Tiny (facebook/deit-tiny-patch16-224) [9] all have a documented Apache-2.0 license.
For a Raspberry Pi deployment, those models provide a starting point when an explicit license is a selection requirement. Treat licensing as a separate check from runtime suitability: a documented license does not establish inference speed, memory use, or compatibility with your chosen inference engine. Keep the exact repository identifier in your deployment records so the license check refers to the checkpoint you intend to distribute or run.
License status remains unconfirmed here for Google’s MobileNetV2 (google/mobilenet_v2_1.0_224) [13], NVIDIA’s Mix Transformer B2 (nvidia/mit-b2) [15], and NVIDIA’s Mix Transformer B0 (nvidia/mit-b0) [11]. An unconfirmed license should not be read as a finding that a model prohibits commercial use or has no license. Check the repository’s license file and checkpoint terms before making that decision.
For deployment review, record the checkpoint revision, retain its applicable license and notices, and review any separately licensed runtime or conversion tools. Choose hardware and quantization settings through a separate compatibility check.
What benchmark information is missing when choosing a classifier for Raspberry Pi?
Raspberry Pi latency, throughput, peak runtime memory, and classification accuracy after quantization are missing from this comparison, leaving actual device performance unresolved. Every candidate has “no public benchmark yet” here; no shared public leaderboard scores these candidates against each other. Publication dates and download counts cannot fill that gap.[1][3][5][7][9][11][13][15]
A useful device benchmark would specify the Raspberry Pi hardware, operating system, inference runtime, thread settings, cooling, and image preprocessing. Request both single-image latency and sustained throughput, with the batch size stated. Include startup time separately so loading a model does not get confused with processing an image.
Weight storage also needs separating from runtime memory. Google’s MobileNetV2 (google/mobilenet_v2_1.0_224) has 4M parameters; its weight-only footprint at 4-bit precision is approximately 2 MB, estimated (params x 0.5 bytes).[13] That calculation does not establish Raspberry Pi compatibility or total memory use. A deployment benchmark should measure peak memory and identify the actual weight format and execution backend.
Accuracy testing should name the dataset, evaluation split, preprocessing, and metric, then compare the original and quantized checkpoints under matching conditions. Any model-card result should be labeled self-reported. Until comparable device measurements are available, treat the candidates as options to evaluate rather than a demonstrated Raspberry Pi performance ranking.
Frequently Asked Questions
Which image classification model ranks first for Raspberry Pi?
Facebook’s ConvNeXt V2 (facebook/convnextv2-base-22k-384) ranks first because its Hugging Face publication date, February 19, 2023, is the newest among the listed candidates [7][13][5][3][1][9][15][11]. The checkpoint has 89M parameters and an Apache-2.0 license [7]. Its benchmark status is “no public benchmark yet.” Treat its position as a recency-based selection, rather than a demonstrated Raspberry Pi performance result.
How is the ranking ordered?
Ranking note: this ranking admits only image classification models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: image-classification). No public benchmark scores any candidate, so the rule gives: newest release first, then downloads [7][13][5][3][1][9][15][11]. Too few candidates meet the requested twelve-month release window, so the newest available entries are listed [7][13][5][3][1][9][15][11]. Downloads only break ties; model size does not determine placement.
How much memory would quantized model weights need?
Google’s MobileNetV2 (google/mobilenet_v2_1.0_224) has 4M parameters: 2 MB of weight storage, estimated (params x 0.5 bytes) [13]. Facebook’s ConvNeXt V2 has 89M parameters: 44.5 MB, estimated (params x 0.5 bytes) [7]. Both calculations describe hypothetical quantized weight storage. Neither establishes a working quantization path or a Raspberry Pi RAM requirement. Measure total memory with your chosen runtime before selecting hardware.
Which checkpoints have a documented permissive license?
Apache-2.0 is documented for Microsoft’s ResNet (ResNet-50 (Microsoft)) [3], Microsoft’s Swin Transformer V2 (microsoft/swinv2-tiny-patch4-window16-256) [5], and Google’s Vision Transformer (google/vit-base-patch16-224) [1]. Facebook’s ConvNeXt V2 [7] and Facebook’s Data-efficient Image Transformer (facebook/deit-tiny-patch16-224) also carry Apache-2.0 [9]. Check the exact checkpoint’s license when packaging an application, and review the license terms before distributing model weights with your Raspberry Pi software.
Which models count as established picks?
The established picks are Google’s Vision Transformer, Microsoft’s ResNet, and Microsoft’s Swin Transformer V2: the three most-downloaded candidates [1][3][5]. Their recorded download totals are 8,229,454 [1], 589,550 [3], and 402,958 [5], respectively, for the preceding thirty days as of September 25, 2026 [1][3][5]. Established picks bypass the recency gate but retain the same ordering rules. The exemption cannot give them first place, and popularity does not establish Raspberry Pi performance.
What should I check before deploying a classifier on Raspberry Pi?
Check image preprocessing, runtime compatibility, latency and peak RAM with your intended workload. No context-window value is specified for this comparison; verify the checkpoint’s image input requirements instead. NVIDIA’s Mix Transformer checkpoints nvidia/mit-b2 [15] and nvidia/mit-b0 [11], like the other candidates, have “no public benchmark yet” in this ranking. Require a reproducible Raspberry Pi measurement before treating any placement as evidence of deployment performance.
Sources
- google/vit-base-patch16-224 model card (Hugging Face) — 2026-09-25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — 2020-10-22
- microsoft/resnet-50 model card (Hugging Face) — 2026-09-25
- Deep Residual Learning for Image Recognition — 2015-12-10
- microsoft/swinv2-tiny-patch4-window16-256 model card (Hugging Face) — 2026-09-25
- Swin Transformer V2: Scaling Up Capacity and Resolution — 2021-11-18
- facebook/convnextv2-base-22k-384 model card (Hugging Face) — 2026-09-25
- ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders — 2023-01-02
- facebook/deit-tiny-patch16-224 model card (Hugging Face) — 2026-09-25
- Training data-efficient image transformers & distillation through attention — 2020-12-23
- nvidia/mit-b0 model card (Hugging Face) — 2026-09-25
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers — 2021-05-31
- google/mobilenet_v2_1.0_224 model card (Hugging Face) — 2026-09-25
- MobileNetV2: Inverted Residuals and Linear Bottlenecks — 2018-01-13
- nvidia/mit-b2 model card (Hugging Face) — 2026-09-25