Quick Answer
RMBG-1.4 (briaai) is the top pick as of September 2026 because it is the newest listed release, although the candidates fall outside the requested release window and have no public benchmark yet [1][3][5][6][7][9][10][11]. In order, the ranking is RMBG-1.4 (briaai), Mask2Former Swin Large Cityscapes Semantic (facebook), Mask2Former Swin Large ADE Semantic (facebook), CLIPSeg RD64 Refined (CIDAS), SegFormer B0 Fine-Tuned ADE 512-512 (nvidia), SegFormer B2 Fine-Tuned ADE 512-512 (nvidia), BEiT Large Fine-Tuned ADE 640-640 (microsoft), and BEiT Base Fine-Tuned ADE 640-640 (microsoft) [1][3].
Key Takeaways
- RMBG-1.4 from briaai ranks first because its Hugging Face publication date, December 12, 2023, is the newest among the admitted candidates; the position does not establish superior background-removal quality. [10][1][6][11][3][5][7][9]
- Ranking note: no public benchmark scores any candidate, so the ordering is newest release first, then downloads; every entry has no public benchmark yet. The shortlist uses the newest available entries because recent releases are too few; all listed publication dates fall outside the last 12 months. [10][1][6][11][3][5][7][9]
- Eligibility covers only Hugging Face image-segmentation models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days. [2][4][8][12][10]
- Established picks are CIDAS’s CLIPSeg RD64 Refined, facebook’s Mask2Former Swin Large Cityscapes Semantic, and NVIDIA’s SegFormer B0 Fine-Tuned ADE 512-512—the shortlist’s three most-downloaded models; their recency exemption does not grant first place, and downloads only break release-date ties. [11][1][3]
- Weight storage is not a hardware requirement: RMBG-1.4 lists 0.2 GB of F32 weights, CLIPSeg RD64 Refined lists 0.6 GB, and Mask2Former Swin Large Cityscapes Semantic lists 0.9 GB; those figures alone do not establish runtime RAM or GPU memory needs. [10][11][1]
- License terms differ: RMBG-1.4 uses bria-rmbg-1.4, while CLIPSeg RD64 Refined and Microsoft’s BEiT Large Fine-Tuned ADE 640-640 and BEiT Base Fine-Tuned ADE 640-640 list Apache-2.0. [10][11][7][9]
How do local segmentation models compare on specifications and public benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| RMBG-1.4 [10] | briaai [10] | 44M [10] | F32 weights: 0.2 GB [10]; VRAM: not published | 2023-12-12 [10] | bria-rmbg-1.4 [10] | no public benchmark yet |
| Mask2Former Swin Large Cityscapes Semantic [1] | facebook [1] | 216M [1] | F32 weights: 0.9 GB [1]; VRAM: not published | 2023-01-05 [1] | — | no public benchmark yet |
| Mask2Former Swin Large ADE Semantic [6] | facebook [6] | 216M [6] | F32 weights: 0.9 GB [6]; VRAM: not published | 2023-01-05 [6] | — | no public benchmark yet |
| CLIPSeg RD64 Refined [11] | CIDAS [11] | 151M [11] | F32 weights: 0.6 GB [11]; VRAM: not published | 2022-11-01 [11] | Apache-2.0 [11] | no public benchmark yet |
| SegFormer B0 Fine-Tuned ADE 512-512 [3] | nvidia [3] | 4M [3] | F32 weights: 0.0 GB [3]; VRAM: not published | 2022-03-02 [3] | — | no public benchmark yet |
| SegFormer B2 Fine-Tuned ADE 512-512 [5] | nvidia [5] | — | — | 2022-03-02 [5] | — | no public benchmark yet |
| BEiT Large Fine-Tuned ADE 640-640 [7] | microsoft [7] | 503M [7] | F32 weights: 2.3 GB [7]; VRAM: not published | 2022-03-02 [7] | Apache-2.0 [7] | no public benchmark yet |
| BEiT Base Fine-Tuned ADE 640-640 [9] | microsoft [9] | — | — | 2022-03-02 [9] | Apache-2.0 [9] | no public benchmark yet |
Which segmentation models should you consider for local background removal?
1. RMBG-1.4
RMBG-1.4 by briaai ranks first because its Hugging Face publication date, December 12, 2023, is the newest among the listed candidates.[10][1][3][5][6][7][9][11] The ordering is newest release first, then downloads; no public benchmark yet establishes its background-removal quality against these alternatives. The selection uses the newest available candidates because too few releases meet the recency gate. RMBG-1.4 has 44 million parameters and 0.2 GB of F32 weights, with the bria-rmbg-1.4 license.[10]
For local hardware planning, treat the 0.2 GB weight footprint as a starting point, not a complete RAM or VRAM requirement.[10] A specific GPU recommendation or minimum system-memory figure remains unverified. Consider the model for background-removal trials on your own images before committing it to a workflow. The practical caveat is licensing: review the bria-rmbg-1.4 terms against your intended use before deployment.[10]
2. Mask2Former Swin Large Cityscapes Semantic
Mask2Former Swin Large Cityscapes Semantic by facebook ranks second under the newest-release-first, then-downloads ordering.[1][6][10] Its Hugging Face publication date is January 5, 2023,[1] behind the first-place release.[10] Its 483,243 downloads over the preceding 30 days, measured September 25, 2026,[1] break the publication-date tie with facebook’s Mask2Former Swin Large ADE Semantic.[6] Background-removal performance has no public benchmark yet, so this position does not establish cutout quality.
The model has 216 million parameters and 0.9 GB of F32 weights.[1] Those weights provide a starting point for local hardware planning, but not a complete runtime memory requirement. A validated GPU configuration or RAM requirement is not specified; confirm memory use with your intended inference setup before committing hardware.
Consider this semantic-segmentation model for workflows where selecting scene regions is part of background removal.[1][2] The practical caveat is unverified background-removal quality: evaluate foreground boundaries on representative images before adopting it for finished cutouts.
3. Mask2Former Swin Large ADE Semantic
Mask2Former Swin Large ADE Semantic by facebook ranks third under the ordering rule “newest release first, then downloads.”[1][6][10] Its Hugging Face publication date is January 5, 2023, with 279,927 downloads over the last 30 days as of September 25, 2026.[6] The newer background-removal entry takes precedence, while its same-date Cityscapes sibling wins the download tiebreak.[1][10] Background-removal quality does not determine this position: no public benchmark yet compares the ranked candidates.
The model has 216 million parameters and listed F32 weights of 0.9 GB.[6] Local hardware requirements remain unspecified: the weight listing does not establish a GPU recommendation or total RAM or VRAM requirement. Plan to measure memory use with your intended image inputs before committing hardware.
Consider this checkpoint for a semantic-segmentation workflow where you can evaluate its masks for background removal.[6] The practical caveat is licensing: a license is not specified in the available model details, so verify the terms before deployment.
4. CLIPSeg RD64 Refined
CLIPSeg RD64 Refined by CIDAS ranks fourth under the ordering rule: newest release first, then downloads. Its Hugging Face publication date is November 1, 2022.[11] Background-removal performance does not determine this position: no public benchmark yet. Treat the placement as a way to order candidates, rather than evidence of better mask quality.
The model has 151 million parameters, a context length of 77 tokens, and F32 weights listed at 0.6 GB.[11] For local hardware planning, the weight size is a starting point, not a verified memory requirement. A specific GPU, minimum VRAM allocation, or system RAM requirement cannot be established from those specifications alone.
Use CLIPSeg for segmentation workflows guided by text or image prompts, the task described in its accompanying paper.[12] The checkpoint carries an Apache-2.0 license.[11] The practical caveat for background removal is the lack of a public task benchmark: evaluate its masks on your intended images before adoption.
5. SegFormer B0 Fine-Tuned ADE 512-512
SegFormer B0 Fine-Tuned ADE 512-512 by nvidia was first published on Hugging Face on March 2, 2022.[3] Its placement follows newest release first, then downloads. Among candidates sharing its publication date, its 368,092 downloads over the preceding 30 days, measured on September 25, 2026, put it ahead of the other listed SegFormer and BEiT checkpoints.[3][5][7][9] Popularity settles that tie; it does not establish background-removal quality.
The model has 4M parameters and F32 weights.[3] Treat local hardware sizing as unresolved: no validated GPU configuration or runtime RAM requirement is specified. The parameter count alone does not justify promising that a particular machine can run it.
Consider it for evaluating semantic segmentation within a local background-removal workflow.[4] The caveat is task performance: no public benchmark yet establishes its background-removal standing in this comparison, so its position should not be read as an accuracy recommendation.
6. SegFormer B2 Fine-Tuned ADE 512-512
SegFormer B2 Fine-Tuned ADE 512-512 by nvidia ranks sixth under the ordering rule: newest release first, then downloads.[5] Its first Hugging Face publication was March 2, 2022, matching the B0 checkpoint’s date; its 291,568 downloads over the last 30 days, as of September 25, 2026, place it behind B0’s 368,092.[3][5] Download counts determine that tie, not background-removal quality.
The checkpoint belongs to the SegFormer family for semantic segmentation.[4][5] Parameter count, weight size, and verified local GPU or RAM requirements are unspecified, leaving hardware sizing unresolved. Before committing to a local deployment, measure memory use and latency on representative images with your intended hardware.
Treat this checkpoint as a semantic-segmentation candidate to evaluate for background removal. The caveat is no public benchmark yet for this comparison: its position does not establish cleaner foreground boundaries or better background removal than another candidate.
7. BEiT Large Fine-Tuned ADE 640-640
BEiT Large Fine-Tuned ADE 640-640 by microsoft ranks seventh under the ordering rule: newest release first, then downloads.[7][9] First published on Hugging Face on March 2, 2022, the checkpoint has 29,850 downloads over the reporting window ending September 25, 2026, placing it ahead of microsoft’s BEiT Base Fine-Tuned ADE 640-640 among releases sharing that date.[7][9] Background-removal performance remains unranked: no public benchmark yet.
The model has 503 million parameters, with listed F32 weights of 2.3 GB, and uses the Apache-2.0 license.[7] Those weights provide a starting point for local deployment planning, but do not establish a GPU VRAM or system RAM requirement. Confirm runtime memory use on your target hardware before committing to deployment.
Use the checkpoint as an evaluation candidate when its ADE fine-tuning matches your segmentation task.[7] For background removal, the caveat is practical: its position does not establish cutout quality. Evaluate the resulting masks on your intended images before selecting it.
8. BEiT Base Fine-Tuned ADE 640-640
BEiT Base Fine-Tuned ADE 640-640 by microsoft ranks here under the ordering rule: newest release first, then downloads. Its Hugging Face publication date is March 2, 2022, and it recorded 28,849 downloads over the last 30 days as of September 25, 2026.[9] The same-date large variant has more downloads, placing it ahead.[7][9] Background-removal quality does not determine this position: no public benchmark yet.
The checkpoint carries the Apache-2.0 license.[9] A verified parameter count, weight footprint and local RAM or GPU-memory requirement are unavailable, so a specific hardware recommendation would be speculative. Before choosing hardware, measure memory consumption with your intended image dimensions and inference configuration.
Consider this checkpoint for a local segmentation workflow where you can evaluate the resulting masks against your own background-removal requirements. The practical caveat is the absence of a public background-removal score: validate foreground boundaries and retained detail before adopting it.
How are segmentation models ranked when no public benchmark compares them?
Segmentation models are ranked by newest release first, then downloads when no public benchmark scores any candidate against the others. Each candidate is marked “no public benchmark yet”; the ordering does not establish background-removal quality.
The intended eligibility window is the last 12 months, but this pool has too few qualifying releases, so the newest available candidates are listed instead; even briaai’s RMBG-1.4 was first published on Hugging Face on December 12, 2023.[10] RMBG-1.4 takes first place because its publication date is newer than every other admitted candidate’s.[1][3][5][6][7][9][10][11]
The ranking admits only image segmentation models tagged image-segmentation on Hugging Face from labs with a published paper or leaderboard record, or other publishers exceeding 10,000 downloads in the last 30 days.[1][3][5][6][7][9][10][11] Publication dates determine the order; downloads break ties between candidates published on the same date.[1][3][5][6][7][9] Parameter count does not determine rank.
The established picks—the pool’s three most-downloaded models—are CIDAS’s CIDAS/clipseg-rd64-refined, Facebook’s facebook/mask2former-swin-large-cityscapes-semantic, and NVIDIA’s nvidia/segformer-b0-finetuned-ade-512-512.[11][1][3] Their exemption bypasses the recency gate, but the same ordering applies, and the exemption cannot award first place.
A public leaderboard can compare only models it actually evaluates. An older scored model cannot displace a newer unscored candidate under this policy. Model-card benchmark results must be labeled “self-reported.” Hardware requirements and licenses remain separate selection checks; neither weight-file size nor download count demonstrates cutout accuracy.
What hardware requirements can you estimate from segmentation model weights?
Segmentation model weights provide a starting estimate for weight-storage memory, but they do not establish total system RAM or GPU memory requirements.
The RMBG-1.4 model from briaai (RMBG-1.4 (briaai)) has 44 million parameters and listed F32 weights of 0.2 GB.[10] Facebook’s facebook/mask2former-swin-large-cityscapes-semantic has 216 million parameters and listed F32 weights of 0.9 GB.[1] Microsoft’s microsoft/beit-large-finetuned-ade-640-640 has 503 million parameters and listed F32 weights of 2.3 GB.[7] Use those weight sizes as starting points for memory budgeting, rather than complete hardware requirements.
Quantization calculations describe hypothetical weight storage. For briaai’s 44 million parameters, 4-bit weight storage would be approximately 22 MB—estimated (params x 0.5 bytes).[10] For Facebook’s 216 million parameters, the corresponding figure would be approximately 108 MB—estimated (params x 0.5 bytes).[1] Neither calculation establishes that a compatible quantized implementation is available or that inference will fit within that memory.
A practical hardware decision still needs a measured memory requirement for the chosen implementation and workload. Weight size alone cannot justify a specific GPU recommendation, system RAM capacity, or processing-speed claim. Keep the distinction explicit when comparing models: listed weight storage is available here; verified total inference memory and hardware-specific performance are not.
Which segmentation model licenses should you check before deployment?
Check briaai’s RMBG-1.4 (briaai) license before deployment: the checkpoint lists bria-rmbg-1.4, so verify whether its terms permit your intended use.[10] Review commercial use, redistribution and deployment conditions before approving the model for a product. Treat those as questions to resolve from the license text, rather than assuming permission from the availability of downloadable weights.
CIDAS’s CIDAS/clipseg-rd64-refined lists Apache-2.0.[11] Microsoft’s microsoft/beit-large-finetuned-ade-640-640 and microsoft/beit-base-finetuned-ade-640-640 also list Apache-2.0.[7][9] Record the license attached to the exact checkpoint you deploy, and check the applicable requirements for notices, attribution and distributing modifications. Keep that review alongside the model revision in your deployment documentation.
Also verify the license for Facebook’s facebook/mask2former-swin-large-cityscapes-semantic and facebook/mask2former-swin-large-ade-semantic.[1][6] Apply the same check to NVIDIA’s nvidia/segformer-b0-finetuned-ade-512-512 and nvidia/segformer-b2-finetuned-ade-512-512.[3][5] Confirm the terms for each checkpoint individually before approving deployment.
Make the review specific to how you will ship the system: internal inference, a hosted service or weights bundled with an application. Review any accompanying code and dependencies separately. Save the applicable license text with the release record so future model updates trigger a deliberate licensing check.
Frequently Asked Questions
Which segmentation model should I evaluate first for background removal?
briaai’s RMBG-1.4 (briaai) takes first place because its Hugging Face publication date, 2023-12-12, is the newest among the eligible candidates.[10][1][3][5][6][7][9][11] Its benchmark status is “no public benchmark yet”; the position reflects release order, not demonstrated background-removal superiority. The model has 44M parameters, listed F32 weights of 0.2 GB, and the bria-rmbg-1.4 license.[10]
What is the ranking order, and how was it decided?
Ranking note: no public benchmark scores any candidate, so the rule gives: newest release first, then downloads. The order is RMBG-1.4 (briaai) [10]; facebook’s facebook/mask2former-swin-large-cityscapes-semantic [1]; facebook/mask2former-swin-large-ade-semantic [6]; CIDAS’s CIDAS/clipseg-rd64-refined [11]; nvidia’s nvidia/segformer-b0-finetuned-ade-512-512 [3]; nvidia/segformer-b2-finetuned-ade-512-512 [5]; microsoft’s microsoft/beit-large-finetuned-ade-640-640 [7]; microsoft/beit-base-finetuned-ade-640-640.[9] Downloads break publication-date ties; they do not establish background-removal quality. Parameter count does not determine placement.
Why does the selection include older models and established picks?
Ranking note: eligibility admits only image segmentation models tagged image-segmentation from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days.[10] Too few recent candidates qualify, so the selection uses the newest available entries in its candidate set.[10][1][11][3] Established picks are CIDAS/clipseg-rd64-refined, facebook/mask2former-swin-large-cityscapes-semantic and nvidia/segformer-b0-finetuned-ade-512-512 by downloads; their recency exemption never grants first place, and the same ordering rule applies.[11][1][3]
How much GPU memory do these models need locally?
A verified GPU-memory requirement is not available for these candidates. Listed F32 weight sizes are 0.2 GB for RMBG-1.4 (briaai), 0.6 GB for CIDAS/clipseg-rd64-refined, 0.9 GB for each listed facebook Mask2Former checkpoint, and 2.3 GB for microsoft/beit-large-finetuned-ade-640-640.[10][11][1][6][7] Treat those figures as weight sizes, not GPU-fit guarantees. Measure runtime memory with your intended image dimensions and inference configuration before choosing hardware.
Which candidates have an Apache license for a commercial project?
CIDAS/clipseg-rd64-refined, microsoft/beit-large-finetuned-ade-640-640 and microsoft/beit-base-finetuned-ade-640-640 list Apache-2.0.[11][7][9] RMBG-1.4 (briaai) instead lists bria-rmbg-1.4.[10] Read the applicable license before incorporating a checkpoint into a commercial product. License details are unspecified for the listed facebook and nvidia candidates in this comparison; treat commercial clearance for those checkpoints as unresolved rather than assuming permission from public availability.
Do dated benchmarks prove which model removes backgrounds better?
No public benchmark yet scores these candidates against one another for this ranking. The model-card snapshot is dated 2026-09-25; that date is not a background-removal benchmark date.[1][3][5][6][7][9][10][11] No comparative task scores are available here to establish a quality winner. Evaluate candidates on your intended images, and label any model-card benchmark you subsequently use as self-reported.
Sources
- facebook/mask2former-swin-large-cityscapes-semantic model card (Hugging Face) — 2026-09-25
- Masked-attention Mask Transformer for Universal Image Segmentation — 2021-12-02
- nvidia/segformer-b0-finetuned-ade-512-512 model card (Hugging Face) — 2026-09-25
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers — 2021-05-31
- nvidia/segformer-b2-finetuned-ade-512-512 model card (Hugging Face) — 2026-09-25
- facebook/mask2former-swin-large-ade-semantic model card (Hugging Face) — 2026-09-25
- microsoft/beit-large-finetuned-ade-640-640 model card (Hugging Face) — 2026-09-25
- BEiT: BERT Pre-Training of Image Transformers — 2021-06-15
- microsoft/beit-base-finetuned-ade-640-640 model card (Hugging Face) — 2026-09-25
- briaai/RMBG-1.4 model card (Hugging Face) — 2026-09-25
- CIDAS/clipseg-rd64-refined model card (Hugging Face) — 2026-09-25
- Image Segmentation Using Text and Image Prompts — 2021-12-18