Quick Answer
LiquidAI’s LFM2.5-Encoder-350M-PII-Detector is the top pick as of September 2026 because it is the newest release among the listed candidates, with no public benchmark yet [6]. In order, the ranking is LFM2.5-Encoder-350M-PII-Detector (LiquidAI), privacy-filter (OpenAI), gliner-PII (NVIDIA), llmlingua-2-bert-base-multilingual-cased-meetingbank (Microsoft), and llmlingua-2-xlm-roberta-large-meetingbank (Microsoft) [6][3].
Key Takeaways
- LiquidAI’s LFM2.5-Encoder-350M-PII-Detector ranks first on release recency, having been published on June 12, 2026.[6] Its 355M parameters imply 177.5 MB of 4-bit weight storage, estimated (params x 0.5 bytes); context is 128K tokens and the license is lfm1.0.[6]
- OpenAI’s privacy-filter ranks second and offers 128K-token context under Apache-2.0.[3] Its 1.4B parameters imply 700 MB of 4-bit weight storage, estimated (params x 0.5 bytes).[3]
- NVIDIA’s gliner-PII ranks third, with an August 29, 2025 publication date and the nvidia-open-model-license.[5] Parameter count and context length are unspecified here, leaving its weight-storage estimate unresolved.[5]
- Microsoft’s llmlingua-2-bert-base-multilingual-cased-meetingbank ranks fourth: 177M parameters imply 88.5 MB of 4-bit weight storage, estimated (params x 0.5 bytes), with 512-token context and Apache-2.0 licensing.[1] Its associated paper concerns prompt compression.[2]
- Microsoft’s llmlingua-2-xlm-roberta-large-meetingbank ranks fifth: 559M parameters imply 279.5 MB of 4-bit weight storage, estimated (params x 0.5 bytes), with 514-token context and MIT licensing.[4] Its associated paper also concerns prompt compression.[2]
- Ranking note: Scope admits only token classification models from labs with a published paper or leaderboard record, using Hugging Face’s
token-classificationtag.[1][3][4][5][6] Too few releases meet the recency gate, so the newest available are included.[1][3][4][5][6] Every candidate has no public benchmark yet; ordering is newest release first, then downloads, never size.[1][3][4][5][6] Established picks—the three most-downloaded candidates, privacy-filter and both Microsoft entries—bypass the recency gate, follow the same ordering, and cannot take first through that exemption.[1][3][4]
How do these NER models compare on memory estimates, context, licenses and benchmark availability?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| LFM2.5-Encoder-350M-PII-Detector [6] | LiquidAI | 355M [6] | 4-bit weights: 177.5 MB, estimated (params x 0.5 bytes) [6]; runtime memory excluded | 2026-06-12 [6] | lfm1.0 [6] | no public benchmark yet |
| privacy-filter [3] | OpenAI | 1.4B [3] | 4-bit weights: 700 MB, estimated (params x 0.5 bytes) [3]; runtime memory excluded | 2026-04-17 [3] | Apache-2.0 [3] | no public benchmark yet |
| gliner-PII [5] | NVIDIA | — | — | 2025-08-29 [5] | nvidia-open-model-license [5] | no public benchmark yet |
| llmlingua-2-bert-base-multilingual-cased-meetingbank [1] | Microsoft | 177M [1] | 4-bit weights: 88.5 MB, estimated (params x 0.5 bytes) [1]; runtime memory excluded | 2024-03-17 [1] | Apache-2.0 [1] | no public benchmark yet |
| llmlingua-2-xlm-roberta-large-meetingbank [4] | Microsoft | 559M [4] | 4-bit weights: 279.5 MB, estimated (params x 0.5 bytes) [4]; runtime memory excluded | 2024-03-17 [4] | MIT [4] | no public benchmark yet |
Which NER models should you consider for an Apple Silicon Mac?
1. LFM2.5-Encoder-350M-PII-Detector
LFM2.5-Encoder-350M-PII-Detector by LiquidAI ranks first because its June 12, 2026 release is the newest among the eligible candidates.[6][3][5][1][4] The model has 355M parameters, a 128K-token context window, and 1.4 GB of F32 weights.[6] Distribution uses the lfm1.0 license.[6] Its status is “no public benchmark yet,” so its position reflects release order rather than a demonstrated accuracy advantage.
For an Apple Silicon Mac, the 355M parameters imply approximately 177.5 MB of 4-bit weights, estimated (params x 0.5 bytes).[6] Actual hardware requirements remain unverified: that estimate covers weights alone and does not establish total memory consumption or compatibility with a Mac inference runtime. Consider the model for local personally identifiable information detection. The practical caveat is that the advertised 128K-token context does not establish memory use or throughput at that length on Apple Silicon.[6]
Ranking note: this ranking admits only token classification models from labs with a published paper or leaderboard record, using the Hugging Face token-classification pipeline tag. The family has too few recent releases, so the newest available are included.[1][3][4][5][6] With no public benchmark scoring any candidate, the order is “newest release first, then downloads.” Established picks—the three most-downloaded candidates—bypass the recency gate, follow the same ordering rule, and cannot take first place through that exemption.[1][3][4]
2. privacy-filter
privacy-filter by OpenAI ranks second because its April release precedes Liquid AI’s June release of LFM2.5-Encoder-350M-PII-Detector.[3][6] The ordering basis is newest release first, then downloads; privacy-filter has no public benchmark yet, so its position does not establish an accuracy advantage. Its first Hugging Face publication was April 17, 2026, and its license is Apache-2.0.[3]
The model has 1.4B parameters and a context limit of 128K tokens.[3] Quantized weight storage would be approximately 0.7 GB at 4-bit, estimated (params x 0.5 bytes) from the cited 1.4B parameter count.[3] That calculation covers weights alone. For local hardware planning, treat the estimate as a starting point rather than a complete Apple Silicon memory requirement: a specific Mac configuration cannot be recommended from weight storage alone.
Consider privacy-filter for privacy-oriented entity filtering in documents where its 128K-token context could be useful.[3] The practical caveat is deployment uncertainty: the weight estimate does not establish a working quantized runtime, full-context memory consumption or measured Mac performance. Evaluate compatibility and memory use with your intended document lengths before choosing hardware. Its ranking reflects release order, while suitability for your workload still requires validation.
3. gliner-PII
gliner-PII by NVIDIA ranks #3 under the release-date ordering, behind the newer LiquidAI and OpenAI releases.[3][5][6] The ordering is newest release first, then downloads; no public benchmark scores any candidate. The ranking admits only token classification models from labs with a published paper or leaderboard record, using Hugging Face’s token-classification pipeline tag. The shortlist includes older releases because too few recent candidates qualify; gliner-PII was first published on Hugging Face on August 29, 2025.[5]
NVIDIA distributes gliner-PII under the nvidia-open-model-license, and its Hugging Face repository recorded 23,327 downloads over the preceding 30 days as of September 25, 2026.[5] Parameter count, context length and weight size are unavailable for this comparison. Consequently, neither quantized weight memory nor a specific Apple Silicon RAM requirement can be stated responsibly. Local deployment needs validation on the intended Mac before committing to a hardware configuration.
Consider gliner-PII for a local personally identifiable information detection evaluation, with representative documents and entity labels determining its usefulness. The central caveat is “no public benchmark yet”: its position reflects release order, not demonstrated accuracy or Mac performance. Check the license terms and validate installation, memory consumption and detection quality before adopting it.
4. llmlingua-2-bert-base-multilingual-cased-meetingbank
llmlingua-2-bert-base-multilingual-cased-meetingbank by Microsoft ranks fourth as an established pick, with 292,286 downloads in the reporting window ending September 25, 2026.[1] The model has 177M parameters, a context limit of 512 tokens, and an Apache-2.0 license.[1] Weight storage at 4-bit is approximately 88.5 MB, estimated (params x 0.5 bytes) from the 177M parameter count.[1] For Apple Silicon Mac planning, treat that estimate as weight storage only; a total RAM requirement, compatible quantized runtime, and measured Mac performance are not established.
Ranking note: the ordering is newest release first, then downloads, because no public benchmark scores any candidate. The scope admits only token classification models from labs with a published paper or leaderboard record, using Hugging Face’s token-classification pipeline tag. Too few recent releases are available, so the newest available models are included alongside established picks. Microsoft’s model qualifies through the established-pick exemption for the family’s three most-downloaded models; its release date is March 17, 2024.[1][3][4]
The practical use is prompt compression: Microsoft’s LLMLingua-2 work targets efficient, faithful, task-agnostic compression.[2] The caveat for an NER shortlist is that prompt compression does not establish entity-recognition quality. Benchmark status is “no public benchmark yet”; the placement should not be read as demonstrated NER accuracy or a verified Apple Silicon deployment recommendation.
5. llmlingua-2-xlm-roberta-large-meetingbank
llmlingua-2-xlm-roberta-large-meetingbank by Microsoft ranks fifth as an established pick.[4] The ordering is newest release first, then downloads; the family has too few recent releases, so the newest available models and established picks are included.[1][3][4][5][6] Eligibility is limited to token classification models from labs with a published paper or leaderboard record. Microsoft’s model has no public benchmark yet. Its publication date matches the BERT variant, but its lower download count places it behind that model.[1][4]
The model has 559M parameters, a context limit of 514 tokens, F32 weights totaling 2.2 GB, and an MIT license.[4] Quantized weight storage would be approximately 279.5 MB at 4-bit—estimated (params x 0.5 bytes), using the cited 559M parameters.[4] Treat that figure as a weight-storage estimate when planning an Apple Silicon deployment. Allow additional memory for execution and inputs; neither a verified quantized runtime nor a total Mac RAM requirement is established here.
The practical use is task-agnostic prompt compression, the purpose described in Microsoft’s “LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression.”[2] Choose it when the task is shortening prompts while preserving useful information. The caveat for an NER shortlist is that prompt compression does not establish entity-extraction quality: a token-classification label alone does not demonstrate suitability for your entity schema. Validate entity-level behavior before adopting it for extraction.
How can you estimate NER weight memory from parameter counts?
Estimate NER weight memory by multiplying the parameter count by the storage per parameter: for 4-bit weights, label the result “estimated (params x 0.5 bytes)” alongside the cited parameter count. [1][3][4][6] Use a consistent unit when comparing the resulting weight budgets.
Liquid AI’s LFM2.5-Encoder-350M-PII-Detector (LiquidAI) has 355M parameters: estimated (params x 0.5 bytes), its weight memory is 177.5 MB. [6] OpenAI’s privacy-filter (OpenAI) has 1.4B parameters: estimated (params x 0.5 bytes), its weight memory is 700 MB. [3] Both calculations use decimal megabytes and assume the stated storage rate applies across the parameter count.
Microsoft’s llmlingua-2-bert-base-multilingual-cased-meetingbank (Microsoft) has 177M parameters, giving 88.5 MB estimated (params x 0.5 bytes). [1] Microsoft’s llmlingua-2-xlm-roberta-large-meetingbank (Microsoft) has 559M parameters, giving 279.5 MB estimated (params x 0.5 bytes). [4]
Use those estimates as a starting weight budget when planning a local deployment on an Apple Silicon Mac. Keep the calculation separate from a claim about total application memory or whether a particular Mac can run the model. A parameter-based estimate alone does not establish either result. Verify the actual deployment’s memory use before treating the weight budget as a hardware requirement.
Which NER model licenses apply to local deployment?
The listed models use Apache-2.0, MIT, nvidia-open-model-license or lfm1.0; the applicable license depends on the exact model selected for local deployment.[1][3][4][5][6]
Microsoft’s llmlingua-2-bert-base-multilingual-cased-meetingbank (Microsoft) and OpenAI’s privacy-filter (OpenAI) both list Apache-2.0.[1][3] Microsoft’s llmlingua-2-xlm-roberta-large-meetingbank (Microsoft) lists MIT.[4] Record the repository alongside the license in your deployment documentation: the Microsoft models share the LLMLingua-2 family name but list different licenses.[1][4]
NVIDIA’s gliner-PII (NVIDIA) lists nvidia-open-model-license.[5] Liquid AI’s LFM2.5-Encoder-350M-PII-Detector (LiquidAI) lists lfm1.0.[6] Review those named agreements directly when assessing your intended use. Avoid treating either license label as interchangeable with Apache-2.0 or MIT.
For an internal application, describe how the model will be used before reviewing the terms. For a distributed application, include whether the delivery contains model weights or a modified copy. Check the applicable agreement for permissions, conditions and any notice requirements relevant to that plan.
Keep licensing and Apple Silicon compatibility as separate deployment checks. A license label identifies the agreement to review; it does not establish runtime support, quantization support or memory requirements. Choose the exact model first, then document both the license review and the technical deployment requirements.
How are recent releases and established picks ranked without public benchmarks?
Recent releases and established picks are ranked by newest release first, then downloads, because no public benchmark scores any candidate.[1][3][4][5][6]
The ranking admits only token classification models from labs with a published paper or leaderboard record, using Hugging Face’s token-classification pipeline tag. The release window covers the last 12 months, but too few qualifying releases are available, so the newest available candidates are included.[1][3][4][5][6]
The order is LiquidAI’s LFM2.5-Encoder-350M-PII-Detector (LiquidAI),[6] OpenAI’s privacy-filter (OpenAI),[3] NVIDIA’s gliner-PII (NVIDIA),[5] Microsoft’s llmlingua-2-bert-base-multilingual-cased-meetingbank (Microsoft),[1] then Microsoft’s llmlingua-2-xlm-roberta-large-meetingbank (Microsoft).[4] LiquidAI’s detector takes first place because its publication date of June 12, 2026 is the newest among the listed candidates.[1][3][4][5][6]
The established picks are Microsoft’s BERT-based and XLM-RoBERTa-based LLMLingua-2 models and OpenAI’s privacy-filter: the three most-downloaded candidates in the supplied snapshot.[1][3][4][5][6] Established status bypasses the release window, follows the same ordering rule, and cannot by itself secure first place. Both Microsoft models were published on March 17, 2024; their downloads break that date tie.[1][4] Popularity is a tiebreak only, and model size does not determine placement.
Every entry carries “no public benchmark yet.”[1][3][4][5][6] A future leaderboard may compare only models it actually scores and may never elevate an older scored release above a newer unscored release. Model-card benchmark results must be labeled “self-reported.” Read this ordering as a release-based shortlist for evaluation on Apple Silicon Macs, rather than evidence of comparative recognition accuracy or local execution speed.
Frequently Asked Questions
How are the models ranked for Apple Silicon Macs?
Ranking note: this ranking admits only token classification models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: token-classification). Ordering: newest release first, then downloads. Order: LiquidAI’s LFM2.5-Encoder-350M-PII-Detector (LiquidAI) [6]; OpenAI’s privacy-filter (OpenAI) [3]; NVIDIA’s gliner-PII (NVIDIA) [5]; Microsoft’s llmlingua-2-bert-base-multilingual-cased-meetingbank (Microsoft) [1]; Microsoft’s llmlingua-2-xlm-roberta-large-meetingbank (Microsoft) [4]. All have no public benchmark yet. Established picks—privacy-filter and both Microsoft models—are the three most-downloaded candidates; they bypass the recency gate but cannot take first through that exemption. [1][3][4]
Why does LiquidAI’s detector rank first?
LiquidAI’s detector ranks first because its June release follows OpenAI’s April release and the older NVIDIA and Microsoft releases. [6][3][5][1][4] The family has too few releases within the last twelve months, so the newest available are included alongside established picks. [1][3][4][5][6] LiquidAI’s detector has no public benchmark yet; its position reflects release order, not demonstrated superiority in entity recognition.
How much memory would quantized model weights need?
For 4-bit weights: LiquidAI, 355M parameters → 177.5 MB estimated (params x 0.5 bytes) [6]; OpenAI, 1.4B parameters → 700 MB estimated (params x 0.5 bytes) [3]; Microsoft BERT, 177M parameters → 88.5 MB estimated (params x 0.5 bytes) [1]; Microsoft XLM-RoBERTa, 559M parameters → 279.5 MB estimated (params x 0.5 bytes) [4]. Treat these figures as weight estimates, not confirmed application memory requirements or evidence of working Mac quantization support.
Which candidates support long documents?
OpenAI’s privacy-filter and LiquidAI’s detector each list a context length of 128K tokens. [3][6] Microsoft’s BERT variant lists 512 tokens, while its XLM-RoBERTa variant lists 514 tokens. [1][4] Use those limits to shortlist candidates for your document lengths. Avoid treating a published context limit as a measured Apple Silicon speed result or a guarantee that a complete workload fits your Mac.
What licenses do these models use?
OpenAI’s privacy-filter and Microsoft’s BERT variant use Apache-2.0. [3][1] Microsoft’s XLM-RoBERTa variant uses MIT. [4] NVIDIA’s gliner-PII uses nvidia-open-model-license, while LiquidAI’s detector uses lfm1.0. [5][6] Check the applicable license text against your intended deployment and redistribution plans. Avoid assuming that every downloadable model grants identical permissions, especially when choosing a model for integration into a shipped product.
Are the Microsoft models intended specifically for entity extraction?
Microsoft’s accompanying paper, “LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression,” describes prompt compression. [2] Their inclusion follows the token-classification scope; avoid interpreting that inclusion as demonstrated NER accuracy. Both Microsoft candidates have no public benchmark yet for this ranking. Evaluate their outputs against your required entity labels before selecting either for an extraction workflow.
Sources
- microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank model card (Hugging Face) — 2026-09-25
- LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression — 2024-03-19
- openai/privacy-filter model card (Hugging Face) — 2026-09-25
- microsoft/llmlingua-2-xlm-roberta-large-meetingbank model card (Hugging Face) — 2026-09-25
- nvidia/gliner-PII model card (Hugging Face) — 2026-09-25
- LiquidAI/LFM2.5-Encoder-350M-PII-Detector model card (Hugging Face) — 2026-09-25