Best NER Models for CPU-Only PCs in 2026

Rankings 2026-09-25 Last updated 2026-09-25 14 min read By Q4KM

Quick Answer

GLiNER2.5-multi-Decide (fastino) is the top pick as of September 2026 because it is the newest eligible release; the ranking admits only token classification models (Hugging Face pipeline tag: token-classification), with no public benchmark yet for any candidate, so the ordering is newest release first, then downloads. [1][2][3][5][6][7][10][11]

The ranking is: 1. In order, the ranking is GLiNER2.5-multi-Decide (fastino), gliformer-large-v1 (knowledgator), gliner2.5-multi-v1 (fastino), rampart (nationaldesignstudio), LFM2.5-Encoder-350M-PII-Detector (LiquidAI), privacy-filter-nemotron-v2 (OpenMed), privacy-filter-nemotron (OpenMed), and privacy-filter (openai) [1][2].

Key Takeaways

How do these NER models compare on memory, context, licenses and published benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
GLiNER2.5-multi-Decide [5] fastino 287M [5] 4-bit weights: 143.5 MB, estimated (params x 0.5 bytes) [5] 2026-09-24 [5] Apache-2.0 [5] no public benchmark yet
gliformer-large-v1 [7] knowledgator not published not published 2026-09-11 [7] Apache-2.0 [7] no public benchmark yet
gliner2.5-multi-v1 [3] fastino 287M [3] 4-bit weights: 143.5 MB, estimated (params x 0.5 bytes) [3] 2026-08-14 [3] Apache-2.0 [3] no public benchmark yet
rampart [6] nationaldesignstudio not published not published 2026-06-28 [6] CC-BY-4.0 [6] no public benchmark yet
LFM2.5-Encoder-350M-PII-Detector [2] LiquidAI 355M [2] 4-bit weights: 177.5 MB, estimated (params x 0.5 bytes) [2] 2026-06-12 [2] lfm1.0 [2] no public benchmark yet
privacy-filter-nemotron-v2 [10] OpenMed 1.4B [10] 4-bit weights: 700 MB, estimated (params x 0.5 bytes) [10] 2026-05-03 [10] not published no public benchmark yet
privacy-filter-nemotron [11] OpenMed 1.4B [11] 4-bit weights: 700 MB, estimated (params x 0.5 bytes) [11] 2026-04-24 [11] Apache-2.0 [11] no public benchmark yet
privacy-filter [1] openai 1.4B [1] 4-bit weights: 700 MB, estimated (params x 0.5 bytes) [1] 2026-04-17 [1] Apache-2.0 [1] no public benchmark yet

Which NER models should you consider for a CPU-only PC?

1. GLiNER2.5-multi-Decide

GLiNER2.5-multi-Decide by fastino ranks first on release recency, with a Hugging Face publication date of 2026-09-24.[5] The model has 287M parameters, F32 weights occupying 1.1 GB, and an Apache-2.0 license.[5] Consider it for local named entity recognition when permissive licensing matters and you can validate extraction quality against your own documents.

For CPU-only planning, its 4-bit weight footprint is approximately 143.5 MB, estimated (params x 0.5 bytes) from 287M parameters.[5] Treat that as a weight-only estimate, not a tested RAM requirement or confirmation of CPU quantization support. A specific processor and total system RAM recommendation cannot be established from the available specifications. Context length is unspecified, and the model has no public benchmark yet; CPU throughput also remains unverified.

Ranking note: this ranking admits only token classification models with the Hugging Face pipeline tag token-classification. The ordering is newest release first, then downloads; placement therefore reflects release timing rather than demonstrated extraction accuracy or CPU speed.

2. gliformer-large-v1

gliformer-large-v1 by knowledgator ranks second under the ordering rule “newest release first, then downloads,” with a Hugging Face publication date of September 11, 2026.[5][7] The ranking admits only token classification models with the Hugging Face token-classification pipeline tag. Its position reflects release timing; gliformer-large-v1 has no public benchmark yet.

The model uses the Apache-2.0 license.[7] Parameter count, context length and weight size are unspecified here, so a quantized weight-memory estimate would be speculative. A concrete CPU or RAM requirement also remains unestablished. Before committing to local deployment, check the model’s runtime requirements and measure memory use on your intended PC.

Consider gliformer-large-v1 for evaluating token classification in a project that requires Apache-2.0 licensing.[7] The caveat is the missing performance evidence: its ranking does not establish extraction accuracy or CPU throughput. Validate both against representative documents before selecting it for production.

3. gliner2.5-multi-v1

gliner2.5-multi-v1 by fastino ranks third under the ordering rule “newest release first, then downloads,” following its Hugging Face publication on August 14, 2026.[3][5][7] The ranking admits only token classification models with the Hugging Face pipeline tag token-classification. Its benchmark status is “no public benchmark yet,” so its position does not establish an accuracy or CPU-speed advantage.

The model has 287M parameters, F32 weights totaling 1.1 GB, and an Apache-2.0 license.[3] Its 4-bit weight payload would be 143.5 MB, estimated (params x 0.5 bytes) from that parameter count.[3] For local CPU deployment, use this estimate for preliminary capacity planning; it does not establish total system RAM requirements or confirm that a compatible quantized build exists.

Consider it for schema-driven information extraction, the workflow described in GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface.[4] Before committing a CPU-only PC, verify runtime compatibility, context limits, and throughput on representative documents. The practical caveat is that weight size alone cannot establish usable inference speed.

4. rampart

rampart by nationaldesignstudio takes fourth place with a release date of June 28, 2026 [6]. The ordering basis is “newest release first, then downloads”; rampart has no public benchmark yet, so its position does not establish an accuracy or CPU-speed advantage. The ranking admits only token classification models with Hugging Face’s token-classification pipeline tag.

rampart has a context limit of 512 tokens and uses the CC-BY-4.0 license [6]. Short text segments that fit within that limit are a practical starting point for local evaluation. Longer documents require a plan for handling text beyond the context window; document-level suitability should not be assumed from the model’s inclusion here.

Hardware sizing remains unresolved: the available specifications do not establish a parameter count, quantized weight footprint, or required system RAM [6]. A specific CPU-only PC configuration therefore cannot be recommended confidently. Evaluate memory use and processing time on your intended machine before adopting rampart.

5. LFM2.5-Encoder-350M-PII-Detector

LFM2.5-Encoder-350M-PII-Detector by LiquidAI ranks here because of its June 12, 2026 release date [2], under the ordering rule “newest release first, then downloads.” The ranking admits only token classification models with the Hugging Face pipeline tag token-classification. Benchmark status: no public benchmark yet; its position does not establish measured CPU performance.

The model has 355M parameters [2], a 128K-token context window [2], and published F32 weights occupying 1.4 GB [2]. Weight storage at 4-bit precision would be approximately 177.5 MB, estimated (params x 0.5 bytes) from the 355M parameter count [2]. For a CPU-only PC, treat that estimate as a weight-storage budget, not a complete RAM requirement or confirmation that a compatible quantized runtime is available.

Consider it for local personally identifiable information detection when a long context window matters. The hardware caveat is that total runtime RAM and CPU throughput are unspecified, so validate both with representative documents before deployment. Distribution uses the lfm1.0 license [2]; review its terms before integration.

6. privacy-filter-nemotron-v2

privacy-filter-nemotron-v2 by OpenMed ranks sixth under the rule “newest release first, then downloads,” with a release date of May 3, 2026.[10] The ranking admits only token classification models with the Hugging Face pipeline tag token-classification. The model has no public benchmark yet, so its position does not establish an accuracy or CPU-speed advantage.

OpenMed lists 1.4B parameters, a 128K-token context window and BF16 weights occupying 2.8 GB.[10] For CPU-only planning, weight storage at 4-bit is approximately 0.7 GB—estimated (params x 0.5 bytes), using the cited 1.4B parameter count.[10] That calculation describes weight storage alone; it does not establish total RAM requirements or confirm a working quantized runtime.

Consider privacy-filter-nemotron-v2 for evaluating privacy filtering on long inputs, given its documented 128K-token context window.[10] The practical caveat is deployment uncertainty: a specific CPU or RAM recommendation cannot be justified from the listed specifications, and the license is not specified in the available model details.[10]

7. privacy-filter-nemotron

privacy-filter-nemotron by OpenMed ranks seventh under the “newest release first, then downloads” ordering, with a Hugging Face publication date of 2026-04-24.[11] The ranking admits only token classification models with Hugging Face’s token-classification pipeline tag. Its position reflects release order rather than demonstrated CPU performance: no public benchmark yet.

The model has 1.4B parameters, a 128K-token context window and an Apache-2.0 license.[11] Published BF16 weights occupy 2.8 GB.[11] For CPU memory planning, the 4-bit weight payload is approximately 0.7 GB, estimated (params x 0.5 bytes) from the cited 1.4B parameters.[11] Treat that estimate as a weight-only calculation; total runtime RAM requirements and a compatible CPU quantization setup remain unverified.

Consider it for local privacy-filtering workflows where the documented context window and license fit your requirements. The practical caveat is its successor: OpenMed’s privacy-filter-nemotron-v2 was published on 2026-05-03.[10] Evaluate that newer release before committing to this version; the ordering does not establish an accuracy or speed advantage.

8. privacy-filter

privacy-filter by openai ranks eighth under the “newest release first, then downloads” ordering, following its April 17, 2026 Hugging Face debut.[1] Its position reflects release timing rather than demonstrated CPU performance: no public benchmark yet.

The model has 1.4B parameters, a 128K-token context window and an Apache-2.0 license.[1] Published BF16 weights occupy 5.6 GB.[1] For hardware planning, 4-bit weight storage would be 0.7 GB, estimated (params x 0.5 bytes) from the cited 1.4B parameters.[1] Treat that calculation as a weight-storage estimate, not a validated system RAM requirement or confirmation that a compatible quantized CPU runtime is available.

Consider privacy-filter for local document privacy filtering when its advertised 128K-token context is relevant.[1] Before committing a CPU-only PC, verify runtime compatibility and measure RAM use at your intended input length; the available specifications do not establish a tested hardware minimum.

How can you estimate NER model weight memory from parameter counts?

Estimate NER model weight memory by multiplying its parameter count by the assumed storage per parameter. For a four-bit calculation, use half a byte per parameter and label the result as an estimate, rather than a measured memory requirement.

For fastino’s GLiNER2.5-multi-Decide (fastino), the parameter count is 287M [5], giving 143.5 MB estimated (params x 0.5 bytes) [5]. LiquidAI’s LFM2.5-Encoder-350M-PII-Detector (LiquidAI) has 355M parameters [2], giving 177.5 MB estimated (params x 0.5 bytes) [2]. OpenAI’s privacy-filter (openai) has 1.4B parameters [1], giving 700 MB estimated (params x 0.5 bytes) [1]. Those estimates use decimal megabytes.

Keep the estimated quantized weight size separate from the published weight size. The fastino model lists F32 weights at 1.1 GB [5], while the LiquidAI model lists F32 weights at 1.4 GB [2]. Neither published size should be presented as a measured four-bit footprint.

Use the calculation as a weight-storage planning figure, not a guarantee that the model will fit within an equivalent amount of system RAM. Before choosing a model for a CPU-only PC, verify the actual quantized artifact and measure memory use with your intended runtime and input length. Parameter arithmetic alone does not establish quantization availability, CPU speed, or extraction accuracy.

Which NER models publish context limits and license terms?

Published context limits and license terms are available for openai’s privacy-filter [1], LiquidAI’s LFM2.5-Encoder-350M-PII-Detector [2], OpenMed’s privacy-filter-nemotron [11], and nationaldesignstudio’s rampart [6].

The openai privacy-filter supports a context length of 128K tokens and uses Apache-2.0 [1]. OpenMed privacy-filter-nemotron also lists 128K tokens and Apache-2.0 [11]. LiquidAI LFM2.5-Encoder-350M-PII-Detector lists the same 128K-token context limit, but its license is lfm1.0 [2]. Review that license’s terms against your intended use before deployment.

The nationaldesignstudio rampart lists a context limit of 512 tokens and a CC-BY-4.0 license [6]. For document processing, check whether your intended input fits that published limit before selecting the model.

OpenMed’s privacy-filter-nemotron-v2 publishes a 128K-token context limit [10]. Verify its license separately before deployment; do not assume the Apache-2.0 terms listed for OpenMed privacy-filter-nemotron also cover this release [11].

Among the remaining candidates, fastino’s GLiNER2.5-multi-Decide [5], fastino’s gliner2.5-multi-v1 [3], and knowledgator’s gliformer-large-v1 [7] list Apache-2.0 licenses. The zilliz organization’s semantic-highlight-bilingual-v1 lists MIT [8]. Check context limits separately for those models before committing to an input format or document-processing workflow.

How are NER models ranked when there is no public benchmark yet?

NER models with no public benchmark yet are ordered by newest release first, then downloads; the order indicates release recency, not demonstrated accuracy or CPU speed.

The ranking admits only token classification models with the Hugging Face pipeline tag token-classification. Eligibility is restricted to recent releases in the candidate list. Model size does not determine position, and popularity serves only as a tiebreaker.

The first position goes to GLiNER2.5-multi-Decide (fastino) because its September 24, 2026 publication date puts it first under the release-date rule.[5] gliformer-large-v1 (knowledgator) follows with a September 11, 2026 release, ahead of gliner2.5-multi-v1 (fastino), published August 14, 2026.[7][3]

Every candidate carries the label “no public benchmark yet.” A leaderboard may compare only models evaluated on that board. An older scored model cannot displace a newer unscored candidate under this ranking policy. Any model-card benchmark must be identified as “self-reported.”

For CPU-only selection, read the order alongside parameter counts, published weight sizes, context limits and license terms. Calculated quantized weight memory must be labeled as an estimate with its parameter-based calculation. Missing specifications remain unspecified. The ranking provides a starting order for evaluation; choosing a deployment still requires checking the intended workload and hardware.

Frequently Asked Questions

Which NER model comes first for CPU-only PCs?

fastino’s GLiNER2.5-multi-Decide (fastino) takes the first position because its September 24, 2026 release is the newest among the admitted candidates [5]. Ranking note: newest release first, then downloads; no public benchmark scores any candidate against the others. Scope: this ranking admits only token classification models with the Hugging Face pipeline tag token-classification. The first position reflects release recency, rather than demonstrated CPU speed or recognition accuracy.

What is the complete model ranking?

The ranking is fastino’s GLiNER2.5-multi-Decide (fastino) [5]; knowledgator’s gliformer-large-v1 (knowledgator) [7]; fastino’s gliner2.5-multi-v1 (fastino) [3]; nationaldesignstudio’s rampart (nationaldesignstudio) [6]; LiquidAI’s LFM2.5-Encoder-350M-PII-Detector (LiquidAI) [2]; OpenMed’s privacy-filter-nemotron-v2 (OpenMed) [10]; OpenMed’s privacy-filter-nemotron (OpenMed) [11]; openai’s privacy-filter (openai) [1]; and zilliz’s zilliz/semantic-highlight-bilingual-v1 [8]. Every entry has the status “no public benchmark yet.” Release recency determines this order; downloads serve only as a tiebreaker, and parameter count does not determine placement.

How much memory would quantized model weights need?

For 4-bit weights, GLiNER2.5-multi-Decide (fastino) has 287M parameters: 143.5 MB estimated (params x 0.5 bytes) [5]. LFM2.5-Encoder-350M-PII-Detector (LiquidAI) has 355M parameters: 177.5 MB estimated (params x 0.5 bytes) [2]. privacy-filter (openai) has 1.4B parameters: 700 MB estimated (params x 0.5 bytes) [1]. Treat those calculations as weight-storage estimates, not measured application RAM requirements or confirmation that compatible quantized packages are available.

Which models have a documented context length for long documents?

privacy-filter (openai) [1], LFM2.5-Encoder-350M-PII-Detector (LiquidAI) [2], privacy-filter-nemotron-v2 (OpenMed) [10], and privacy-filter-nemotron (OpenMed) [11] each list a context length of 128K tokens [1][2][10][11]. rampart (nationaldesignstudio) lists 512 tokens [6]. Use those limits when assessing document fit. A documented context limit does not establish practical CPU latency or memory consumption, so avoid treating long-context support as evidence that full-length documents will run comfortably on your PC.

Which licenses should I check before choosing a model?

Apache-2.0 is listed for GLiNER2.5-multi-Decide (fastino) [5], gliformer-large-v1 (knowledgator) [7], gliner2.5-multi-v1 (fastino) [3], privacy-filter-nemotron (OpenMed) [11], and privacy-filter (openai) [1]. LFM2.5-Encoder-350M-PII-Detector (LiquidAI) lists lfm1.0 [2]; rampart (nationaldesignstudio) lists CC-BY-4.0 [6]; zilliz/semantic-highlight-bilingual-v1 lists MIT [8]. Check the applicable terms against your intended deployment. Verify privacy-filter-nemotron-v2 (OpenMed)’s license separately rather than assuming it inherits the license of another OpenMed release.

Does the ranking show which model is fastest or most accurate on my CPU?

No public benchmark yet is the status of every ranked candidate, so the ordering establishes neither CPU throughput nor comparative recognition accuracy. For example, GLiNER2.5-multi-Decide (fastino) lists 287M parameters [5], but parameter count alone is not a measured performance result. Evaluate candidates with representative documents on your own PC, recording recognition quality, processing time, and peak RAM before choosing a deployment model.

Sources

  1. openai/privacy-filter model card (Hugging Face) — 2026-09-25
  2. LiquidAI/LFM2.5-Encoder-350M-PII-Detector model card (Hugging Face) — 2026-09-25
  3. fastino/gliner2.5-multi-v1 model card (Hugging Face) — 2026-09-25
  4. GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface — 2025-07-24
  5. fastino/GLiNER2.5-multi-Decide model card (Hugging Face) — 2026-09-25
  6. nationaldesignstudio/rampart model card (Hugging Face) — 2026-09-25
  7. knowledgator/gliformer-large-v1 model card (Hugging Face) — 2026-09-25
  8. zilliz/semantic-highlight-bilingual-v1 model card (Hugging Face) — 2026-09-25
  9. Provence: efficient and robust context pruning for retrieval-augmented generation — 2025-01-27
  10. OpenMed/privacy-filter-nemotron-v2 model card (Hugging Face) — 2026-09-25
  11. OpenMed/privacy-filter-nemotron model card (Hugging Face) — 2026-09-25

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog