Best Reranker Models for CPU-Only PCs in 2026: jina-reranker-v3.5 and R3-rerank-0.6b

Rankings 2026-09-26 Last updated 2026-09-26 14 min read By Q4KM

Quick Answer

jina-reranker-v3.5 (jinaai) is the top pick for CPU-only PCs as of September 2026 because it is the newest eligible release [7][9][2][4][5][1]. In order, the ranking is jina-reranker-v3.5 (jinaai), R3-rerank-0.6b (tencent), Qwen3-VL-Reranker-2B (Qwen), Qwen3-VL-Reranker-8B (Qwen), llama-nemotron-rerank-vl-1b-v2 (nvidia), and llama-nemotron-rerank-1b-v2 (nvidia) [7][9].

Key Takeaways

How do local rerankers compare on weight memory, context, licenses and benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
jina-reranker-v3.5 [7] jinaai [7] 597M [7] 4-bit weights: 298.5 MB, estimated (597M params x 0.5 bytes) [7] 2026-07-14 [7] cc-by-nc-4.0 [7] no public benchmark yet [7]
R3-rerank-0.6b [9] tencent [9] 596M [9] 4-bit weights: 298 MB, estimated (596M params x 0.5 bytes) [9] 2026-07-08 [9] Apache-2.0 [9] no public benchmark yet [9]
Qwen3-VL-Reranker-2B [2] Qwen [2] 2.1B [2] 4-bit weights: 1.05 GB, estimated (2.1B params x 0.5 bytes) [2] 2026-01-07 [2] Apache-2.0 [2] no public benchmark yet [2]
Qwen3-VL-Reranker-8B [4] Qwen [4] 8.8B [4] 4-bit weights: 4.4 GB, estimated (8.8B params x 0.5 bytes) [4] 2026-01-07 [4] Apache-2.0 [4] no public benchmark yet [4]
llama-nemotron-rerank-vl-1b-v2 [5] nvidia [5] 1.7B [5] 4-bit weights: 850 MB, estimated (1.7B params x 0.5 bytes) [5] 2025-12-04 [5] nvidia-open-model-license [5] no public benchmark yet [5]
llama-nemotron-rerank-1b-v2 [1] nvidia [1] 1.2B [1] 4-bit weights: 600 MB, estimated (1.2B params x 0.5 bytes) [1] 2025-10-16 [1] openmdw-1.1 [1] no public benchmark yet [1]

Which reranker models should you consider for a CPU-only PC?

1. jina-reranker-v3.5

jina-reranker-v3.5 by jinaai ranks first because its July 14, 2026 release is the newest among the eligible candidates.[1][2][4][5][7][9] Ranking note: newest release first, then downloads; no public benchmark yet, so placement does not establish a performance advantage. The ranking admits only rerankers from labs with a published paper or leaderboard record, using Hugging Face’s text-ranking pipeline tag; embedding models are excluded. The accompanying paper describes a listwise reranker with hybrid attention and self-distillation.[8]

The model has 597M parameters, a context length of 128K tokens, and BF16 weights totaling 1.2 GB.[7] Its 4-bit weight footprint is approximately 298.5 MB, estimated (params x 0.5 bytes) from the cited 597M parameters.[7] Treat that estimate as a weight-storage budget when assessing a CPU-only PC. A complete RAM requirement and CPU execution speed remain unestablished; check CPU-runtime and quantization support before committing to a local deployment.

Consider jina-reranker-v3.5 for local, noncommercial reranking experiments where its listwise design is relevant to the retrieval workflow.[7][8] Evaluate ranking quality and latency with your own queries before choosing it for routine use. The licensing caveat matters: the model uses cc-by-nc-4.0, so a commercial deployment needs a separate licensing assessment.[7]

2. R3-rerank-0.6b

R3-rerank-0.6b by tencent ranks second under the ordering rule: newest release first, then downloads.[7][9] Its publication date falls behind jinaai’s jina-reranker-v3.5, so the position reflects release order rather than demonstrated ranking quality.[7][9] Published on July 8, 2026, the model has 596M parameters, a 40K-token context window and an Apache-2.0 license.[9] Its benchmark status is “no public benchmark yet”; the position should not be read as a CPU performance result.

For CPU-only planning, weight storage would be 298 MB at 4-bit, estimated (params x 0.5 bytes) from the 596M parameter count.[9] The published F32 weights occupy 2.4 GB.[9] Treat the quantized figure as a weight-storage budget, not a complete RAM requirement. Before choosing hardware, confirm that your intended runtime supports the model on CPU and measure total memory use with your workload. Neither a specific RAM capacity nor CPU throughput is established.

A practical use is evaluating local reranking for a project that requires Apache-2.0 licensing and wants to investigate a compact parameter footprint.[9] Start with representative queries and candidate documents, then check relevance and latency together. The caveat is deployment uncertainty: estimated weight storage alone does not establish usable CPU speed or confirm that a compatible quantized implementation is available.

3. Qwen3-VL-Reranker-2B

Qwen3-VL-Reranker-2B by Qwen ranks here under the ordering rule: newest release first, then downloads. Its January 7, 2026 release follows the July releases from jinaai and tencent; its download count places it ahead of Qwen3-VL-Reranker-8B by Qwen, released on the same date.[2][4][7][9] The ranking admits only rerankers from labs with a published paper or leaderboard record, using Hugging Face’s text-ranking pipeline tag; embedding models are excluded. Benchmark status: no public benchmark yet.

The model has 2.1B parameters, a 256K-token context window and an Apache-2.0 license.[2] Published BF16 weights occupy 4.3 GB.[2] For CPU-only hardware planning, 4-bit weight storage is approximately 1.05 GB, estimated (params x 0.5 bytes) from the cited 2.1B parameters.[2] Treat that estimate as a weight-storage budget, not a total system-RAM requirement. A specific RAM capacity or CPU configuration cannot be recommended from these specifications alone.

Consider the model for local multimodal reranking when its context allowance and license match your application; Qwen describes the family as a framework for multimodal retrieval and ranking.[3] The practical caveat is that its position does not establish CPU speed or ranking quality. Before committing to deployment, check runtime compatibility and measure latency and peak memory with your intended documents, candidate count and input lengths.

4. Qwen3-VL-Reranker-8B

Qwen3-VL-Reranker-8B by Qwen ranks fourth under the ordering rule “newest release first, then downloads”: its publication date matches Qwen3-VL-Reranker-2B by Qwen, but its download count is lower.[2][4] Both appeared on January 7, 2026.[2][4] Benchmark status: no public benchmark yet. The ranking admits only rerankers from labs with a published paper or leaderboard record, using Hugging Face’s text-ranking pipeline tag; embedding models are excluded.

The model has 8.8 billion parameters, a 256K-token context window, and an Apache-2.0 license.[4] Published BF16 weights occupy 17.5 GB.[4] Weight storage at 4-bit precision is approximately 4.4 GB, estimated (params x 0.5 bytes) from the cited 8.8 billion parameters.[4] Treat that calculation as a planning estimate for quantized weights, not a measured download size or a complete system-RAM requirement.

For a CPU-only PC, consider it for multimodal ranking experiments where Qwen’s unified retrieval and ranking framework matches the workload.[3] Local hardware planning must leave room beyond the estimated weights for execution; a specific RAM configuration cannot be recommended from weight size alone. The practical caveat is that CPU latency and total runtime memory are unverified, so the advertised context window should not be treated as a promise of comfortable desktop performance.

5. llama-nemotron-rerank-vl-1b-v2

llama-nemotron-rerank-vl-1b-v2 by nvidia ranks fifth under the ordering rule “newest release first, then downloads”: its publication date falls between the Qwen releases and nvidia’s text reranker.[1][2][4][5] Published on 2025-12-04, the model has no public benchmark yet, so its position does not establish a ranking-quality advantage.[5] Consider the placement a recency-based selection order, rather than evidence of CPU performance.

The model has 1.7B parameters, a 128K-token context window and 3.4 GB of BF16 weights.[5] Alongside that parameter count, the weight-only memory budget at 4-bit is 0.85 GB, estimated (params x 0.5 bytes).[5] For a local CPU-only PC, that estimate is a starting point for memory planning. Available RAM must also accommodate the runtime and workload; the weight calculation does not establish a complete system requirement or confirm a compatible quantized runtime.

A practical use to evaluate is local document reranking when a long context window matters: the advertised limit is 128K tokens.[5] Choose a representative workload and check memory consumption and response time before committing to deployment. The central caveat is that context capacity does not establish acceptable CPU latency. Distribution uses the nvidia-open-model-license, which should be reviewed against the intended deployment.[5]

6. llama-nemotron-rerank-1b-v2

llama-nemotron-rerank-1b-v2 by nvidia ranks sixth because its first Hugging Face publication, October 16, 2025, precedes the other eligible releases.[1][2][4][5][7][9] The ordering basis is newest release first, then downloads; its benchmark status is no public benchmark yet, so its position does not establish relative ranking quality. The ranking admits only rerankers from labs with a published paper or leaderboard record and the Hugging Face text-ranking pipeline tag; embedding models are excluded.

The model has 1.2B parameters, a context length of 128K tokens, and BF16 weights totaling 2.5 GB.[1] Its license is openmdw-1.1.[1] For CPU-only planning, the 4-bit weight footprint is approximately 0.60 GB, estimated (params x 0.5 bytes) from the cited 1.2B parameters.[1] Treat that calculation as a weight-storage estimate rather than a complete system RAM requirement. A specific RAM configuration or CPU speed recommendation cannot be established from the listed specifications.

Consider nvidia’s model for local text reranking when a context allowance of 128K tokens matters to your workload.[1] Start with representative query-and-document inputs and measure memory use and latency before choosing deployment settings. The practical caveat is that neither the context specification nor the estimated weight footprint establishes acceptable CPU performance; validate both with your intended runtime and workload.

How can you estimate reranker weight memory from parameter counts?

Estimate reranker weight memory by multiplying the published parameter count by the assumed storage per parameter; for 4-bit weights, label the calculation “estimated (params x 0.5 bytes).” [7][9] Keep the precision assumption explicit so readers can distinguish a calculated weight budget from a published download size.

For jinaai’s jina-reranker-v3.5, the parameter count is 597M [7]: weight memory is 298.5 MB, estimated (params x 0.5 bytes). [7] For tencent’s R3-rerank-0.6b, the parameter count is 596M [9]: weight memory is 298 MB, estimated (params x 0.5 bytes). [9] Use decimal megabytes and gigabytes consistently when presenting these calculations.

The same calculation applies to larger parameter counts. Qwen’s Qwen3-VL-Reranker-2B has 2.1B parameters [2]: weight memory is 1.05 GB, estimated (params x 0.5 bytes). [2] Qwen’s Qwen3-VL-Reranker-8B has 8.8B parameters [4]: weight memory is 4.4 GB, estimated (params x 0.5 bytes). [4]

Treat each result as a weight-storage estimate rather than a total RAM requirement. Keep runtime memory, context handling and CPU execution speed as separate checks when evaluating a local setup. A parameter-count calculation alone does not establish that a compatible quantized artifact exists, that a CPU runtime supports it, or that the model will run comfortably on a particular PC.

Which reranker licenses suit commercial local deployments?

The Apache-2.0 licenses on Qwen’s Qwen3-VL-Reranker-2B, Qwen’s Qwen3-VL-Reranker-8B and Tencent’s R3-rerank-0.6b suit commercial local deployments, subject to compliance with their license terms.[2][4][9] For a commercial shortlist, start with those models and assess hardware requirements separately. A shared license does not establish equal suitability for your application.

Jina AI’s jina-reranker-v3.5 uses cc-by-nc-4.0, which includes a noncommercial restriction.[7] Treat commercial deployment as requiring separate permission from the rights holder. Running the model locally does not remove that restriction, so avoid selecting it for a commercial product on the assumption that downloadable weights provide commercial authorization.

NVIDIA’s llama-nemotron-rerank-1b-v2 uses openmdw-1.1, while NVIDIA’s llama-nemotron-rerank-vl-1b-v2 uses nvidia-open-model-license.[1][5] Evaluate those licenses separately before approving either model. Similar model names do not justify carrying a license decision from one repository to the other. Check the applicable terms against your intended internal use, customer-facing service or redistribution.

For deployment review, record the exact model repository, selected revision and accompanying license. Distinguish permission to use the model commercially from obligations attached to distributing weights or a packaged application. Keep the license decision alongside the deployment configuration so a later model replacement receives its own review.

What do context limits and missing benchmarks tell you about CPU suitability?

Context limits tell you how much input a reranker can accept; missing benchmarks leave CPU latency and comparative ranking quality unresolved. A supported context window is a capacity specification, not evidence that processing a full window is practical on your PC.

Qwen’s Qwen3-VL-Reranker-2B and Qwen3-VL-Reranker-8B each specify 256K tokens.[2][4] NVIDIA’s llama-nemotron-rerank-vl-1b-v2 and llama-nemotron-rerank-1b-v2 each specify 128K tokens.[5][1] Those limits alone cannot establish a CPU performance advantage or tell you how long a reranking request will take.

Jina AI’s jina-reranker-v3.5 specifies 128K tokens with 597M parameters; its 4-bit weight storage is approximately 298.5 MB, estimated (params x 0.5 bytes).[7] Tencent’s R3-rerank-0.6b specifies 40K tokens with 596M parameters; its 4-bit weight storage is approximately 298 MB, estimated (params x 0.5 bytes).[9] Similar parameter counts therefore do not imply similar context limits. Both calculations describe weight storage, not measured total RAM use or confirmed quantized CPU support.

Every candidate has no public benchmark yet for the comparison here. Release order consequently provides no evidence of CPU speed or ranking accuracy. Any model-card benchmark should be described as self-reported, without treating it as an independent comparison.

For a CPU-only deployment, use context limits to check whether your intended inputs are supported. Then validate ranking quality, latency and peak RAM with your own documents and candidate counts before choosing a model. Neither a context specification nor a weight-storage estimate establishes practical CPU suitability.

Frequently Asked Questions

Which reranker should I consider first for a CPU-only PC?

jinaai’s jina-reranker-v3.5 takes first place because its July 14, 2026 release is the newest eligible release. [7][9][2][4][5][1] Ranking note: newest release first, then downloads. The ranking admits only rerankers from labs with a published paper or leaderboard record (Hugging Face pipeline tags: text-ranking); embedding models are excluded. Every candidate has no public benchmark yet, so the ordering does not establish a CPU performance winner. [7][9][2][4][5][1]

How are the eligible rerankers ordered?

The order is jinaai’s jina-reranker-v3.5 [7], Tencent’s R3-rerank-0.6b [9], Qwen’s Qwen3-VL-Reranker-2B [2], Qwen’s Qwen3-VL-Reranker-8B [4], NVIDIA’s llama-nemotron-rerank-vl-1b-v2 [5], then NVIDIA’s llama-nemotron-rerank-1b-v2 [1]. The Qwen models share a January 7, 2026 publication date; their downloads over the last 30 days, as of September 25, 2026, break that tie: 686,616 versus 114,106. [2][4] Downloads determine only that tie, not a claim of superior reranking quality.

How much memory would quantized weights need?

For jina-reranker-v3.5, 597M parameters imply 0.2985 GB of 4-bit weights, estimated (params x 0.5 bytes). [7] For R3-rerank-0.6b, 596M parameters imply 0.298 GB, estimated (params x 0.5 bytes). [9] Both calculations describe weight storage only. Neither calculation establishes total application RAM requirements, confirms an available quantized package, or demonstrates acceptable execution speed on your CPU.

Which licenses should I check before choosing a model?

R3-rerank-0.6b, Qwen3-VL-Reranker-2B and Qwen3-VL-Reranker-8B use Apache-2.0. [9][2][4] jina-reranker-v3.5 uses cc-by-nc-4.0. [7] NVIDIA’s llama-nemotron-rerank-vl-1b-v2 uses nvidia-open-model-license, while llama-nemotron-rerank-1b-v2 uses openmdw-1.1. [5][1] Check the exact license attached to your chosen repository before adopting it. Treat licensing as a separate selection decision from publication date, download count, context length and estimated weight storage.

Which models offer enough context for long documents?

Qwen3-VL-Reranker-2B and Qwen3-VL-Reranker-8B each list a context length of 256K tokens. [2][4] jina-reranker-v3.5, llama-nemotron-rerank-vl-1b-v2 and llama-nemotron-rerank-1b-v2 list 128K tokens. [7][5][1] R3-rerank-0.6b lists 40K tokens. [9] Use those limits to screen candidates against your intended inputs. A listed context limit alone does not establish whether processing that input length will meet your CPU latency or memory requirements.

Do published benchmarks show which model runs well on a CPU?

Every candidate has no public benchmark yet in the eligible comparison. [7][9][2][4][5][1] No benchmark score therefore supports a performance ordering among them. The ranking uses release recency and download ties; it should not be read as measured CPU speed or reranking quality. For deployment, measure latency and memory with your own queries, documents and chosen runtime before committing.

Sources

  1. nvidia/llama-nemotron-rerank-1b-v2 model card (Hugging Face) — 2026-09-25
  2. Qwen/Qwen3-VL-Reranker-2B model card (Hugging Face) — 2026-09-25
  3. Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — 2026-01-08
  4. Qwen/Qwen3-VL-Reranker-8B model card (Hugging Face) — 2026-09-25
  5. nvidia/llama-nemotron-rerank-vl-1b-v2 model card (Hugging Face) — 2026-09-25
  6. Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models — 2025-01-20
  7. jinaai/jina-reranker-v3.5 model card (Hugging Face) — 2026-09-25
  8. jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation — 2026-07-20
  9. tencent/R3-rerank-0.6b model card (Hugging Face) — 2026-09-25
  10. Skill Is Not Document: A Query-Conditional Benchmark and Two-Stage Retriever for LLM Agent Skill Routing — 2026-06-14

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog