Best Reranker Models to Run Locally in 2026

Rankings 2026-09-25 Last updated 2026-09-25 13 min read By Q4KM

Quick Answer

jina-reranker-v3.5 (jinaai) is the top pick as of September 2026 for its July 2026 release, 597M parameters and 128K-token context, with no public benchmark yet [7]. In order, the ranking is jina-reranker-v3.5 (jinaai), R3-rerank-0.6b (tencent), Qwen3-VL-Reranker-2B (Qwen), Qwen3-VL-Reranker-8B (Qwen), llama-nemotron-rerank-vl-1b-v2 (nvidia), zerank-2-reranker (zeroentropy), and llama-nemotron-rerank-1b-v2 (nvidia) [7][9].

Key Takeaways

How do local reranker models compare on parameters, memory, context, licenses, and benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
jina-reranker-v3.5 [7] jinaai [7] 597M [7] BF16 weights: 1.2 GB [7]; quantized VRAM: not published 2026-07-14 [7]. Release recency puts it first [1][2][4][5][7][9][11]. cc-by-nc-4.0 [7] no public benchmark yet
R3-rerank-0.6b [9] tencent [9] 596M [9] F32 weights: 2.4 GB [9]; quantized VRAM: not published 2026-07-08 [9] Apache-2.0 [9] no public benchmark yet
Qwen3-VL-Reranker-2B [2] Qwen [2] 2.1B [2] BF16 weights: 4.3 GB [2]; quantized VRAM: not published 2026-01-07 [2] Apache-2.0 [2] no public benchmark yet
Qwen3-VL-Reranker-8B [4] Qwen [4] 8.8B [4] BF16 weights: 17.5 GB [4]; quantized VRAM: not published 2026-01-07 [4] Apache-2.0 [4] no public benchmark yet
llama-nemotron-rerank-vl-1b-v2 [5] nvidia [5] 1.7B [5] BF16 weights: 3.4 GB [5]; quantized VRAM: not published 2025-12-04 [5] nvidia-open-model-license [5] no public benchmark yet
zerank-2-reranker [11] zeroentropy [11] 4B [11] BF16 weights: 8.0 GB [11]; quantized VRAM: not published 2025-11-19 [11] Apache-2.0 [11] no public benchmark yet
llama-nemotron-rerank-1b-v2 [1] nvidia [1] 1.2B [1] BF16 weights: 2.5 GB [1]; quantized VRAM: not published 2025-10-16 [1] openmdw-1.1 [1] no public benchmark yet

Which reranker models should you run locally?

1. jina-reranker-v3.5

jina-reranker-v3.5 by jinaai ranks first here for its recent release and compact parameter count, rather than a demonstrated benchmark lead: it was published on July 14, 2026, and has 597M parameters.[7] Its benchmark status is no public benchmark yet, so treat the placement as a recommendation to evaluate it, not a proven quality advantage.

The model supports a 128K-token context and ships with 1.2 GB of BF16 weights.[7] Hardware sizing remains an open question: that download size does not establish a minimum system RAM or GPU VRAM requirement. Memory use at 4-bit precision is not specified, so a concrete quantized hardware recommendation would be premature.[7]

Consider it for local listwise reranking, where its hybrid-attention and self-distillation approach is relevant to evaluating retrieved candidates together.[8] The practical caveat is its cc-by-nc-4.0 license; check compatibility with your intended deployment before adopting it.[7]

2. R3-rerank-0.6b

R3-rerank-0.6b by tencent ranks here as a recent, compact option with an Apache-2.0 license: published on July 8, 2026, it has 596M parameters and a 40K-token context window.[9] Its placement reflects those practical attributes; ranking quality remains unverified, with no public benchmark yet.

For local hardware planning, the published F32 weights occupy 2.4 GB.[9] Weight size alone does not establish the RAM or VRAM needed for inference. A measured 4-bit memory requirement is not specified,[9] so a particular GPU or minimum-memory configuration cannot be recommended confidently. Check runtime memory on your intended workload before committing hardware.

Consider it for evaluating query-dependent routing to agent skills, the application described in its accompanying paper, “Skill Is Not Document: A Query-Conditional Benchmark and Two-Stage Retriever for LLM Agent Skill Routing.”[10] The caveat is the missing public benchmark: validate relevance on your own queries before selecting it for deployment.

3. Qwen3-VL-Reranker-2B

Qwen3-VL-Reranker-2B by Qwen ranks here as a multimodal reranking option with a smaller weight footprint than Qwen3-VL-Reranker-8B by Qwen: published BF16 weights occupy 4.3 GB versus 17.5 GB.[2][4] Consider it for local multimodal retrieval and ranking when weight memory influences model selection.[3] Benchmark status: no public benchmark yet, so its placement does not establish a measured quality advantage.

The model has 2.1 billion parameters, supports a 256K-token context, and uses the Apache-2.0 license.[2] First published on Hugging Face on January 7, 2026, it falls within the release window for this selection.[2]

For hardware planning, the published 4.3 GB BF16 weight size is a starting point, not a verified total runtime memory requirement.[2] The caveat is hardware certainty: neither a measured 4-bit memory requirement nor a validated local hardware configuration is specified here. Choose hardware only after checking the intended runtime’s requirements.

4. Qwen3-VL-Reranker-8B

Qwen3-VL-Reranker-8B by Qwen belongs here as an option for local multimodal retrieval and ranking, with no public benchmark yet to establish a performance advantage over the other candidates.[3][4] Its position should be read as a shortlist placement, not a demonstrated benchmark result.

Published on January 7, 2026, the model has 8.8 billion parameters, a context length of 256K tokens, and BF16 weights totaling 17.5 GB.[4] Qwen distributes it under the Apache-2.0 license.[4] Multimodal retrieval pipelines are its intended use: consider it when your application needs ranking across text and visual content.[3]

For local hardware planning, budget memory around the published 17.5 GB weight payload, with additional room for execution.[4] Choose hardware after checking your runtime’s actual memory demand. The caveat is capacity uncertainty: quantized memory use and total runtime memory are unspecified, so the weight size alone cannot establish whether a particular machine will fit the model.

5. llama-nemotron-rerank-vl-1b-v2

llama-nemotron-rerank-vl-1b-v2 by nvidia occupies the fifth position provisionally: no public benchmark yet establishes its quality relative to the other candidates.[5] Its practical appeal is the combination of a 128K-token context window and 1.7B parameters.[5] The model was first published on Hugging Face on December 4, 2025.[5]

For local hardware planning, the published BF16 weights occupy 3.4 GB.[5] Budget memory for those weights plus runtime overhead; the weight-file size alone does not establish a sufficient GPU capacity. A measured 4-bit memory requirement is not specified, so a precise quantized hardware recommendation remains unverified.[5] Choose hardware after measuring memory use with your intended input lengths.

Consider the model for local reranking workflows involving long inputs, where its published context limit is relevant.[5] The caveat is licensing: distribution uses the nvidia-open-model-license, so review those terms against your intended deployment before adopting it.[5]

6. zerank-2-reranker

zerank-2-reranker by zeroentropy holds a provisional place in this shortlist: no public benchmark yet establishes its quality relative to the other candidates. Its Hugging Face publication date is November 19, 2025.[11] Download activity does not establish a performance advantage, so this placement should guide evaluation rather than imply a measured quality ranking.

The model has 4B parameters, a 40K-token context window and 8.0 GB of BF16 weights, with an Apache-2.0 license.[11] For local deployment, choose hardware with memory headroom beyond the weight footprint. The published weight size alone does not establish total RAM or VRAM needs; quantized runtime memory remains unverified.

Consider it for local reranking when its license and context window fit your application.[11] The practical caveat is evaluation uncertainty: validate relevance and memory consumption on your own workload before committing hardware.

7. llama-nemotron-rerank-1b-v2

llama-nemotron-rerank-1b-v2 by nvidia is a local reranker with 1.2B parameters, first published on Hugging Face on October 16, 2025.[1] Its placement here is provisional: no public benchmark yet establishes its standing against the other candidates. Choose it for its deployment characteristics, with retrieval quality still requiring evaluation on your own documents.

The model supports a context length of 128K tokens and ships with 2.5 GB of BF16 weights.[1] For local hardware planning, allow memory for those weights plus runtime overhead. The weight download size alone does not establish a complete RAM or VRAM requirement, and a verified quantized memory requirement is unavailable.

Consider it for reranking workflows that need room for long inputs within that context limit.[1] The practical caveat is licensing: the model uses openmdw-1.1, so check its terms against your intended deployment before adoption.[1]

How much memory do quantized rerankers need?

Quantized rerankers need roughly 0.30–4.4 GB for raw four-bit weights across these candidates, calculated from their published parameter counts; total runtime memory is additional.[1][2][4][5][7][9][11] The calculation assumes every parameter occupies four bits, or half a byte, and uses decimal gigabytes. Treat the results as theoretical weight payloads, not measured RAM or VRAM requirements.

Tencent’s R3-rerank-0.6b (tencent) and jinaai’s jina-reranker-v3.5 (jinaai) each work out to approximately 0.30 GB for weights alone.[9][7] NVIDIA’s llama-nemotron-rerank-1b-v2 (nvidia) comes to 0.60 GB, while NVIDIA’s llama-nemotron-rerank-vl-1b-v2 (nvidia) comes to 0.85 GB under the same assumption.[1][5]

Qwen’s Qwen3-VL-Reranker-2B (Qwen) works out to 1.05 GB, zeroentropy’s zerank-2-reranker (zeroentropy) to 2.0 GB, and Qwen’s Qwen3-VL-Reranker-8B (Qwen) to 4.4 GB.[2][11][4] Published checkpoint sizes use BF16 for these models, so those download sizes describe a different weight representation.[2][11][4]

Leave room beyond the calculated payload for quantization metadata, parameters retained at higher precision, and runtime allocations. A parameter-count calculation cannot establish whether a model will fit your hardware during inference. Validate peak memory with your intended input lengths and batch size before committing to deployment. The estimates also do not establish that a compatible quantized checkpoint or loader is available.

Which local rerankers use the Apache-2.0 license ?

Qwen’s Qwen3-VL-Reranker-2B [2] and Qwen3-VL-Reranker-8B [4], tencent’s R3-rerank-0.6b [9], and zeroentropy’s zerank-2-reranker [11] use the Apache-2.0 license.

Qwen3-VL-Reranker-2B has 2.1B parameters, a context length of 256K tokens, and BF16 weights totaling 4.3 GB.[2] The model was first published on Hugging Face on January 7, 2026.[2] Within the Apache-licensed Qwen pair, it offers the same stated context length with a smaller weight download.[2][4]

Qwen3-VL-Reranker-8B has 8.8B parameters and the same 256K-token context length, with BF16 weights totaling 17.5 GB.[4] Its first Hugging Face publication date was also January 7, 2026.[4] The shared license and context length therefore come with substantially different parameter counts and weight sizes.[2][4]

R3-rerank-0.6b has 596M parameters and a context length of 40K tokens.[9] Its published weights total 2.4 GB in F32 format, and its first Hugging Face publication date was July 8, 2026.[9] Note the weight format when comparing its download size with the BF16 releases.

zerank-2-reranker has 4B parameters, a context length of 40K tokens, and BF16 weights totaling 8.0 GB.[11] The model was first published on Hugging Face on November 19, 2025.[11] Its stated context length matches R3-rerank-0.6b’s, while its parameter count and published weight size are larger.[9][11]

How should you choose a reranker without a public benchmark?

Choose a reranker without a public benchmark by checking its license and hardware fit, then evaluating relevance and latency on queries from your own application. Treat “no public benchmark yet” as an evidence gap, not a quality verdict.

Start with deployment constraints. Tencent’s R3-rerank-0.6b (tencent) has 596M parameters, a 40K-token context window and an Apache-2.0 license.[9] Jinaai’s jina-reranker-v3.5 (jinaai) has 597M parameters and a 128K-token context window, but its cc-by-nc-4.0 license requires attention if you plan commercial use.[7] Choose a license compatible with your deployment before spending time on evaluation.

Check memory with the runtime and workload you intend to deploy. Published weight sizes describe stored weights; use measured runtime memory for capacity planning. Test your intended quantization, document lengths and batch settings, and record relevance alongside latency and memory consumption. Avoid assuming that a supported context window guarantees practical throughput at that length.

Build an evaluation set from representative queries and relevance judgments. Give each candidate the same retrieved documents, then inspect ordering errors and whether reranking improves the passages passed to your generator. Label model-card benchmark results “self-reported.” Use leaderboard scores only to compare models evaluated together, and never place an older scored model above a newer unscored candidate on that basis. Downloads can break a tie; they cannot establish retrieval quality.

Frequently Asked Questions

Which local reranker should I evaluate first?

jina-reranker-v3.5 (jinaai) is our first evaluation pick because it combines a July 14, 2026 release, 597M parameters and a 128K-token context window [7]. Benchmark status: no public benchmark yet. Treat that recommendation as a starting point for evaluation, rather than a measured quality ranking. Check its cc-by-nc-4.0 license against your intended use before adopting it [7].

Which small reranker should I consider for an Apache-licensed deployment?

R3-rerank-0.6b (tencent) has 596M parameters, a 40K-token context window and an Apache-2.0 license [9]. Its original F32 weights occupy 2.4 GB, so parameter count alone should not determine your memory budget [9]. NVIDIA’s llama-nemotron-rerank-1b-v2 (nvidia) offers 1.2B parameters and a 128K-token context window under openmdw-1.1, making it a separate licensing choice [1]. Evaluate retrieval quality on your own queries before committing.

How much memory should I budget for quantized rerankers?

A verified quantized runtime memory figure is not available here. Original weight sizes provide a reference: jina-reranker-v3.5 (jinaai) has 1.2 GB of BF16 weights [7], while zerank-2-reranker (zeroentropy) has 8.0 GB of BF16 weights [11]. Neither figure establishes quantized runtime memory consumption. Measure memory with your intended quantization, input lengths and workload before choosing deployment hardware.

Which rerankers should I consider for long inputs?

Qwen3-VL-Reranker-2B (Qwen) and Qwen3-VL-Reranker-8B (Qwen) each advertise a 256K-token context window [2][4]. NVIDIA’s llama-nemotron-rerank-vl-1b-v2 (nvidia) advertises 128K tokens [5], while zerank-2-reranker (zeroentropy) advertises 40K tokens [11]. Use those limits to shortlist candidates for your document lengths. Evaluate ranking quality and memory at the lengths you actually expect to process before selecting a model.

How do the Qwen vision-language reranker sizes compare?

Qwen3-VL-Reranker-2B (Qwen) has 2.1B parameters and 4.3 GB of BF16 weights [2]. Qwen3-VL-Reranker-8B (Qwen) has 8.8B parameters and 17.5 GB of BF16 weights [4]. Both were published on January 7, 2026, with Apache-2.0 licensing and 256K-token context windows [2][4]. Choose an evaluation candidate around your memory budget; the larger parameter count does not establish a quality advantage. Benchmark status for both: no public benchmark yet.

Can download counts tell me which reranker is better?

Download counts describe adoption, not ranking quality. As of September 25, 2026, llama-nemotron-rerank-1b-v2 (nvidia) recorded 920,905 downloads over the preceding 30 days [1], compared with 648,100 for zerank-2-reranker (zeroentropy) [11]. Use popularity only as a tiebreak after release eligibility and comparable benchmark results. Label model-card benchmarks self-reported, and leave unscored candidates marked no public benchmark yet rather than inferring performance from adoption.

Sources

  1. nvidia/llama-nemotron-rerank-1b-v2 model card (Hugging Face) — 2026-09-25
  2. Qwen/Qwen3-VL-Reranker-2B model card (Hugging Face) — 2026-09-25
  3. Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking — 2026-01-08
  4. Qwen/Qwen3-VL-Reranker-8B model card (Hugging Face) — 2026-09-25
  5. nvidia/llama-nemotron-rerank-vl-1b-v2 model card (Hugging Face) — 2026-09-25
  6. Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models — 2025-01-20
  7. jinaai/jina-reranker-v3.5 model card (Hugging Face) — 2026-09-25
  8. jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation — 2026-07-20
  9. tencent/R3-rerank-0.6b model card (Hugging Face) — 2026-09-25
  10. Skill Is Not Document: A Query-Conditional Benchmark and Two-Stage Retriever for LLM Agent Skill Routing — 2026-06-14
  11. zeroentropy/zerank-2-reranker model card (Hugging Face) — 2026-09-25
  12. zELO: ELO-inspired Training Method for Rerankers and Embedding Models — 2025-09-16

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog