Quick Answer
jina-reranker-v3.5 (jinaai) is the top pick as of September 2026 because its July release places it first under the recency-based ordering [9]. In order, the ranking is jina-reranker-v3.5 (jinaai), R3-rerank-0.6b (tencent), llama-nemotron-rerank-vl-1b-v2 (nvidia), zerank-2-reranker (zeroentropy), llama-nemotron-rerank-1b-v2 (nvidia), Qwen3-Reranker-4B (Qwen), Qwen3-Reranker-0.6B (Qwen), and gte-reranker-modernbert-base (Alibaba-NLP) [9][11].
Key Takeaways
- jinaai’s jina-reranker-v3.5 takes the first position because its July 14, 2026 release is the newest eligible release; its position does not establish benchmark superiority.[9][11]
- Ranking note: newest release first, then downloads; every candidate has “no public benchmark yet.”[1][3][5][6][7][9][11][13] Scope admits only text-ranking rerankers from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days; embedding models are excluded.
- tencent’s R3-rerank-0.6b combines an Apache-2.0 license with 596M parameters, a 40K-token context and 2.4 GB of F32 weights—a concrete starting point for local deployment planning.[11]
- nvidia’s llama-nemotron-rerank-vl-1b-v2 and llama-nemotron-rerank-1b-v2 both list 128K-token contexts, with BF16 weights of 3.4 GB and 2.5 GB respectively; their licenses differ: nvidia-open-model-license for the former and openmdw-1.1 for the latter.[7][6]
- zeroentropy’s zerank-2-reranker and Qwen’s Qwen3-Reranker-4B each list 4B parameters, 40K-token contexts, 8.0 GB of BF16 weights and Apache-2.0 licensing.[13][3] jina-reranker-v3.5 instead lists cc-by-nc-4.0, making license review a practical selection criterion.[9]
- Qwen3-Reranker-4B, Qwen’s Qwen3-Reranker-0.6B and Alibaba-NLP’s gte-reranker-modernbert-base are established picks admitted through the family popularity exemption; that exemption does not grant the first position.[3][5][1] Their ordering follows release date, rather than parameter count or download volume.[3][5][1]
How do local reranker models compare on specs and benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| jina-reranker-v3.5 [9] | jinaai [9] | 597M [9] | BF16 weights: 1.2 GB [9]; runtime VRAM: not published | 2026-07-14 [9] | cc-by-nc-4.0 [9] | no public benchmark yet |
| R3-rerank-0.6b [11] | tencent [11] | 596M [11] | F32 weights: 2.4 GB [11]; runtime VRAM: not published | 2026-07-08 [11] | Apache-2.0 [11] | no public benchmark yet |
| llama-nemotron-rerank-vl-1b-v2 [7] | nvidia [7] | 1.7B [7] | BF16 weights: 3.4 GB [7]; runtime VRAM: not published | 2025-12-04 [7] | nvidia-open-model-license [7] | no public benchmark yet |
| zerank-2-reranker [13] | zeroentropy [13] | 4B [13] | BF16 weights: 8.0 GB [13]; runtime VRAM: not published | 2025-11-19 [13] | Apache-2.0 [13] | no public benchmark yet |
| llama-nemotron-rerank-1b-v2 [6] | nvidia [6] | 1.2B [6] | BF16 weights: 2.5 GB [6]; runtime VRAM: not published | 2025-10-16 [6] | openmdw-1.1 [6] | no public benchmark yet |
| Qwen3-Reranker-4B — established pick [3] | Qwen [3] | 4B [3] | BF16 weights: 8.0 GB [3]; runtime VRAM: not published | 2025-06-03 [3] | Apache-2.0 [3] | no public benchmark yet |
| Qwen3-Reranker-0.6B — established pick [5] | Qwen [5] | 596M [5] | BF16 weights: 1.2 GB [5]; runtime VRAM: not published | 2025-05-29 [5] | Apache-2.0 [5] | no public benchmark yet |
| gte-reranker-modernbert-base — established pick [1] | Alibaba-NLP [1] | 150M [1] | F32 weights: 0.6 GB [1]; runtime VRAM: not published | 2025-01-20 [1] | Apache-2.0 [1] | no public benchmark yet |
Which reranker models should you consider for a local RAG pipeline?
1. jina-reranker-v3.5
jina-reranker-v3.5 by jinaai ranks first under the newest-release-first ordering, following its July 14, 2026 publication on Hugging Face [9]. Benchmark status: no public benchmark yet. Placement reflects release recency, not a demonstrated retrieval-quality advantage. The model has 597M parameters and a context length of 128K tokens [9].
For local hardware planning, the published BF16 weights occupy 1.2 GB [9]. Treat that as a weights-only baseline, not a total GPU VRAM or system RAM requirement. Measure peak memory with your intended input lengths and batching before choosing hardware; the weight footprint alone does not establish a reliable fit.
Consider it for RAG pipelines that need listwise reranking of long candidate passages: its design uses hybrid attention and self-distillation [10]. The deployment caveat is its Creative Commons Attribution-NonCommercial 4.0 license, which restricts commercial use [9].
2. R3-rerank-0.6b
R3-rerank-0.6b by tencent ranks second under the ordering rule of newest release first, then downloads: its Hugging Face publication date is July 8, 2026, behind jinaai’s jina-reranker-v3.5, published July 14, 2026.[11][9] Its position reflects release timing; no public benchmark yet establishes its retrieval quality against the other candidates.
The model has 596M parameters, a 40K-token context window, and F32 weights occupying 2.4 GB.[11] For local hardware planning, use that weight footprint as a starting point, not a verified total RAM or VRAM requirement. A specific GPU recommendation would require a runtime memory measurement that is not available here.
Consider it for a local retrieval-augmented generation pipeline where Apache-2.0 licensing and a 40K-token context match your requirements.[11] The benchmark gap is the caveat: evaluate relevance on your own queries before choosing it over another reranker.
3. llama-nemotron-rerank-vl-1b-v2
llama-nemotron-rerank-vl-1b-v2 by nvidia ranks third under the ordering rule: newest release first, then downloads.[7][9][11] Its first Hugging Face publication was 2025-12-04, after the jinaai and tencent entries and before zeroentropy’s entry.[7][9][11][13] The benchmark status is “no public benchmark yet”; its position does not establish a retrieval-quality advantage.
The model has 1.7B parameters, a 128K-token context window and 3.4 GB of BF16 weights.[7] Distribution uses the nvidia-open-model-license.[7] Consider it for local RAG evaluations where a long context window is a selection requirement: the documented allowance is 128K tokens.[7] Evaluate relevance on your own queries before choosing it for production.
For local hardware planning, start with the 3.4 GB BF16 weight footprint.[7] Treat that figure as weight storage, not a complete runtime memory requirement. A specific GPU or system-RAM fit remains an estimate; a measured hardware requirement is unavailable. The practical caveat is that the documented context allowance does not establish memory use or throughput at that length.
4. zerank-2-reranker
zerank-2-reranker by zeroentropy ranks fourth under the ordering “newest release first, then downloads,” with a Hugging Face publication date of November 19, 2025.[13] Its position reflects release recency; benchmark status is “no public benchmark yet,” so the ranking does not establish comparative RAG quality.
The model has 4B parameters, a 40K-token context window, and published BF16 weights occupying 8.0 GB.[13] For local hardware planning, quantized weight storage would be 2.0 GB at 4-bit precision, estimated (params x 0.5 bytes) from the parameter count.[13] Neither weight figure establishes a complete GPU-memory or system-RAM requirement, so a specific hardware fit remains unverified.
Consider it for a local RAG pipeline whose selection criteria include a 40K-token context window and the Apache-2.0 license.[13] The practical caveat is unverified comparative quality: evaluate relevance on your own queries and retrieved documents before adopting it.
5. llama-nemotron-rerank-1b-v2
llama-nemotron-rerank-1b-v2 by nvidia ranks fifth under the ordering rule of newest release first, then downloads, with a Hugging Face publication date of October 16, 2025.[6] The position reflects release timing rather than demonstrated retrieval quality: no public benchmark yet.[6] Treat it as a candidate to evaluate on your RAG workload, rather than a proven quality winner.
The model has 1.2 billion parameters, a 128K-token context window, and 2.5 GB of BF16 weights.[6] For local hardware planning, use that weight footprint as a starting point. The listed weight size does not establish a complete GPU VRAM or system RAM requirement, so a specific GPU fit cannot be stated here.
Consider it for RAG pipelines where long query-and-document inputs make its 128K-token context relevant.[6] Check relevance quality and latency on your own documents before deployment. A deployment caveat is licensing: the model uses openmdw-1.1, which should be reviewed against your intended use.[6]
6. Qwen3-Reranker-4B
Qwen3-Reranker-4B by Qwen ranks here as an established pick, with a Hugging Face publication date of June 3, 2025.[3] Its exemption from the recency gate allows inclusion; ordering follows release date, then downloads. Benchmark status: no public benchmark yet. Its position therefore reflects the ordering rule rather than a demonstrated retrieval-quality advantage.
Qwen lists 4B parameters, a 40K-token context window and an Apache-2.0 license.[3] Consider it for local RAG pipelines where long candidate passages and that license are selection priorities. The context specification provides a reason to evaluate it on your documents, but does not establish ranking accuracy.
Local hardware planning starts with an 8.0 GB BF16 weight footprint.[3] Allow additional memory for inference; weight storage alone does not establish a total RAM or VRAM requirement. The practical caveat is that a GPU-fit recommendation remains unverified: measure memory use with your intended passage lengths and batching before committing hardware.
7. Qwen3-Reranker-0.6B
Qwen3-Reranker-0.6B by Qwen is an established pick, first published on Hugging Face on May 29, 2025.[5] Its inclusion bypasses the recency gate because it is among its family’s three most-downloaded models.[5] The ordering is newest release first, then downloads; its position does not establish a retrieval-quality advantage.
The model has 596M parameters, a 40K-token context length, and BF16 weights occupying 1.2 GB.[5] For local hardware planning, treat the published weight size as a starting point for RAM or VRAM allocation, with additional memory needed for inference. A specific GPU fit cannot be established from the weight size alone.
Consider it for local RAG pipelines that need long query–passage inputs and Apache-2.0 licensing.[5] The caveat is evaluation: no public benchmark yet under this ranking’s criteria. Validate relevance and latency on your own candidate passages before choosing it for deployment; context capacity alone does not establish reranking quality.
8. gte-reranker-modernbert-base
gte-reranker-modernbert-base by Alibaba-NLP is an established pick that bypasses the recency gate and ranks last under “newest release first, then downloads,” with a first Hugging Face publication date of January 20, 2025.[1] Its established status reflects its place among the family’s three most-downloaded models.[1] Benchmark status: no public benchmark yet. Treat its position as release ordering rather than a measured quality comparison.
The model has 150M parameters, an 8K-token context window, and 0.6 GB of F32 weights, under the Apache-2.0 license.[1] Consider it for a local RAG pipeline whose query and candidate text fit that context limit.[1] For hardware planning, use the published weight footprint as a starting point; measure runtime RAM or VRAM consumption before committing to a machine. The practical caveat is that weight storage does not establish total inference memory, so a specific hardware fit remains unverified.
How can you estimate memory requirements for a local reranker?
Estimate memory requirements by starting with the published weight size, then budgeting additional memory for inference at your intended input length and batch size; treat weight storage as a baseline, not a hardware fit guarantee.
Qwen’s Qwen3-Reranker-4B has 4B parameters and published BF16 weights occupying 8.0 GB.[3] A hypothetical 4-bit version would require approximately 2.0 GB for weights alone, estimated (params x 0.5 bytes).[3] That calculation does not establish that a compatible quantized release exists or account for the memory needed to run inference.
Alibaba-NLP’s gte-reranker-modernbert-base has 150M parameters and published F32 weights occupying 0.6 GB.[1] Compare the actual weight formats you intend to load: a published file size and a calculated quantized weight size describe different deployment assumptions. Neither figure should become a GPU recommendation without checking runtime memory use.
For a practical capacity check, load the chosen model with your intended inference software and measure peak memory while reranking representative query–document inputs. Test with input lengths and batch sizes your application will use, then leave memory headroom before choosing hardware. Record the weight format, input length and batch size alongside the measurement so the estimate remains useful when your pipeline changes.
Which reranker licenses suit your deployment?
Choose the Apache-2.0 candidates when your deployment policy requires that license; jinaai’s and NVIDIA’s candidates need separate license review.[1][3][5][6][7][9][11][13] Make license compatibility a selection requirement before comparing model quality or hardware needs.
Alibaba-NLP’s gte-reranker-modernbert-base uses Apache-2.0.[1] Qwen’s Qwen3-Reranker-4B and Qwen3-Reranker-0.6B carry the same license.[3][5] Tencent’s R3-rerank-0.6b and ZeroEntropy’s zerank-2-reranker also use Apache-2.0.[11][13] Those candidates give an Apache-only shortlist across several publishers. Check the license obligations against your intended distribution and deployment before shipping.
Jina AI’s jina-reranker-v3.5 carries cc-by-nc-4.0.[9] Treat its noncommercial restriction as a deployment constraint: confirm permission for your intended use before selecting it for a commercial service. Running the weights locally should not substitute for that review.
NVIDIA’s llama-nemotron-rerank-1b-v2 uses openmdw-1.1, while NVIDIA’s llama-nemotron-rerank-vl-1b-v2 uses nvidia-open-model-license.[6][7] Review each model’s terms individually. The shared publisher and similar names should not replace checking the license attached to the exact repository.
Document the selected model, its license, and the planned use in your deployment review. Include whether you will operate an internal service, expose an API, or redistribute weights. Recheck that decision whenever you switch models, including switches within a publisher’s reranker family.
How should you choose a reranker with no public benchmark yet?
Choose a reranker with no public benchmark yet by evaluating it on your own retrieval tasks, then checking its license and local resource requirements before deployment. Treat “no public benchmark yet” as an evidence gap, not a quality rating. Label any model-card benchmark “self-reported” and keep it separate from your own evaluation.
Build an evaluation set from representative queries, retrieved passages and relevance judgments. Include ambiguous queries, near-duplicate passages and cases where retrieval misses the answer. Compare reranked results with your existing retrieval order. Keep the candidate passages fixed so you can attribute changes to reranking. Record relevance, end-to-end latency and peak memory on your intended hardware.
Use published weight sizes as planning inputs, not deployment memory requirements. Qwen’s Qwen3-Reranker-4B lists BF16 weights of 8.0 GB [3], while Alibaba-NLP’s gte-reranker-modernbert-base lists F32 weights of 0.6 GB [1]. Measure memory under your intended passage lengths and batch settings before selecting hardware.
Check license compatibility early: jinaai’s jina-reranker-v3.5 lists cc-by-nc-4.0 [9], while tencent’s R3-rerank-0.6b lists Apache-2.0 [11]. Set acceptance criteria before testing, and choose based on measured retrieval quality within your deployment constraints. Release dates and download counts can organize an evaluation queue; neither substitutes for results on your workload.
Frequently Asked Questions
Which local reranker should I evaluate first?
jinaai’s jina-reranker-v3.5 (jinaai) takes first place under the release-date rule, with a Hugging Face publication date of 2026-07-14.[9] Its benchmark status is “no public benchmark yet,” so the placement does not establish superior retrieval quality. The model has 597M parameters, a 128K-token context window and 1.2 GB of BF16 weights.[9] Check its cc-by-nc-4.0 license before choosing it for your deployment.[9]
How are the rerankers ordered?
Ranking note: no public benchmark scores any candidate, so the order is newest release first, then downloads: jina-reranker-v3.5 (jinaai)[9], Tencent’s R3-rerank-0.6b (tencent)[11], NVIDIA’s llama-nemotron-rerank-vl-1b-v2 (nvidia)[7], ZeroEntropy’s zerank-2-reranker (zeroentropy)[13], NVIDIA’s llama-nemotron-rerank-1b-v2 (nvidia)[6], Qwen’s Qwen3-Reranker-4B (Qwen)[3], Qwen’s Qwen3-Reranker-0.6B (Qwen)[5], Alibaba-NLP’s gte-reranker-modernbert-base (Alibaba-NLP)[1]. Scope: admission requires a lab with a published paper or leaderboard record, or another publisher above 10,000 downloads in the last 30 days[13], using Hugging Face’s text-ranking tag; embedding models are excluded.
Why include older Qwen and Alibaba-NLP rerankers?
Qwen3-Reranker-4B (Qwen)[3], Qwen3-Reranker-0.6B (Qwen)[5] and gte-reranker-modernbert-base (Alibaba-NLP)[1] are established picks: each belongs to its family’s three most-downloaded models.[3][5][1] Established picks bypass the twelve-month recency gate but follow the same ordering rule; the exemption cannot put them first. Their benchmark status remains “no public benchmark yet.” Download popularity does not establish retrieval quality.
How much memory do I need to run these rerankers locally?
Use the published weight sizes as a starting point, rather than a hardware-fit guarantee. gte-reranker-modernbert-base (Alibaba-NLP) has 0.6 GB of F32 weights.[1] Qwen3-Reranker-0.6B (Qwen) and jina-reranker-v3.5 (jinaai) each have 1.2 GB of BF16 weights.[5][9] Qwen3-Reranker-4B (Qwen) and zerank-2-reranker (zeroentropy) each have 8.0 GB of BF16 weights.[3][13] Those figures describe weights; they do not establish total runtime memory or confirm that a particular GPU can run your workload.
Which rerankers have an Apache license?
Tencent’s R3-rerank-0.6b (tencent)[11], ZeroEntropy’s zerank-2-reranker (zeroentropy)[13], both listed Qwen rerankers[3][5] and gte-reranker-modernbert-base (Alibaba-NLP)[1] carry Apache-2.0 licenses.[11][13][3][5][1] The remaining candidates use different licenses: jina-reranker-v3.5 (jinaai) uses cc-by-nc-4.0[9], llama-nemotron-rerank-1b-v2 (nvidia) uses openmdw-1.1[6], and llama-nemotron-rerank-vl-1b-v2 (nvidia) uses nvidia-open-model-license.[7] If Apache-2.0 is a deployment requirement, start your evaluation with the models explicitly carrying that license.[11][13][3][5][1]
Which rerankers support long context windows?
jina-reranker-v3.5 (jinaai) and both listed NVIDIA rerankers specify 128K-token context windows.[9][6][7] Tencent’s R3-rerank-0.6b (tencent), ZeroEntropy’s zerank-2-reranker (zeroentropy) and both listed Qwen rerankers specify 40K tokens.[11][13][3][5] gte-reranker-modernbert-base (Alibaba-NLP) specifies 8K tokens.[1] Treat context length as a compatibility filter for your inputs. A context specification alone does not demonstrate ranking quality, latency or memory consumption on your hardware.
Sources
- Alibaba-NLP/gte-reranker-modernbert-base model card (Hugging Face) — 2026-09-25
- Towards General Text Embeddings with Multi-stage Contrastive Learning — 2023-08-07
- Qwen/Qwen3-Reranker-4B model card (Hugging Face) — 2026-09-25
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models — 2025-06-05
- Qwen/Qwen3-Reranker-0.6B model card (Hugging Face) — 2026-09-25
- nvidia/llama-nemotron-rerank-1b-v2 model card (Hugging Face) — 2026-09-25
- nvidia/llama-nemotron-rerank-vl-1b-v2 model card (Hugging Face) — 2026-09-25
- Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models — 2025-01-20
- jinaai/jina-reranker-v3.5 model card (Hugging Face) — 2026-09-25
- jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation — 2026-07-20
- tencent/R3-rerank-0.6b model card (Hugging Face) — 2026-09-25
- Skill Is Not Document: A Query-Conditional Benchmark and Two-Stage Retriever for LLM Agent Skill Routing — 2026-06-14
- zeroentropy/zerank-2-reranker model card (Hugging Face) — 2026-09-25
- zELO: ELO-inspired Training Method for Rerankers and Embedding Models — 2025-09-16