Quick Answer
Google’s gemma-4-26B-A4B-it is the top pick for local RAG as of September 2026 under the ordering “newest release first, then downloads,” because it shares the newest release date and wins the download tiebreak; all candidates have no public benchmark yet [1][2][3][5][7]. In order, the ranking is gemma-4-26B-A4B-it (Google), gemma-4-31B-it (Google), NVIDIA-Nemotron-3-Nano-4B-BF16 (NVIDIA), Qwen3.5-0.8B (Qwen), and Qwen3.5-9B (Qwen) [1][2].
Key Takeaways
- Ranking note: newest release first, then downloads; all candidates have no public benchmark yet, so ordering does not establish RAG quality.[1][2][3][5][7] Scope admits general chat LLMs with chat templates, tagged text-generation or image-text-to-text, from labs with a published paper or leaderboard record, or other publishers exceeding 10,000 downloads in the last 30 days; any-to-any generators, vision-first models, OCR/grounding/driving models and other specialists are excluded.
- Google’s
gemma-4-26B-A4B-it (Google)ranks first because its March 11, 2026 publication date ties with its sibling, and higher downloads break that tie.[3][7] BF16 weights total 51.6 GB, context is 256K tokens, and licensing is Apache-2.0.[3] - Google’s
gemma-4-31B-it (Google)ranks second, sharing its sibling’s publication date, 256K-token context and Apache-2.0 license, with 62.5 GB of BF16 weights.[3][7] The ordering provides no demonstrated RAG accuracy advantage for either model.[3][7] - NVIDIA’s
NVIDIA-Nemotron-3-Nano-4B-BF16 (NVIDIA)ranks third: BF16 weights total 7.9 GB and context is 256K tokens.[5] Deployment review should account for its NVIDIA Nemotron Open Model License.[5] - Qwen’s
Qwen3.5-0.8B (Qwen)ranks fourth: its February 28, 2026 publication places it ahead of its larger sibling under the recency rule.[1][2] BF16 weights total 1.7 GB, context is 256K tokens, and licensing is Apache-2.0.[1] - Qwen’s
Qwen3.5-9B (Qwen)ranks fifth despite higher downloads than its smaller sibling; its publication date is February 27, 2026.[1][2] BF16 weights total 19.3 GB, context is 256K tokens, and licensing is Apache-2.0.[2]
How do these local RAG models compare on specifications and benchmark availability?
| Model | Org | Params | Quant/VRAM | Context | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| gemma-4-26B-A4B-it [3] | Google [3] | 25.8B [3] | BF16 weights: 51.6 GB [3] | 256K tokens [3] | Apache-2.0 [3] | no public benchmark yet |
| gemma-4-31B-it [7] | Google [7] | 31.3B [7] | BF16 weights: 62.5 GB [7] | 256K tokens [7] | Apache-2.0 [7] | no public benchmark yet |
| NVIDIA-Nemotron-3-Nano-4B-BF16 [5] | NVIDIA [5] | 4B [5] | BF16 weights: 7.9 GB [5] | 256K tokens [5] | nvidia-nemotron-open-model-license [5] | no public benchmark yet |
| Qwen3.5-0.8B [1] | Qwen [1] | 873M [1] | BF16 weights: 1.7 GB [1] | 256K tokens [1] | Apache-2.0 [1] | no public benchmark yet |
| Qwen3.5-9B [2] | Qwen [2] | 9.7B [2] | BF16 weights: 19.3 GB [2] | 256K tokens [2] | Apache-2.0 [2] | no public benchmark yet |
Which open-weight LLMs should you consider for local RAG?
1. gemma-4-26B-A4B-it
gemma-4-26B-A4B-it by Google ranks first because its release date ties the newest candidates, and its download count breaks that tie: published on March 11, 2026, it recorded 11,073,424 downloads over the preceding 30 days as of September 25, 2026.[3][7] For local retrieval-augmented generation (RAG), consider it for document-grounded chat when you want a 256K-token context window and Apache-2.0 licensing.[3] Treat that use as an evaluation target, not a demonstrated advantage in answer quality.
The model contains 25.8B parameters, with BF16 weights totaling 51.6 GB.[3] At 4-bit precision, weight storage would be 12.9 GB, estimated (params x 0.5 bytes) from the cited 25.8B parameters.[3] Hardware planning must distinguish weight storage from total runtime memory; that estimate does not establish a GPU or system RAM requirement. Choose hardware after checking the memory requirements of your intended runtime and workload.
The ranking note uses newest release first, then downloads: no public benchmark yet, so placement does not establish superior RAG performance. Scope admits only general chat LLMs shipping a chat template, tagged text-generation or image-text-to-text, from labs with a published paper or leaderboard record, or other publishers with documented download activity; any-to-any generators, vision-first models, OCR, grounding, driving, and other specialists are excluded. Google qualifies through the Gemma 4 Technical Report.[4]
2. gemma-4-31B-it
gemma-4-31B-it by Google ranks second behind Google’s gemma-4-26B-A4B-it: both were first published on Hugging Face on March 11, 2026, but the former has fewer downloads over the last 30 days—9,263,106 versus 11,073,424 as of September 25, 2026.[7][3] Ordering follows newest release first, then downloads. Its benchmark status is “no public benchmark yet,” so the position does not establish superior retrieval-augmented generation (RAG) quality.
The model has 31.3 billion parameters, a 256K-token context window, and 62.5 GB of BF16 weights under the Apache-2.0 license.[7] For local deployment, treat that weight footprint as a starting point for memory planning. For the cited 31.3 billion parameters, a 4-bit weight footprint is 15.65 GB, estimated (params x 0.5 bytes).[7] Neither figure establishes total GPU memory or system RAM requirements; execution and context also need memory.
Consider it for a local RAG evaluation where a long context window and Apache-2.0 licensing match your deployment requirements.[7] Test it with retrieved passages from your own corpus, checking whether answers stay grounded, attach correct citations, and decline unsupported questions. The caveat is unproven comparative RAG performance: its placement reflects release timing and a download tiebreak, so validate answer quality before committing hardware.
3. NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA-Nemotron-3-Nano-4B-BF16 by NVIDIA[5] ranks third under the ordering rule: newest release first, then downloads.[1][2][3][5][7] Its Hugging Face publication date is March 7, 2026,[5] placing it behind Google’s Gemma releases and ahead of Qwen’s entries.[1][2][3][7] Its benchmark status is “no public benchmark yet,” so the position does not establish superior retrieval-augmented generation performance. Treat it as a candidate to evaluate against your documents and questions.
The model has 4B parameters, a 256K-token context window and 7.9 GB of BF16 weights.[5] For local hardware planning, its 4B parameters[5] imply a weight footprint of 2 GB at 4-bit precision, estimated (params x 0.5 bytes).[5] Treat that calculation as a weight-storage budget, then allow additional memory for execution and context. Choose GPU memory or system RAM capacity around the complete workload; the weight estimate alone is insufficient for a hardware-fit claim.
A practical use is a local RAG pilot where you can measure answer grounding before committing to deployment. Start with a bounded document collection, require supporting passages in answers, and test questions the collection cannot answer. The license caveat is NVIDIA’s nvidia-nemotron-open-model-license.[5] Review its terms against your intended deployment before adopting the model, and make acceptance depend on your own RAG evaluation.
4. Qwen3.5-0.8B
Qwen3.5-0.8B by Qwen ranks fourth under the ordering rule: newest release first, then downloads.[1][2][3][5][7] Published on Hugging Face on February 28, 2026, it follows NVIDIA’s NVIDIA-Nemotron-3-Nano-4B-BF16, published on March 7, 2026, and precedes Qwen’s Qwen3.5-9B, published on February 27, 2026.[1][5][2] The position reflects release timing; RAG quality remains unverified, with no public benchmark yet.
The model has 873 million parameters, a context window of 256K tokens, and BF16 weights totaling 1.7 GB.[1] Quantized weight storage at 4-bit precision is approximately 0.44 GB, estimated (params x 0.5 bytes) from the cited parameter count.[1] For local hardware planning, treat weight storage as a starting budget: leave additional RAM or VRAM for the runtime, working memory, and retrieved context. Avoid treating the weight estimate as a complete machine-memory requirement or a guarantee that a particular device will run the model.
Consider the model for a local RAG prototype where a compact weight footprint matters and you can evaluate answers against your own documents. Start with focused retrieved passages and check whether answers stay grounded before expanding context. Apache-2.0 licensing supports deployment planning.[1] The caveat is evaluation: the advertised context capacity alone does not establish reliable retrieval-grounded answering.
5. Qwen3.5-9B
Qwen3.5-9B by Qwen ranks fifth: its first Hugging Face publication date is February 27, 2026.[2] The ordering basis is newest release first, then downloads, placing it behind the other eligible releases.[1][3][5][7] Benchmark status: no public benchmark yet. Treat its position as a release-order decision rather than evidence of weaker retrieval-augmented generation performance.
The model has 9.7 billion parameters, a 256K-token context window, and 19.3 GB of BF16 weights, with an Apache-2.0 license.[2] For local hardware planning, the quantized weight budget is approximately 4.85 GB at 4-bit precision, estimated (params x 0.5 bytes) from the parameter count.[2] Budget additional memory for runtime overhead and the KV cache; the weight estimate alone does not establish whether a particular GPU or RAM configuration will fit.
Use Qwen3.5-9B as a candidate for a local document assistant when its Apache-2.0 license suits your deployment requirements.[2] Evaluate it with your own retrieved passages, checking whether answers follow the documents, attribute claims correctly, and decline unsupported questions. The caveat is the missing public benchmark: neither the advertised context window nor the release date demonstrates answer accuracy on your RAG workload.
Which models qualify, and why is the ordering newest release first, then downloads?
Google’s Gemma 4 26B A4B IT [3] and Gemma 4 31B IT [7], NVIDIA’s NVIDIA-Nemotron-3-Nano-4B-BF16 [5], and Qwen’s Qwen3.5-0.8B [1] and Qwen3.5-9B [2] qualify, in that order, because the ordering basis is newest release first, then downloads.
Scope admits only general chat LLMs shipping a chat template, tagged text-generation or image-text-to-text on Hugging Face, from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days; any-to-any generators, vision-first models, OCR, grounding, driving and other specialists are excluded. [1][2][3][4][5][6][7]
Each candidate meets the release-window requirement. For this ranking, every candidate carries the status “no public benchmark yet”: no public leaderboard scores these candidates against one another, so the ordering does not establish comparative RAG quality. [1][2][3][5][7] Model-card benchmarks, when discussed, must be labeled self-reported.
Gemma 4 26B A4B IT takes first place because both Google candidates were published on March 11, 2026, and its 11,073,424 downloads exceed its sibling’s 9,263,106 over the preceding 30 days, measured September 25, 2026. [3][7] Downloads resolve that date tie; popularity does not demonstrate retrieval-grounded answer quality.
NVIDIA’s candidate follows with publication on March 7, 2026 [5], then Qwen’s smaller candidate on February 28, 2026 [1], and its larger candidate on February 27, 2026. [2] Parameter count does not determine placement, and no established-pick exemption changes this list.
How much memory should you budget to run these models locally?
Budget for weights plus runtime headroom: BF16 weight sizes range from 1.7 GB for Qwen’s Qwen3.5-0.8B [1] to 62.5 GB for Google’s Gemma 4 31B IT [7]. Treat the weight figures as a starting point, not a complete RAM or VRAM requirement.
Google’s Gemma 4 26B A4B IT has 25.8B parameters and 51.6 GB of BF16 weights [3]. Its quantized weight budget is approximately 12.9 GB at 4-bit, estimated (params x 0.5 bytes) [3]. Gemma 4 31B IT has 31.3B parameters [7], giving approximately 15.65 GB at 4-bit, estimated (params x 0.5 bytes) [7].
NVIDIA’s NVIDIA-Nemotron-3-Nano-4B-BF16 has 4B parameters and 7.9 GB of BF16 weights [5]. Allow approximately 2 GB for 4-bit weights, estimated (params x 0.5 bytes) [5].
Qwen3.5-0.8B has 873M parameters [1], giving approximately 0.44 GB at 4-bit, estimated (params x 0.5 bytes) [1]. Qwen’s Qwen3.5-9B has 9.7B parameters and 19.3 GB of BF16 weights [2]; its 4-bit weight budget is approximately 4.85 GB, estimated (params x 0.5 bytes) [2].
For a local RAG deployment, reserve additional capacity for inference and retrieval, then measure total usage with your intended context and concurrency. Avoid choosing hardware whose entire memory capacity merely matches the estimated weight budget.
Which licenses apply to these models?
The Google and Qwen candidates use Apache-2.0, while NVIDIA’s candidate uses nvidia-nemotron-open-model-license.[1][2][3][5][7] Check the license attached to the exact repository you intend to deploy when selecting a model for local retrieval-augmented generation (RAG).
Google’s gemma-4-26B-A4B-it (Google) and gemma-4-31B-it (Google) both list Apache-2.0.[3][7] Qwen’s Qwen3.5-0.8B (Qwen) and Qwen3.5-9B (Qwen) also list Apache-2.0.[1][2] Those candidates therefore share a license designation. Keep licensing review separate from your evaluation of answer quality, retrieval grounding and deployment requirements.
NVIDIA’s NVIDIA-Nemotron-3-Nano-4B-BF16 (NVIDIA) lists nvidia-nemotron-open-model-license.[5] Review that license independently before approving deployment. Avoid carrying assumptions from an Apache-licensed candidate into your assessment of the NVIDIA model.
For an engineering handoff, record the repository, the selected revision and its accompanying license text. Describe your intended use concretely: internal document search, a customer-facing service, redistribution of weights or distribution of a modified model. Check the applicable terms against that use, and document any obligations before packaging or shipping. Use the specific license name in your deployment inventory rather than treating “open-source” as a sufficient licensing description.
Frequently Asked Questions
Which model should I try first for local RAG?
Google’s gemma-4-26B-A4B-it (Google) leads because its March 11, 2026 release ties for newest among the eligible candidates, with downloads breaking the tie [3][7]. The order is Google’s gemma-4-26B-A4B-it (Google) [3], Google’s gemma-4-31B-it (Google) [7], NVIDIA’s NVIDIA-Nemotron-3-Nano-4B-BF16 (NVIDIA) [5], Qwen’s Qwen3.5-0.8B (Qwen) [1], then Qwen’s Qwen3.5-9B (Qwen) [2]. Treat that order as a shortlist for evaluation rather than evidence of superior RAG accuracy.
How are these models ranked?
Ranking note: no public benchmark yet; newest release first, then downloads [1][2][3][5][7]. The ranking admits only general chat LLMs shipping a chat template from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days [1][2][3][5][7], using Hugging Face pipeline tags text-generation or image-text-to-text; any-to-any generators, vision-first models, OCR, grounding, driving and other specialists are excluded.
How much memory should I budget to run these models locally?
Start with the published BF16 weight sizes: Qwen3.5-0.8B (Qwen) uses 1.7 GB [1], NVIDIA-Nemotron-3-Nano-4B-BF16 (NVIDIA) uses 7.9 GB [5], and Qwen3.5-9B (Qwen) uses 19.3 GB [2]. gemma-4-26B-A4B-it (Google) lists 51.6 GB [3], while gemma-4-31B-it (Google) lists 62.5 GB [7]. Use those figures to plan deployment, then measure total memory with your intended runtime, retrieved documents and concurrency before choosing hardware.
How much could quantization reduce the weight memory?
Qwen3.5-0.8B (Qwen) has 873M parameters [1]: approximately 0.44 GB for 4-bit weights, estimated (params x 0.5 bytes) [1]. gemma-4-26B-A4B-it (Google) has 25.8B parameters [3]: approximately 12.9 GB for 4-bit weights, estimated (params x 0.5 bytes) [3]. Treat those calculations as weight-only estimates. Verify the actual quantized artifact and measure total runtime memory before committing to a GPU or system RAM configuration.
Do all of these models use the same open-source license?
License terms differ. Qwen3.5-0.8B (Qwen) and Qwen3.5-9B (Qwen) use Apache-2.0 [1][2]. gemma-4-26B-A4B-it (Google) and gemma-4-31B-it (Google) also use Apache-2.0 [3][7]. NVIDIA’s NVIDIA-Nemotron-3-Nano-4B-BF16 (NVIDIA) uses nvidia-nemotron-open-model-license [5]. Check the applicable license against your intended deployment and redistribution plans. Avoid treating the phrase “open-weight” as a substitute for reviewing the actual license.
Does a long context window prove a model is good at RAG?
A context specification alone does not establish RAG answer quality. All ranked candidates list a 256K-token context window [1][2][3][5][7], but their ranking benchmark status is “no public benchmark yet” [1][2][3][5][7]. Evaluate each candidate against your own retrieved passages. Check whether answers follow the supplied documents, attach accurate citations, handle conflicting passages and acknowledge when the documents cannot answer the question.
Sources
- Qwen/Qwen3.5-0.8B model card (Hugging Face) — 2026-09-25
- Qwen/Qwen3.5-9B model card (Hugging Face) — 2026-09-25
- google/gemma-4-26B-A4B-it model card (Hugging Face) — 2026-09-25
- Gemma 4 Technical Report — 2026-07-02
- nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 model card (Hugging Face) — 2026-09-25
- Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs — 2025-11-20
- google/gemma-4-31B-it model card (Hugging Face) — 2026-09-25
- meta-llama/Llama-3.2-11B-Vision-Instruct model card (Hugging Face) — 2024-09-18
- moonshotai/Kimi-K2.6 model card (Hugging Face) — 2026-04-14