Best Image Captioning Models to Run Locally in 2026

Rankings 2026-09-26 Last updated 2026-09-26 14 min read By Q4KM

Quick Answer

PP-OCRv6_medium_det_safetensors (PaddlePaddle) is the top pick as of September 2026 because its June release is the newest among the eligible candidates [11]. In order, the ranking is PP-OCRv6_medium_det_safetensors (PaddlePaddle), NuExtract3 (numind), Zero-To-CAD-Qwen3-VL-2B (ADSKAILab), nemotron-ocr-v2 (nvidia), FireRed-OCR (FireRedTeam), Falcon-OCR (tiiuae), GLM-OCR (zai-org), and LightOnOCR-1B-1025 (lightonai) [11][6].

Key Takeaways

How do local image captioning models compare on specifications and benchmark availability?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
PP-OCRv6_medium_det_safetensors [11] PaddlePaddle 22M [11] 4-bit weights: 11 MB, estimated (params x 0.5 bytes) [11]; runtime VRAM not published 2026-06-09 [11] Apache-2.0 [11] no public benchmark yet
NuExtract3 [6] numind 4.5B [6] 4-bit weights: 2.25 GB, estimated (params x 0.5 bytes) [6]; runtime VRAM not published 2026-04-29 [6] Apache-2.0 [6] no public benchmark yet
Zero-To-CAD-Qwen3-VL-2B [13] ADSKAILab 2.1B [13] 4-bit weights: 1.05 GB, estimated (params x 0.5 bytes) [13]; runtime VRAM not published 2026-04-10 [13] Apache-2.0 [13] no public benchmark yet
nemotron-ocr-v2 [3] nvidia not published not published 2026-04-01 [3] nvidia-open-model-license [3] no public benchmark yet
FireRed-OCR [9] FireRedTeam 2.1B [9] 4-bit weights: 1.05 GB, estimated (params x 0.5 bytes) [9]; runtime VRAM not published 2026-02-28 [9] Apache-2.0 [9] no public benchmark yet
Falcon-OCR [4] tiiuae 270M [4] 4-bit weights: 135 MB, estimated (params x 0.5 bytes) [4]; runtime VRAM not published 2026-02-22 [4] Apache-2.0 [4] no public benchmark yet
GLM-OCR [1] zai-org 1.3B [1] 4-bit weights: 650 MB, estimated (params x 0.5 bytes) [1]; runtime VRAM not published 2026-01-30 [1] MIT [1] no public benchmark yet
LightOnOCR-1B-1025 [7] lightonai 1.2B [7] 4-bit weights: 600 MB, estimated (params x 0.5 bytes) [7]; runtime VRAM not published 2025-10-20 [7] Apache-2.0 [7] no public benchmark yet

Which image captioning models can you run locally?

1. PP-OCRv6_medium_det_safetensors

PP-OCRv6_medium_det_safetensors by PaddlePaddle ranks first on release recency, with a Hugging Face publication date of June 9, 2026 [11]. Ranking note: newest release first, then downloads. Eligibility is limited to image to text models with the Hugging Face pipeline tag image-to-text. Benchmark status: no public benchmark yet; the placement does not establish superior caption quality.

The model has 22M parameters and 0.1 GB of F32 weights under the Apache-2.0 license [11]. Quantized weight storage would be approximately 11 MB at 4-bit, estimated (params x 0.5 bytes) from the cited 22M parameters [11]. Treat that calculation as a weights-only estimate, not a complete hardware requirement. Total runtime RAM, GPU memory requirements and context length are unspecified, so a particular local hardware configuration cannot be recommended here.

The practical use case is OCR-focused work, consistent with the family’s technical report [12]. The caveat for image captioning is the absence of a public benchmark establishing descriptive-caption performance.

2. NuExtract3

NuExtract3 by numind ranks second under the ordering “newest release first, then downloads,” with its Hugging Face publication dated 2026-04-29.[6] PaddlePaddle’s PP-OCRv6_medium_det_safetensors has the later publication date of 2026-06-09.[11] NuExtract3 has no public benchmark yet, so its position does not establish superior caption quality.

NuExtract3 has 4.5B parameters, a 256K-token context window and an Apache-2.0 license.[6] Published BF16 weights occupy 9.3 GB.[6] For local hardware planning, its 4-bit weight footprint is approximately 2.25 GB, estimated (params x 0.5 bytes) from the cited 4.5B parameters.[6] Treat that figure as a weight-storage estimate, not a complete runtime memory requirement or confirmation that quantized execution is supported.

Consider NuExtract3 when your image-to-text workflow calls for a long context window and an Apache-2.0 license.[6] The practical caveat is hardware sizing: a specific GPU or system RAM recommendation cannot be established from the published weight size alone. Validate memory use with your intended runtime and workload before committing hardware.

3. Zero-To-CAD-Qwen3-VL-2B

Zero-To-CAD-Qwen3-VL-2B by ADSKAILab was first published on Hugging Face on April 10, 2026.[13] Its third-place position follows the later publication dates of the preceding entries.[6][11][13] The ranking admits only models tagged image-to-text on Hugging Face; the ordering basis is newest release first, then downloads.

The model has 2.1B parameters, a 256K-token context window, BF16 weights totaling 4.3 GB, and an Apache-2.0 license.[13] Quantized weight storage would be approximately 1.05 GB at 4-bit, estimated (params x 0.5 bytes) from the cited 2.1B parameters.[13] Treat that calculation as a weight-storage estimate, not a verified GPU or system RAM requirement. A concrete local hardware recommendation remains unverified.

Consider the model for evaluating interpretable CAD program generation, the focus of its associated paper, “Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data.”[14] The key caveat is evaluation: no public benchmark yet establishes its position on captioning quality. Its placement reflects release order.

4. nemotron-ocr-v2

nemotron-ocr-v2 by nvidia ranks fourth under the ordering rule of newest release first, then downloads. Its first Hugging Face publication was April 1, 2026, placing it after ADSKAILab’s Zero-To-CAD-Qwen3-VL-2B and before FireRedTeam’s FireRed-OCR.[3][13][9] The ranking covers only models tagged image-to-text on Hugging Face; placement does not establish captioning quality.

The license is nvidia-open-model-license.[3] Parameter count, context length and weight size are unconfirmed here. Without a parameter count, a defensible quantized-weight memory estimate is unavailable, and no specific GPU or RAM capacity can be recommended for local execution.

Treat nemotron-ocr-v2 as a candidate for a local image-to-text evaluation using representative images and expected outputs. The key caveat is no public benchmark yet: the available information does not establish its accuracy or hardware requirements. Check those requirements before committing hardware or integrating it into a captioning workflow.

5. FireRed-OCR

FireRed-OCR by FireRedTeam ranks fifth under the newest release first, then downloads ordering, with a Hugging Face publication date of 2026-02-28.[9] Its release falls between NVIDIA’s nemotron-ocr-v2 and tiiuae’s Falcon-OCR.[3][4][9] Benchmark status: no public benchmark yet. Its position reflects release timing, with captioning quality still unverified.

FireRed-OCR has 2.1B parameters, a 256K-token context window and 4.3 GB of BF16 weights, under the Apache-2.0 license.[9] The 4-bit weight budget is approximately 1.05 GB, estimated (params x 0.5 bytes) from its 2.1B parameters.[9] Treat that figure as a weight-storage estimate; allow additional RAM or VRAM for execution before choosing local hardware. A specific device fit remains unverified.

Consider FireRed-OCR for local OCR evaluation, consistent with its documented OCR focus.[10] The practical caveat for an image-captioning workflow is the missing public benchmark: validate its descriptions on representative images before adopting it.

6. Falcon-OCR

Falcon-OCR by tiiuae ranks sixth under the newest release first, then downloads ordering, with a first Hugging Face publication date of February 22, 2026 [4]. Its position reflects release timing, not demonstrated captioning quality: no public benchmark yet. Downloads serve only as a tiebreak.

Falcon-OCR has 270M parameters, distributed as 1.1 GB of F32 weights under the Apache-2.0 license [4]. Weight storage at 4-bit would be 135 MB, estimated (params x 0.5 bytes) from the 270M parameter count [4]. Treat that calculation as a starting point for hardware planning: it does not establish total RAM or VRAM requirements, available quantization support, or a specific compatible GPU.

Consider Falcon-OCR for local OCR evaluation, consistent with its association with the “Falcon Perception” paper [5]. The practical caveat is the lack of a public benchmark: evaluate recognition accuracy on your own images before choosing it for a captioning workflow.

7. GLM-OCR

GLM-OCR by zai-org ranks seventh under the newest-release-first ordering: its Hugging Face publication date is January 30, 2026,[1] after tiiuae’s Falcon-OCR[4] and before lightonai’s LightOnOCR-1B-1025.[7] GLM-OCR has no public benchmark yet, so its position reflects release timing rather than a demonstrated captioning-quality advantage.

GLM-OCR has 1.3B parameters,[1] a 128K-token context window,[1] and an MIT license.[1] Published BF16 weights occupy 2.7 GB.[1] For local hardware planning, quantized weight storage is approximately 0.65 GB at 4-bit—estimated (params x 0.5 bytes), using the cited 1.3B parameter count.[1] Treat that figure as a weight-storage estimate, not a measured runtime memory requirement or confirmation that a particular GPU can run it.

Choose GLM-OCR as a candidate for OCR-focused work, consistent with its documented purpose in the GLM-OCR Technical Report.[2] The practical caveat is that neither general image-captioning quality nor a complete local RAM or VRAM requirement is established here; validate both against your intended workload before committing hardware.

8. LightOnOCR-1B-1025

LightOnOCR-1B-1025 by lightonai ranks eighth under the image-to-text-only ranking’s ordering rule: newest release first, then downloads. Its first publication on Hugging Face was October 20, 2025.[7] Benchmark status: no public benchmark yet, so its position does not establish a captioning-quality gap against the other candidates.

The model has 1.2B parameters, an 8K-token context window, and 2.3 GB of BF16 weights, with an Apache-2.0 license.[7] For local hardware planning, the 1.2B parameters imply approximately 0.6 GB for 4-bit weights, estimated (params x 0.5 bytes).[7] Treat that estimate as a weight-storage budget; a specific GPU or total system-RAM requirement is not established.

Consider LightOnOCR-1B-1025 for multilingual optical character recognition, the task described in its technical report.[8] The practical caveat is task fit: its OCR focus should guide selection, while general image-captioning quality remains unestablished by a public benchmark.

How do you estimate memory requirements from parameter counts?

Estimate weight memory by multiplying the parameter count by the bytes stored per parameter. Treat the result as a weight-storage estimate, not a complete RAM or VRAM requirement. Keep the precision assumption attached to the result so readers can distinguish calculated storage from published weight sizes.

For numind’s NuExtract3 (numind), the parameter count is 4.5B: 4-bit weight memory is estimated (params x 0.5 bytes) at 2.25 GB.[6] For zai-org’s GLM-OCR (zai-org), the parameter count is 1.3B: 4-bit weight memory is estimated (params x 0.5 bytes) at 0.65 GB.[1] Both calculations use decimal gigabytes and describe weights alone.

Compare those estimates with the published weight files carefully. NuExtract3 lists BF16 weights of 9.3 GB,[6] while GLM-OCR lists BF16 weights of 2.7 GB.[1] A parameter-based calculation and a published file size are different measurements; neither establishes the total memory required to run a particular workload.

Choose hardware only after checking the runtime’s total memory requirement for your intended input and output. NuExtract3 lists a context length of 256K tokens,[6] and GLM-OCR lists 128K tokens,[1] but those limits do not establish memory consumption. Report the calculated weight budget separately, and leave hardware-fit claims unconfirmed until the intended configuration has been measured.

What context lengths and licenses do local image captioning models offer?

Local image captioning models offer context windows from 8K tokens in lightonai’s LightOnOCR-1B-1025 [7] to 256K tokens in numind’s NuExtract3 [6], with licenses including Apache-2.0 [6][7], MIT [1] and nvidia-open-model-license [3].

NuExtract3 combines its 256K-token context window with an Apache-2.0 license [6]. FireRedTeam’s FireRed-OCR and ADSKAILab’s Zero-To-CAD-Qwen3-VL-2B also specify 256K-token context windows and Apache-2.0 licensing [9][13]. Context length and license therefore do not distinguish those models; use your own captioning workload to evaluate which meets your requirements.

zai-org’s GLM-OCR specifies a 128K-token context window and uses the MIT license [1]. LightOnOCR-1B-1025 specifies an 8K-token window under Apache-2.0 [7]. For a deployment decision, record the context limit alongside the license rather than treating either as a substitute for evaluating caption quality.

Apache-2.0 also covers tiiuae’s Falcon-OCR [4] and PaddlePaddle’s PP-OCRv6_medium_det_safetensors [11]. nvidia’s nemotron-ocr-v2 uses nvidia-open-model-license [3]. Review the applicable license terms for your intended deployment, modification and redistribution plans before integrating a model. Keep that review tied to the exact repository you intend to use.

How are image-to-text models ranked when no public benchmark scores them?

When no public benchmark scores the eligible models, the ordering is newest release first, then downloads. The ranking admits only image to text models with the Hugging Face pipeline tag image-to-text. Eligibility is limited to releases within the past twelve months, including the October release from lightonai and the June release from PaddlePaddle.[7][11]

PaddlePaddle’s PP-OCRv6_medium_det_safetensors (PaddlePaddle) takes first place because its June 9, 2026 publication date is the newest among the eligible candidates.[11] numind’s NuExtract3 (numind) follows with an April 29, 2026 publication date, ahead of ADSKAILab’s Zero-To-CAD-Qwen3-VL-2B (ADSKAILab), published April 10, 2026.[6][13] Those positions reflect release order; they do not establish comparative caption quality.

Each entry carries “no public benchmark yet.” A public leaderboard can rank only models evaluated against one another on that board, and an older scored model cannot move above a newer unscored model. Any model-card benchmark must be labeled “self-reported.” Downloads break release-date ties; popularity alone cannot establish a quality advantage.

Read the order as a starting point for local evaluation. Compare captions on your own images before choosing a model. Parameters, context length and license help assess deployment requirements, but model size does not determine rank. Any memory figure derived from parameter count must be explicitly labeled as an estimate.

Frequently Asked Questions

Which local image-captioning model ranks first?

PP-OCRv6_medium_det_safetensors (PaddlePaddle), from PaddlePaddle, ranks first under the release-date rule with its June 9, 2026 publication date [11]. Ranking note: this ranking admits only image to text models (Hugging Face pipeline tags: image-to-text). The ordering basis is newest release first, then downloads; downloads serve only as a tiebreak. Placement does not establish superior captioning quality.

What is the full ranking of the eligible models?

The order is PP-OCRv6_medium_det_safetensors (PaddlePaddle) [11], NuExtract3 (numind) [6], Zero-To-CAD-Qwen3-VL-2B (ADSKAILab) [13], nemotron-ocr-v2 (nvidia) [3], FireRed-OCR (FireRedTeam) [9], Falcon-OCR (tiiuae) [4], GLM-OCR (zai-org) [1], and LightOnOCR-1B-1025 (lightonai) [7]. Each has no public benchmark yet. Read the sequence as a release-recency ranking, with downloads reserved for ties, rather than a measured comparison of image-captioning accuracy.

How much memory would quantized model weights need?

PP-OCRv6_medium_det_safetensors (PaddlePaddle) has 22M parameters [11]: 11 MB at 4-bit, estimated (params x 0.5 bytes). GLM-OCR (zai-org) has 1.3B parameters [1]: 650 MB at 4-bit, estimated (params x 0.5 bytes). NuExtract3 (numind) has 4.5B parameters [6]: 2.25 GB at 4-bit, estimated (params x 0.5 bytes). Those calculations cover parameter storage only; they do not establish total runtime memory, quantization availability, or compatibility with a particular GPU.

What context lengths do the models support?

NuExtract3 (numind) supports a 256K-token context [6], as do Zero-To-CAD-Qwen3-VL-2B (ADSKAILab) [13] and FireRed-OCR (FireRedTeam) [9]. GLM-OCR (zai-org) supports 128K tokens [1], while LightOnOCR-1B-1025 (lightonai) supports 8K tokens [7]. Context length is a specification to compare separately from captioning quality: those limits do not provide a benchmark result or establish how much memory a local workload will require.

Which licenses apply to these local models?

Apache-2.0 covers PP-OCRv6_medium_det_safetensors (PaddlePaddle) [11], NuExtract3 (numind) [6], Zero-To-CAD-Qwen3-VL-2B (ADSKAILab) [13], FireRed-OCR (FireRedTeam) [9], Falcon-OCR (tiiuae) [4], and LightOnOCR-1B-1025 (lightonai) [7]. GLM-OCR (zai-org) uses MIT [1]. nemotron-ocr-v2 (nvidia) uses nvidia-open-model-license [3]. Check the applicable license text against your intended deployment and redistribution plans before selecting a model; the NVIDIA entry should receive its own review because it carries a different license.

Are there public benchmarks proving which model captions images better?

No public benchmark yet scores these candidates against each other, so the ranking cannot establish a caption-quality winner. Technical-report publication dates are separate from benchmark dates: the GLM-OCR Technical Report was published March 11, 2026 [2], and the FireRed-OCR Technical Report appeared March 2, 2026 [10]. Any model-card benchmark should be labeled self-reported, and an older scored model must not outrank a newer unscored candidate.

Sources

  1. zai-org/GLM-OCR model card (Hugging Face) — 2026-09-25
  2. GLM-OCR Technical Report — 2026-03-11
  3. nvidia/nemotron-ocr-v2 model card (Hugging Face) — 2026-09-25
  4. tiiuae/Falcon-OCR model card (Hugging Face) — 2026-09-25
  5. Falcon Perception — 2026-03-28
  6. numind/NuExtract3 model card (Hugging Face) — 2026-09-25
  7. lightonai/LightOnOCR-1B-1025 model card (Hugging Face) — 2026-09-25
  8. LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR — 2026-01-20
  9. FireRedTeam/FireRed-OCR model card (Hugging Face) — 2026-09-25
  10. FireRed-OCR Technical Report — 2026-03-02
  11. PaddlePaddle/PP-OCRv6_medium_det_safetensors model card (Hugging Face) — 2026-09-25
  12. PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks — 2026-06-11
  13. ADSKAILab/Zero-To-CAD-Qwen3-VL-2B model card (Hugging Face) — 2026-09-25
  14. Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data — 2026-04-27

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog