Quick Answer
sapiens2-normal-0.4b (facebook) is the top pick as of September 2026 because it shares the newest release date and wins the download tiebreak against sapiens2-normal-1b (facebook) [1][3][5][6][7]. In order, the ranking is sapiens2-normal-0.4b (facebook), sapiens2-normal-1b (facebook), tipsv2-b14-dpt (google), DA3METRIC-LARGE (depth-anything), and DA3NESTED-GIANT-LARGE (depth-anything) [1][3].
Key Takeaways
- Ranking note: This ranking admits only depth estimation models (Hugging Face pipeline tags: depth-estimation); all candidates have “no public benchmark yet,” so the ordering is newest release first, then downloads, with downloads used only to break release-date ties.[1][3][5][6][7]
- sapiens2-normal-0.4b (facebook) ranks first because it shares the newest publication date, 2026-04-23, and has more downloads than sapiens2-normal-1b (facebook).[3][5] Its 453M parameters imply 226.5 MB of 4-bit weight memory, estimated (params x 0.5 bytes); its license is sapiens2-license.[5]
- sapiens2-normal-1b (facebook) has 1.5B parameters: 750 MB of 4-bit weight memory, estimated (params x 0.5 bytes).[3] Published on 2026-04-23, it uses sapiens2-license.[3]
- tipsv2-b14-dpt (google) has 158M parameters: 79 MB of 4-bit weight memory, estimated (params x 0.5 bytes).[1] Published on 2026-04-09, it uses Apache-2.0.[1]
- DA3METRIC-LARGE (depth-anything) has 334M parameters: 167 MB of 4-bit weight memory, estimated (params x 0.5 bytes), under Apache-2.0.[7] Its 285,312 downloads break the publication-date tie with DA3NESTED-GIANT-LARGE (depth-anything), which has 32,857 downloads; both were published on 2025-11-13.[6][7]
- DA3NESTED-GIANT-LARGE (depth-anything) has 1.7B parameters: 850 MB of 4-bit weight memory, estimated (params x 0.5 bytes).[6] Its cc-by-nc-4.0 license makes noncommercial use a constraint when choosing a local deployment.[6]
How do local depth estimation models compare on specifications and public benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| facebook/sapiens2-normal-0.4b [5] | facebook [5] | 453M [5] | 4-bit weights: 226.5 MB, estimated (params x 0.5 bytes) [5]; VRAM not published | 2026-04-23 [5] | sapiens2-license [5] | no public benchmark yet |
| facebook/sapiens2-normal-1b [3] | facebook [3] | 1.5B [3] | 4-bit weights: 750 MB, estimated (params x 0.5 bytes) [3]; VRAM not published | 2026-04-23 [3] | sapiens2-license [3] | no public benchmark yet |
| google/tipsv2-b14-dpt [1] | google [1] | 158M [1] | 4-bit weights: 79 MB, estimated (params x 0.5 bytes) [1]; VRAM not published | 2026-04-09 [1] | Apache-2.0 [1] | no public benchmark yet |
| depth-anything/DA3METRIC-LARGE [7] | depth-anything [7] | 334M [7] | 4-bit weights: 167 MB, estimated (params x 0.5 bytes) [7]; VRAM not published | 2025-11-13 [7] | Apache-2.0 [7] | no public benchmark yet |
| depth-anything/DA3NESTED-GIANT-LARGE [6] | depth-anything [6] | 1.7B [6] | 4-bit weights: 850 MB, estimated (params x 0.5 bytes) [6]; VRAM not published | 2025-11-13 [6] | cc-by-nc-4.0 [6] | no public benchmark yet |
Which depth estimation models should you run locally?
1. sapiens2-normal-0.4b
sapiens2-normal-0.4b by facebook ranks first because its Hugging Face publication date ties the latest release among the eligible candidates, while downloads break that tie in its favor.[3][5] Published on April 23, 2026, the model recorded 2,483 downloads over the last 30 days as of September 25, 2026.[5] Ranking note: this ranking admits only depth estimation models with the Hugging Face pipeline tag depth-estimation; with no public benchmark scoring any candidate, the ordering is newest release first, then downloads.
The model has 453M parameters; estimated (params x 0.5 bytes), its 4-bit weights would occupy 226.5 MB.[5] The listed F32 weights occupy 3.6 GB, and the license is sapiens2-license.[5] Hardware planning needs more than a weight-storage figure: the calculated footprint is an estimate for weights alone, not a measured runtime memory requirement. A specific GPU or system RAM recommendation cannot be established from the supplied specifications.
For local use, consider the model an evaluation candidate when release recency matters to your selection process and its license suits your project. The benchmark status is “no public benchmark yet,” so its position does not establish depth accuracy or inference speed. The practical caveat is that the quantized weight estimate does not establish a supported quantized execution path or confirm that the model will fit your hardware.
2. sapiens2-normal-1b
sapiens2-normal-1b by facebook ranks second because its Hugging Face publication date matches facebook’s sapiens2-normal-0.4b, but its download count is lower: 746 versus 2,483 over the preceding 30 days as of September 25, 2026.[3][5] Both were published on April 23, 2026.[3][5] The ordering basis is newest release first, then downloads; no public benchmark scores any candidate. The ranking admits only depth estimation models with the Hugging Face pipeline tag depth-estimation. Download counts settle the release-date tie, without establishing a quality advantage.
The model has 1.5 billion parameters, a listed F32 weight package of 12.3 GB, and the sapiens2-license.[3] Weight memory at 4-bit is approximately 0.75 GB, estimated (params x 0.5 bytes) from the cited 1.5 billion parameters.[3] Treat that calculation as a weight-only planning figure, not a verified GPU requirement. A local deployment still needs a confirmed runtime memory budget before you can choose hardware confidently; neither the package size nor the parameter calculation establishes total inference memory.
Consider sapiens2-normal-1b for local evaluation when its license and published weight package suit your project. The practical caveat is “no public benchmark yet”: its position does not demonstrate better depth predictions than the other candidates. Evaluate output quality and runtime memory on your intended workload before committing hardware, and check the sapiens2-license against your intended use.[3]
3. tipsv2-b14-dpt
tipsv2-b14-dpt by google ranks third under the prescribed ordering: newest release first, then downloads.[1][3][5][6][7] Its Hugging Face publication date is April 9, 2026, placing it between the newer facebook releases and the earlier depth-anything releases.[1][3][5][6][7] The ranking admits only depth estimation models with the Hugging Face pipeline tag depth-estimation. Benchmark status: no public benchmark yet; this position does not establish a measured quality advantage.
The model has 158M parameters, distributes 0.6 GB of F32 weights, and uses the Apache-2.0 license.[1] For those 158M parameters, 4-bit weight storage would be approximately 79 MB, estimated (params x 0.5 bytes).[1] That calculation covers parameter storage alone and does not establish the GPU memory or system RAM needed for inference. A specific hardware minimum remains unspecified; confirm runtime requirements before choosing a machine.
A practical use is evaluating local depth estimation when Apache-2.0 licensing matters to your project.[1] Start with the published weights and measure memory use in your intended runtime before planning around quantization. The caveat is that the calculated 4-bit footprint is an estimate, not confirmation of a compatible quantized release or execution path.[1] Confirm supported input dimensions and evaluate depth outputs on your own images before deployment.
4. DA3METRIC-LARGE
DA3METRIC-LARGE by depth-anything ranks fourth under the ordering rule: newest release first, then downloads. Its Hugging Face publication date is 2025-11-13 [7], placing it behind the April releases from facebook and google [1][3][5]. It precedes depth-anything’s DA3NESTED-GIANT-LARGE on the shared publication date because its downloads are higher: 285,312 versus 32,857 over the last 30 days, as of 2026-09-25 [6][7]. The ranking admits only depth estimation models with the Hugging Face pipeline tag depth-estimation.
DA3METRIC-LARGE has 334M parameters and listed F32 weights of 1.3 GB [7]. Four-bit weight storage would be 167 MB, estimated (params x 0.5 bytes) from that parameter count [7]. Treat that calculation as a weight-storage budget when planning local hardware. Before choosing a GPU or allocating system RAM, verify runtime compatibility and measure peak memory at your intended input resolution. A specific GPU recommendation or total RAM requirement cannot be established from the listed specifications.
A practical use is evaluating local depth estimation for a project seeking an Apache-2.0-licensed model [7]. The benchmark caveat is straightforward: no public benchmark yet. Its position reflects release timing and download counts, so the ranking does not establish an accuracy advantage. Validate depth quality on representative images before adopting it, and confirm quantization support before budgeting around the estimated weight size.
5. DA3NESTED-GIANT-LARGE
DA3NESTED-GIANT-LARGE by depth-anything ranks fifth under the ordering rule: newest release first, then downloads. Its Hugging Face publication date is November 13, 2025, shared with DA3METRIC-LARGE by depth-anything.[6][7] The tie is resolved by downloads: DA3NESTED-GIANT-LARGE recorded 32,857 against DA3METRIC-LARGE’s 285,312 over the preceding 30 days, as of September 25, 2026.[6][7] The position reflects release timing and adoption, rather than demonstrated depth-estimation quality: no public benchmark yet.
The model has 1.7B parameters and a listed F32 weight size of 6.8 GB.[6] Hypothetical 4-bit weight storage would be 0.85 GB, estimated (params x 0.5 bytes).[6] That calculation covers weights alone; use it as a storage planning figure, not a complete GPU-memory requirement. A specific GPU or system-RAM recommendation cannot be established from the published weight size alone. Before committing hardware, verify the runtime’s supported precision and measure memory consumption with your intended inputs.
Consider DA3NESTED-GIANT-LARGE for local, noncommercial depth-estimation evaluation where you can assess output quality against your own requirements. The practical caveat is its cc-by-nc-4.0 license, which makes commercial deployment a licensing concern.[6] For an engineering evaluation, establish runtime compatibility, input handling and measured memory requirements before treating the model as a deployment candidate.
How can you estimate weight memory from parameter counts?
Estimate weight memory at 4-bit by multiplying the cited parameter count by 0.5 bytes per parameter; for google’s tipsv2-b14-dpt (google), 158M parameters give 79 MB, estimated (params x 0.5 bytes).[1] Use decimal MB or GB consistently when reporting the result.
For facebook’s sapiens2-normal-0.4b (facebook), the cited 453M parameters give 226.5 MB, estimated (params x 0.5 bytes).[5] For facebook’s sapiens2-normal-1b (facebook), 1.5B parameters give 750 MB, estimated (params x 0.5 bytes).[3] Calculate from the parameter count rather than interpreting the repository name as an exact count.
For depth-anything’s DA3METRIC-LARGE (depth-anything), 334M parameters give 167 MB, estimated (params x 0.5 bytes).[7] For depth-anything’s DA3NESTED-GIANT-LARGE (depth-anything), 1.7B parameters give 850 MB, estimated (params x 0.5 bytes).[6] Apply the same calculation across candidates to keep the comparison consistent.
Keep estimated weight storage separate from published checkpoint size. For example, tipsv2-b14-dpt (google) lists F32 weights at 0.6 GB.[1] The calculated value describes a hypothetical quantized representation; the calculation alone does not establish that a compatible quantized checkpoint or inference path is available.
Treat each result as a planning estimate for weights alone. Avoid turning that estimate into a claim about total inference memory or whether a particular GPU can run the model. Verify runtime requirements separately before choosing hardware.
Which licenses apply to local depth estimation models?
Local depth estimation models in this selection use Apache-2.0[1][7], sapiens2-license[3][5], or cc-by-nc-4.0[6], depending on the checkpoint.
Google’s tipsv2-b14-dpt (google) lists Apache-2.0[1], as does depth-anything’s DA3METRIC-LARGE (depth-anything)[7]. For a project that requires that license, those checkpoints are the relevant candidates here. Record the repository and its license together when documenting your model choice, then review the license text against your intended use.
Facebook’s sapiens2-normal-0.4b (facebook)[5] and sapiens2-normal-1b (facebook)[3] both list sapiens2-license.[3][5] Choosing between those checkpoints does not change the listed license.[3][5] Review the named license directly before incorporating either checkpoint into a product or distributing its weights; do not infer permissions from local availability.
Depth-anything’s DA3NESTED-GIANT-LARGE (depth-anything) lists cc-by-nc-4.0[6], while the organization’s metric checkpoint lists Apache-2.0.[7] A shared organization therefore does not establish a shared license.[6][7] Keep the license attached to the exact checkpoint rather than applying a family-wide label.
For deployment review, describe what you intend to do: run inference internally, provide a service, modify weights, or redistribute them. Check those activities against the applicable license before committing to a checkpoint.
What should you check when a model has no public benchmark yet?
Check the license, parameter count, input requirements and measured performance on your own images before choosing a model with “no public benchmark yet.” Treat release dates and downloads as selection signals, not evidence of depth accuracy. Label any model-card benchmark as self-reported.
Start with the license. Google’s tipsv2-b14-dpt (google) uses Apache-2.0.[1] The depth-anything organization’s DA3METRIC-LARGE (depth-anything) also uses Apache-2.0,[7] while its DA3NESTED-GIANT-LARGE (depth-anything) uses cc-by-nc-4.0.[6] Read the applicable terms before committing to a deployment, particularly for commercial work.
Separate parameter storage from measured runtime memory. Google’s tipsv2-b14-dpt (google) has 158M parameters; its weight storage at 4-bit would be approximately 79 MB, estimated (params x 0.5 bytes).[1] Such arithmetic does not establish quantization support or show whether a particular GPU can run the model. Verify the available implementation and measure memory during inference.
Confirm the accepted image dimensions, preprocessing, output format and whether the output matches your need for metric distance or relative depth. Do not substitute a language-model context-window specification for image-input requirements.
Run each candidate on the same representative images and record the configuration, evaluation date, latency and peak memory. Include scenes that matter to your application, such as reflective surfaces or thin objects. Where reference depth is available, measure errors; otherwise, treat visual inspection as a qualitative check. Base deployment decisions on those recorded results.
Frequently Asked Questions
Which local depth estimation models should I shortlist?
Ranking note: this ranking admits only depth estimation models (Hugging Face pipeline tags: depth-estimation). The ordering is newest release first, then downloads: sapiens2-normal-0.4b (facebook) [5], sapiens2-normal-1b (facebook) [3], tipsv2-b14-dpt (google) [1], DA3METRIC-LARGE (depth-anything) [7], and DA3NESTED-GIANT-LARGE (depth-anything) [6]. Downloads resolve publication-date ties; the ordering does not establish comparative depth accuracy.
Why does the smaller Sapiens2 model rank first?
sapiens2-normal-0.4b (facebook) ranks first because it shares the newest publication date, April 23, 2026, with sapiens2-normal-1b (facebook), then wins the download tiebreak with 2,483 versus 746 downloads over the preceding 30 days as of September 25, 2026. [5][3] Popularity determines that tie only. Public benchmark status remains no public benchmark yet, so the position should not be read as evidence of superior depth accuracy.
How much memory would quantized Sapiens2 weights need?
sapiens2-normal-0.4b (facebook) has 453M parameters [5]: 4-bit weight memory is 226.5 MB, estimated (params x 0.5 bytes). [5] sapiens2-normal-1b (facebook) has 1.5B parameters [3]: 4-bit weight memory is 750 MB, estimated (params x 0.5 bytes). [3] Both calculations describe weights only. Neither estimate establishes total runtime memory, confirms a compatible quantized checkpoint, or supports a claim that the model fits a particular GPU.
What are the estimated quantized weight sizes for TIPSv2 and Depth Anything?
tipsv2-b14-dpt (google) has 158M parameters [1]: 4-bit weights require 79 MB, estimated (params x 0.5 bytes). [1] DA3METRIC-LARGE (depth-anything) has 334M parameters [7]: 4-bit weights require 167 MB, estimated (params x 0.5 bytes). [7] DA3NESTED-GIANT-LARGE (depth-anything) has 1.7B parameters [6]: 4-bit weights require 850 MB, estimated (params x 0.5 bytes). [6] Treat these figures as weight-storage estimates, not measured inference memory or hardware-fit guarantees.
Which licenses should I check before choosing a model?
tipsv2-b14-dpt (google) and DA3METRIC-LARGE (depth-anything) list Apache-2.0. [1][7] DA3NESTED-GIANT-LARGE (depth-anything) lists cc-by-nc-4.0. [6] Both facebook Sapiens2 candidates list sapiens2-license. [5][3] Check the applicable license text against your intended deployment before committing to an integration. Open weights alone should not substitute for that review, and the shared Depth Anything organization does not mean its listed models share a license. [6][7]
Are there benchmark results and context limits I can compare?
Public benchmark status for every candidate in this ranking is no public benchmark yet; no dated comparative score is available to justify an accuracy ordering. Context limits and supported input resolutions are also unspecified here. The TIPSv2 paper is dated April 13, 2026 [2], and the Sapiens2 paper April 23, 2026 [4]; those publication dates are not benchmark results.
Sources
- google/tipsv2-b14-dpt model card (Hugging Face) — 2026-09-25
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment — 2026-04-13
- facebook/sapiens2-normal-1b model card (Hugging Face) — 2026-09-25
- Sapiens2 — 2026-04-23
- facebook/sapiens2-normal-0.4b model card (Hugging Face) — 2026-09-25
- depth-anything/DA3NESTED-GIANT-LARGE model card (Hugging Face) — 2026-09-25
- depth-anything/DA3METRIC-LARGE model card (Hugging Face) — 2026-09-25