Quick Answer
Sapiens2 normal-0.4b from facebook is the top pick as of September 2026 under the ordering of newest release first, then downloads, with no public benchmark yet. [7][9] In order, the ranking is Sapiens2 normal-0.4b (facebook), Sapiens2 normal-1b (facebook), TIPSv2 B14 DPT (google), Depth Pro (apple), and DepthCrafter (tencent) [7][9].
Key Takeaways
- Ranking note: this ranking admits only depth estimation models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: depth-estimation). Every candidate has no public benchmark yet; ordering is newest release first, then downloads, with downloads used only to break ties.[3][7][9][1][5] Established picks bypass the recency gate, follow the same ordering rule, and cannot take first place through that exemption.[1][5]
- facebook’s Sapiens2 normal-0.4b ranks first because it shares the newest publication date, April 23, 2026, with Sapiens2 normal-1b and breaks that tie with 2,483 downloads versus 746.[9][7] Its 453M parameters imply 226.5 MB of 4-bit weights, estimated (params x 0.5 bytes); its license is sapiens2-license.[9]
- facebook’s Sapiens2 normal-1b has 1.5B parameters, implying 750 MB of 4-bit weights, estimated (params x 0.5 bytes).[7] Its license is sapiens2-license; the listed F32 weights occupy 12.3 GB.[7]
- google’s TIPSv2 B14 DPT has 158M parameters, implying 79 MB of 4-bit weights, estimated (params x 0.5 bytes).[3] Its Apache-2.0 license and April 9, 2026 publication date distinguish its licensing and release position.[3]
- apple’s Depth Pro is an established pick, published November 27, 2024, that bypasses the recency gate.[1] Its 952M parameters imply 476 MB of 4-bit weights, estimated (params x 0.5 bytes); its license is apple-amlr.[1]
- tencent’s DepthCrafter is an established pick, published September 14, 2024, with listed full-precision weights of 3.0 GB and a license field recorded as “license.”[5] Weight sizes alone do not establish Apple Silicon compatibility, runtime memory requirements, or inference speed.
How do depth estimation models compare on parameters, estimated memory, licenses, and public benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Sapiens2 normal-0.4b [9] | 453M [9] | 4-bit weights: 226.5 MB, estimated (params x 0.5 bytes) [9]; runtime VRAM not published | 2026-04-23 [9] | sapiens2-license [9] | no public benchmark yet | |
| Sapiens2 normal-1b [7] | 1.5B [7] | 4-bit weights: 750 MB, estimated (params x 0.5 bytes) [7]; runtime VRAM not published | 2026-04-23 [7] | sapiens2-license [7] | no public benchmark yet | |
| TIPSv2 B14 DPT [3] | 158M [3] | 4-bit weights: 79 MB, estimated (params x 0.5 bytes) [3]; runtime VRAM not published | 2026-04-09 [3] | Apache-2.0 [3] | no public benchmark yet | |
| Depth Pro — established pick [1] | apple | 952M [1] | 4-bit weights: 476 MB, estimated (params x 0.5 bytes) [1]; runtime VRAM not published | 2024-11-27 [1] | apple-amlr [1] | no public benchmark yet |
| DepthCrafter — established pick [5] | tencent | — | Full-precision weights: 3.0 GB [5]; quantized size and runtime VRAM not published | 2024-09-14 [5] | license [5] | no public benchmark yet |
Which depth estimation models should you consider for an Apple Silicon Mac?
1. Sapiens2 normal-0.4b
Sapiens2 normal-0.4b by facebook ranks first because its release date matches facebook’s Sapiens2 normal-1b, while its download count breaks the tie.[9][7] Both were published on April 23, 2026; their recorded monthly downloads were 2,483 and 746, respectively.[9][7] The ordering basis is newest release first, then downloads; no public benchmark yet establishes a quality advantage. The ranking admits only depth estimation models from labs with a published paper or leaderboard record, using Hugging Face’s depth-estimation pipeline tag.
The model has 453M parameters and lists F32 weights totaling 3.6 GB.[9] Quantized weight storage would be approximately 226.5 MB at 4-bit, estimated (params x 0.5 bytes) from the cited 453M parameters.[9] That calculation describes parameter storage alone, rather than total application memory or an available quantized release. The listed license is sapiens2-license; review its terms before choosing a deployment use.[9]
For Apple Silicon Macs, treat the model as a candidate for local evaluation where you can check output quality and runtime compatibility yourself. A minimum Mac RAM configuration cannot be established from the weight figures alone. Runtime support, working memory requirements and Apple Silicon performance remain unverified here, so the ranking position should guide evaluation rather than a hardware purchase. The practical caveat is straightforward: recency and downloads establish its position, while measured depth quality and local execution still need validation.
2. Sapiens2 normal-1b
Sapiens2 normal-1b by facebook ranks second under the ordering rule “newest release first, then downloads.”[7][9] Both listed Sapiens2 variants were published on April 23, 2026; normal-1b has 746 downloads against normal-0.4b’s 2,483 for the reported month, so downloads break the release-date tie.[7][9] The ranking admits only depth estimation models from labs with a published paper or leaderboard record, using Hugging Face’s depth-estimation pipeline tag. Sapiens2 has a published paper, but normal-1b has no public benchmark yet.[8]
The model card lists 1.5 billion parameters and an F32 weight download of 12.3 GB.[7] Weight-only memory at 4-bit would be 750 MB, estimated (params x 0.5 bytes) from the cited 1.5 billion parameters.[7] Treat that calculation as a planning estimate: usable quantized weights, Apple Silicon execution support, and total runtime memory requirements remain unverified. No particular Mac or unified-memory configuration can be recommended confidently on those specifications alone.
A practical use is a local evaluation candidate for engineers prepared to establish compatibility and measure memory consumption before integrating the model. The release date earns its position here; demonstrated depth quality or Mac performance does not establish the order. The key caveat is licensing: Sapiens2 normal-1b uses sapiens2-license, whose terms need review against the intended deployment.[7] Input limits and Apple Silicon latency also remain unverified.
3. TIPSv2 B14 DPT
TIPSv2 B14 DPT by google ranks third under the ordering rule “newest release first, then downloads”: its Hugging Face publication date is 2026-04-09, earlier than the releases ahead of it.[3][7][9] The model has no public benchmark yet, so its position does not establish a depth-quality advantage. The ranking admits only depth estimation models from labs with a published paper or leaderboard record, using the Hugging Face depth-estimation pipeline tag; google’s accompanying paper is “TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment.”[4]
The model has 158M parameters and a listed F32 weight size of 0.6 GB.[3] Its weight-only memory at 4-bit is approximately 79 MB, estimated (params x 0.5 bytes) from the cited 158M parameters.[3] Treat that calculation as a storage budget, not a measured runtime footprint or confirmation that a compatible quantized version exists. A specific Apple Silicon Mac or minimum unified-memory capacity cannot be recommended from those figures alone.
A practical use is evaluating a depth-estimation component for a local project where the Apache-2.0 license is a selection requirement.[3] Before choosing hardware, verify the intended runtime’s Apple Silicon support, supported input dimensions and actual memory consumption. The deployment caveat is the gap between compact parameter storage and demonstrated execution: the listed specifications do not establish Mac latency, runtime memory requirements or quantization support.
4. Depth Pro
Depth Pro by apple ranks fourth as an established pick, with a Hugging Face release date of 2024-11-27.[1] Its place among its family’s three most-downloaded models allows it to bypass the recency gate.[1] Ranking note: newest release first, then downloads; Depth Pro has no public benchmark yet. The ranking admits only depth estimation models from labs with a published paper or leaderboard record, using Hugging Face’s depth-estimation pipeline tag.
Depth Pro has 952M parameters, distributed as 1.9 GB of F16 weights under the apple-amlr license.[1] Quantized weight storage would be 476 MB at 4-bit, estimated (params x 0.5 bytes) from the cited 952M parameters.[1] Treat that calculation as a weight-storage estimate, not a complete hardware requirement. The figure does not establish total unified-memory needs, runtime overhead or whether a working quantized implementation is available for an Apple Silicon Mac.
The practical use case is monocular metric depth estimation: deriving metric depth from a single image, with sharp depth detail as a stated design aim.[2] The caveat for local deployment is that a compact weight estimate cannot establish Mac compatibility or performance. Before choosing hardware, confirm that your intended runtime supports the model and measure its peak memory at your intended input resolution. No measured Apple Silicon latency or validated RAM requirement can be stated here.
5. DepthCrafter
DepthCrafter by tencent ranks fifth as an established pick, with a Hugging Face publication date of September 14, 2024.[5] Its established-pick status bypasses the recency gate; the ordering is newest release first, then downloads.[1][3][5][7][9] DepthCrafter has no public benchmark yet, so its position does not establish a quality or speed disadvantage. The ranking admits only depth estimation models from labs with a published paper or leaderboard record, using the Hugging Face depth-estimation pipeline tag.
DepthCrafter lists full-precision weights of 3.0 GB.[5] That download size alone cannot establish the hardware needed to run it locally: a supported Apple Silicon runtime, peak unified-memory requirement and quantized memory footprint remain unconfirmed. A specific Mac configuration therefore cannot be recommended on the available specifications. For deployment planning, require a working implementation and a runtime-memory measurement before treating the weight listing as a capacity estimate.
DepthCrafter targets consistent depth sequences for open-world videos, making video depth generation its intended use.[6] Consider it when consistency across a sequence matters to your workflow, while verifying supported sequence lengths in your chosen implementation. One practical caveat is licensing: the repository labels its license simply license.[5] Review the actual terms before adopting it for a product or redistribution; that label alone does not establish permission for either use.
How can you estimate memory requirements from parameter counts?
Estimate weight memory from the parameter count by assuming 4-bit storage and reporting the result as “estimated (params x 0.5 bytes)”; the calculation is a weight-storage estimate, not a measured Apple Silicon memory requirement.[1][3]
Apple’s Depth Pro (apple/DepthPro-hf) has 952M parameters: 476 MB for weights, estimated (params x 0.5 bytes).[1] Google’s TIPSv2 (google/tipsv2-b14-dpt) has 158M parameters: 79 MB for weights, estimated (params x 0.5 bytes).[3] Both calculations use decimal megabytes and assume that the full parameter count uses the stated storage rate. Neither calculation establishes that a compatible quantized implementation is available.
Keep downloaded weight sizes separate from those estimates. Apple lists F16 weights of 1.9 GB for Depth Pro, while Google lists F32 weights of 0.6 GB for its TIPSv2 checkpoint.[1][3] Those files represent different precision formats, so their sizes should not be presented as quantized memory measurements.
For a Mac purchase or deployment decision, use the calculation as an initial storage budget. Verify total memory consumption with the implementation and input you intend to run before claiming that a model fits. A parameter-based estimate alone does not establish runtime memory use, speed, or Apple Silicon compatibility.
What licenses apply to these depth estimation models?
The licenses are sapiens2-license for Facebook’s facebook/sapiens2-normal-0.4b and facebook/sapiens2-normal-1b [9][7], Apache-2.0 for Google’s google/tipsv2-b14-dpt [3], apple-amlr for Apple’s apple/DepthPro-hf [1], and an entry labeled license for Tencent’s DepthCrafter (tencent) [5].
Google’s model is the candidate with an explicitly named Apache license [3]. For a deployment review, record that identifier alongside the exact repository and revision you intend to use. Keep the license review tied to the downloaded artifact.
Facebook’s candidates share the sapiens2-license identifier [9][7]. Apple’s model lists apple-amlr [1]. Read the corresponding terms before deciding whether either family suits a commercial application, an internal service, or redistribution. Avoid treating the availability of downloadable weights as permission for every intended use.
Tencent’s license field is labeled simply license [5]. Treat that label as a pointer requiring inspection, rather than a description of permitted uses. Check the actual license text before approving the model for your project.
For a local Mac deployment, make licensing a separate selection step: verify your intended use, any redistribution plans, and applicable notice requirements before integrating the weights.
What remains unverified about running these models locally on Apple Silicon?
Apple Silicon compatibility, peak unified-memory use, inference speed, and output quality during local execution remain unverified for these candidates. The ranking should therefore guide evaluation, without promising that a particular checkpoint fits your Mac or runs at a useful speed.
Published weight sizes do not establish a working memory requirement. Apple’s Depth Pro (apple/DepthPro-hf) lists F16 weights of 1.9 GB [1], while Google’s TIPSv2 (google/tipsv2-b14-dpt) lists F32 weights of 0.6 GB [3]. Neither figure establishes peak memory during inference. Quantized loading, memory savings, and any resulting quality changes still need validation on the intended hardware and runtime.
Runtime support also needs confirmation: whether each checkpoint loads successfully, executes through the intended Apple Silicon backend, and completes inference without unsupported operations. A reproducible local evaluation should record the Mac configuration, software versions, input dimensions, loading time, inference time, and peak memory.
Input limits remain unresolved, including supported image dimensions and practical video sequence lengths. Treat context capacity as unverified rather than inferring it from checkpoint size or a paper title.
Every candidate has “no public benchmark yet” under the ranking’s comparison rules. Relative accuracy and speed on Macs therefore remain open questions. License identifiers are listed, but permission for your intended use still requires checking the applicable terms; local execution alone does not resolve licensing.
Frequently Asked Questions
Which model comes first in the ranking, and why?
Facebook’s facebook/sapiens2-normal-0.4b ranks first because it shares the newest publication date with Facebook’s facebook/sapiens2-normal-1b and has more downloads, which breaks the tie. [9][7] Both were published on April 23, 2026; their reported monthly downloads are 2,483 and 746, respectively. [9][7] Both have no public benchmark yet. First place therefore reflects release timing and the download tiebreak, rather than demonstrated depth accuracy or Apple Silicon performance.
What is the ranking order, and which models qualify?
The order is Facebook’s facebook/sapiens2-normal-0.4b, facebook/sapiens2-normal-1b, Google’s google/tipsv2-b14-dpt, Apple’s apple/DepthPro-hf and Tencent’s DepthCrafter (tencent). [9][7][3][1][5] The ordering basis is exactly: newest release first, then downloads. No public benchmark scores any candidate. The ranking admits only depth estimation models from labs with a published paper or leaderboard record, using Hugging Face’s depth-estimation pipeline tag. Apple and Tencent are established picks admitted through the recency exemption. [1][5]
How much memory would quantized weights need?
For 4-bit weights, the parameter-based calculations are: facebook/sapiens2-normal-0.4b, 453M parameters, 0.2265 GB estimated (params x 0.5 bytes); facebook/sapiens2-normal-1b, 1.5B parameters, 0.75 GB estimated (params x 0.5 bytes). [9][7] Google’s model has 158M parameters, giving 0.079 GB estimated (params x 0.5 bytes); Apple’s has 952M, giving 0.476 GB estimated (params x 0.5 bytes). [3][1] Those calculations describe weight storage; Mac runtime memory and quantized execution remain unverified.
Can I choose a model for my Mac from those memory estimates?
A Mac fit recommendation remains unverified for every candidate. Published weight files include F16 weights of 1.9 GB for Apple’s model and F32 weights of 0.6 GB for Google’s model. [1][3] Those figures do not establish Apple Silicon execution, peak unified-memory use or latency. Supported input resolution, sequence limits and other context constraints are also unspecified here. Confirm those details in the intended local runtime before choosing hardware.
Which licenses should I check before using a model?
Google’s google/tipsv2-b14-dpt lists Apache-2.0, Apple’s apple/DepthPro-hf lists apple-amlr, and both Facebook candidates list sapiens2-license. [3][1][7][9] Tencent’s DepthCrafter (tencent) lists the identifier “license,” which does not explain its permissions by itself. [5] Read the applicable terms before commercial deployment, redistribution or modification. Treat the custom identifiers as documents to inspect, rather than assuming every open-weight release grants the same permissions.
Why are Depth Pro and DepthCrafter included despite their older releases?
Apple’s apple/DepthPro-hf and Tencent’s DepthCrafter (tencent) are established picks: each qualifies through its family’s three most-downloaded models and bypasses the twelve-month recency gate. [1][5] Their Hugging Face publication dates are November 27, 2024, and September 14, 2024, respectively. [1][5] Both have no public benchmark yet. The exemption keeps them available for comparison without granting either first place; their positions follow the stated release-date ordering.
Sources
- apple/DepthPro-hf model card (Hugging Face) — 2026-09-25
- Depth Pro: Sharp Monocular Metric Depth in Less Than a Second — 2024-10-02
- google/tipsv2-b14-dpt model card (Hugging Face) — 2026-09-25
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment — 2026-04-13
- tencent/DepthCrafter model card (Hugging Face) — 2026-09-25
- DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos — 2024-09-03
- facebook/sapiens2-normal-1b model card (Hugging Face) — 2026-09-25
- Sapiens2 — 2026-04-23
- facebook/sapiens2-normal-0.4b model card (Hugging Face) — 2026-09-25