Best Image-to-3D Models for Apple Silicon Macs in 2026

Rankings 2026-09-27 Last updated 2026-09-27 15 min read By Q4KM

Quick Answer

Déjà View: Looping Transformers for Multi-View 3D Reconstruction (NVIDIA) is the top pick as of September 2026 because it is the newest eligible release, with no public benchmark yet; the ranking admits only image-to-3d models from labs with a published paper or leaderboard record and follows newest release first, then downloads. [15][16]

In order, the ranking is Déjà View: Looping Transformers for Multi-View 3D Reconstruction (NVIDIA), HY-World 2.0 (Tencent), Lyra 2.0 (NVIDIA), Sharp (Apple), TRELLIS.2-4B (Microsoft), HunyuanWorld-Mirror (Tencent), VGGT-1B (established pick) (Facebook), and TRELLIS-image-large (established pick) (Microsoft) [15][16].

Key Takeaways

How do the image-to-3D candidates compare on parameters, memory estimates, licenses and published benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
Déjà View: Looping Transformers for Multi-View 3D Reconstruction [16] NVIDIA [15] 117M [15] F32 weights: 0.5 GB [15]; 4-bit weights: 58.5 MB, estimated (params x 0.5 bytes), from 117M parameters [15]; VRAM: not published 2026-05-29 [15] nvidia-oneway-noncommercial [15] no public benchmark yet
HY-World 2.0 [8] Tencent [7] — Full-precision weights: 174.5 GB [7]; quantized weights/VRAM: not published 2026-04-10 [7] tencent-hy-world-2.0-community [7] no public benchmark yet
Lyra 2.0 [14] NVIDIA [13] — Full-precision weights: 1.1 GB [13]; quantized weights/VRAM: not published 2026-04-06 [13] — no public benchmark yet
Sharp [11] Apple [11] — — 2025-12-12 [11] apple-amlr [11] no public benchmark yet
TRELLIS.2-4B [3] Microsoft [3] — Full-precision weights: 16.2 GB [3]; quantized weights/VRAM: not published 2025-12-01 [3] MIT [3] no public benchmark yet
HunyuanWorld-Mirror [9] Tencent [9] 1.3B [9] F32 weights: 5.1 GB [9]; 4-bit weights: 0.65 GB, estimated (params x 0.5 bytes), from 1.3B parameters [9]; VRAM: not published 2025-10-16 [9] tencent-hunyuanworld-mirror-community [9] no public benchmark yet
VGGT-1B (Visual Geometry Grounded Transformer; established pick) [5][6] Facebook [5] 1.3B [5] F32 weights: 5.0 GB [5]; 4-bit weights: 0.65 GB, estimated (params x 0.5 bytes), from 1.3B parameters [5]; VRAM: not published 2025-03-11 [5] cc-by-nc-4.0 [5] no public benchmark yet
TRELLIS-image-large (established pick) [1] Microsoft [1] — Full-precision weights: 3.3 GB [1]; quantized weights/VRAM: not published 2024-12-02 [1] MIT [1] no public benchmark yet

Which image-to-3D models should you consider for an Apple Silicon Mac?

1. Déjà View: Looping Transformers for Multi-View 3D Reconstruction

Déjà View: Looping Transformers for Multi-View 3D Reconstruction by NVIDIA ranks first because its Hugging Face publication date is the newest among the eligible candidates: May 29, 2026.[15] Ranking note: newest release first, then downloads. Eligibility is limited to image-to-3d models from labs with a published paper or leaderboard record; established picks bypass the recency gate. Déjà View has no public benchmark yet.

NVIDIA’s nvidia/dvlt has 117M parameters and 0.5 GB of F32 weights.[15] Quantized weight storage would be 58.5 MB at 4-bit, estimated (params x 0.5 bytes) from the cited 117M parameters.[15] That calculation covers weights alone. Apple Silicon runtime support, working quantization, context limits and total memory requirements remain unverified, so a specific Mac configuration cannot be recommended.

Multi-view reconstruction is the appropriate use to investigate, matching the model’s stated purpose.[16] Local deployment still requires runtime compatibility and memory validation. The practical licensing caveat is NVIDIA’s nvidia-oneway-noncommercial license: evaluate its terms before adopting the model for a project.[15]

2. HY-World 2.0

HY-World 2.0 by Tencent ranks second under the prescribed ordering: newest release first, then downloads.[7][13][15] Its Hugging Face release date is April 10, 2026.[7] The benchmark status is “no public benchmark yet,” so its position reflects release timing rather than demonstrated quality or speed on Apple Silicon.

The full-precision weights total 174.5 GB, and the license is Tencent’s tencent-hy-world-2.0-community license.[7] Apple Silicon compatibility and runtime memory requirements remain unverified. A specific Mac configuration therefore cannot be recommended, and the weight-file size should not be treated as a RAM requirement. A parameter-based estimate for quantized weights is also unavailable.

The intended use is reconstructing, generating and simulating environments through a multimodal world model.[8] Choose it for investigating those world-building tasks. The practical caveat is local deployment: establish a working Mac runtime and its memory requirements before committing hardware to the project.

3. Lyra 2.0

Lyra 2.0 by NVIDIA takes third place because its Hugging Face release follows NVIDIA’s dvlt and Tencent’s HY-World-2.0 under the ordering rule: newest release first, then downloads.[13][15][7] Published on April 6, 2026, Lyra has no public benchmark yet, so its position does not establish reconstruction quality or Mac performance.[13]

Full-precision weights total 1.1 GB.[13] A parameter count is unavailable for calculating a defensible quantized-weight estimate. The published weight size alone cannot establish how much unified memory a Mac needs: hardware requirements and Apple Silicon execution remain unverified. No specific Mac configuration can therefore be recommended confidently.

Lyra’s stated focus is explorable generative 3D worlds, making world exploration the relevant use case to evaluate.[14] The practical caveat is deployment uncertainty: a compact weight download does not demonstrate a working local Mac setup. Confirm a compatible runtime and license terms before committing to an integration.

4. Sharp

Sharp by Apple ranks fourth because its Hugging Face publication date falls between NVIDIA’s Lyra-2.0 and Microsoft’s TRELLIS.2-4B.[11][13][3] The ordering is newest release first, then downloads. Sharp was first published on December 12, 2025, and has no public benchmark yet; its position does not establish better performance on Apple Silicon.[11]

Sharp targets monocular view synthesis, making work that starts with a single image and needs alternative views the relevant use case.[12] Its license is apple-amlr.[11] Review those terms before adopting Sharp in a project rather than assuming unrestricted use.

The practical caveat is hardware uncertainty. Apple Silicon compatibility, required unified memory, quantized weight memory and input limits remain unverified. A defensible Mac configuration or quantization estimate cannot be specified without further implementation details. Before committing to a local workflow, confirm a compatible runtime and its documented memory requirements.

5. TRELLIS.2-4B

TRELLIS.2-4B by Microsoft ranks fifth under the prescribed ordering: newest release first, then downloads. Its Hugging Face publication date is December 1, 2025 [3], and it has no public benchmark yet. The model provides 16.2 GB of full-precision weights under the MIT license [3]. Its accompanying paper describes native and compact structured latents for generating 3D content [4].

Consider it for image-to-3D projects where permissive licensing matters: the MIT license is a concrete reason to evaluate this candidate [3]. Local hardware planning remains unresolved, however. The 16.2 GB weight size [3] does not establish total runtime memory or a suitable Mac configuration. Confirm Apple Silicon backend support and measured memory use before committing hardware.

The practical caveat is that a verified Apple Silicon configuration and quantized memory footprint are not established. Treat local Mac suitability as unverified; the ranking position reflects release order rather than demonstrated Mac performance.

6. HunyuanWorld-Mirror

HunyuanWorld-Mirror by Tencent ranks sixth under the ordering rule, newest release first, then downloads, following its October 16, 2025 Hugging Face debut.[9] The model has no public benchmark yet, so its position does not establish a quality advantage. Its associated paper, “WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting,” makes reconstruction with prior prompting the relevant use case to evaluate.[10]

HunyuanWorld-Mirror has 1.3 billion parameters and 5.1 GB of F32 weights.[9] A 4-bit weight payload would be approximately 0.65 GB, estimated (params x 0.5 bytes) from the cited parameter count.[9] Actual hardware requirements remain unverified: that calculation covers weights alone and does not establish total unified-memory needs or Apple Silicon compatibility. Before choosing a Mac for local deployment, validate a compatible runtime and measure peak memory. The model uses the tencent-hunyuanworld-mirror-community license; review its terms for your intended use.[9]

7. VGGT-1B (established pick)

VGGT-1B (established pick) by Facebook ranks seventh under the ordering rule: newest release first, then downloads.[5] Its publication date, March 11, 2025, places it behind Tencent’s HunyuanWorld-Mirror and ahead of Microsoft’s TRELLIS-image-large.[5][9][1] Its established-pick status permits inclusion outside the recency window; no public benchmark yet supports a performance advantage.

VGGT-1B has 1.3 billion parameters and 5.0 GB of F32 weights.[5] Weight storage at 4-bit is approximately 0.65 GB, estimated (params x 0.5 bytes) from that parameter count.[5] The estimate covers weights alone, not total runtime memory. Apple Silicon compatibility, a working quantized implementation and a minimum Mac memory configuration remain unconfirmed, so the weight estimate cannot establish which Mac will run it.

Consider VGGT-1B for visual-geometry experimentation, consistent with its Visual Geometry Grounded Transformer formulation.[6] The practical caveat is licensing: its Creative Commons Attribution-NonCommercial 4.0 license restricts commercial use.[5] Engineers evaluating it for a commercial workflow should resolve that restriction before investing in local integration.

8. TRELLIS-image-large (established pick)

TRELLIS-image-large (established pick) by Microsoft sits last under the ordering rule: newest release first, then downloads. Its established-pick exemption admits its December 2, 2024 release despite the recency cutoff.[1] Benchmark status: no public benchmark yet. Its 2,159,277 downloads over the reporting month indicate adoption, but popularity does not override release order.[1]

Microsoft lists full-precision weights of 3.3 GB and an MIT license.[1] The practical use case is image-conditioned generation using structured representations of geometry, as described in Structured 3D Latents for Scalable and Versatile 3D Generation.[2] The permissive license makes it a candidate for projects that need adaptable licensing terms.[1]

For an Apple Silicon Mac, hardware requirements remain unconfirmed: weight storage alone does not establish the unified memory needed to run locally. A supported Mac execution path, runtime memory requirement and quantized weight size are not specified here. Treat local compatibility as a prerequisite to verify before choosing this model.

How can you estimate quantized weight memory from parameter counts?

Estimate quantized weight memory by multiplying the parameter count by the storage per parameter: NVIDIA’s dvlt (nvidia/dvlt) has 117 million parameters [15], giving 58.5 MB for quantized weights at 4-bit, estimated (params x 0.5 bytes) [15].

Apply the same calculation to Facebook’s Visual Geometry Grounded Transformer (VGGT-1B), with 1.3 billion parameters [5]: 650 MB, estimated (params x 0.5 bytes) [5]. Tencent’s HunyuanWorld-Mirror also has 1.3 billion parameters [9], giving 650 MB, estimated (params x 0.5 bytes) [9]. Those results use decimal megabytes and describe weight storage alone.

Use an explicit parameter count rather than inferring one from a model name or download size. Microsoft’s TRELLIS.2-4B has a listed full-precision weight size of 16.2 GB [3], but that entry does not establish the parameter count needed for this calculation. Leave its quantized weight estimate unspecified rather than treating the name as a verified count.

Treat every calculated result as a planning estimate, not a measured application footprint or a guarantee that the model fits an Apple Silicon Mac. Before choosing hardware, verify that the intended runtime supports the model and quantization format, then check total memory use during inference. A weight-storage calculation alone does not establish local compatibility or performance.

What remains unverified about local execution on Apple Silicon Macs?

Local execution on Apple Silicon Macs remains unverified: compatibility, working quantization, peak unified-memory use, supported input limits and practical inference speed still need confirmation for each model.

Published weight sizes do not establish runtime memory requirements. NVIDIA’s dvlt (nvidia/dvlt) lists F32 weights of 0.5 GB [15], while Microsoft’s TRELLIS.2-4B (TRELLIS.2-4B (Microsoft)) lists full-precision weights of 16.2 GB [3]. Neither figure establishes which Mac configuration can complete an inference run. A useful validation must measure peak memory across loading, inference and output generation, with the input settings recorded.

Quantized execution also needs a reproducible implementation. Parameter-based weight estimates cannot establish that a compatible quantized checkpoint exists, that the required operations execute correctly on Apple’s GPU, or that output quality survives quantization. Installation steps, backend selection and any CPU fallback need documentation before a model can receive a practical local-use recommendation.

Context and input capacity remain unresolved. Engineers need verified limits for image resolution, input views and any model-specific context setting, alongside their effect on memory and latency. Each candidate has “no public benchmark yet” under the ranking criteria, leaving comparative quality and Mac performance unverified.

License checks remain separate from execution checks. Microsoft publishes TRELLIS.2-4B under MIT [3]; that permission does not establish Apple Silicon compatibility. A recommendation needs a documented successful run on a specified Mac configuration.

Which model licenses allow commercial use?

Microsoft’s TRELLIS-image-large and TRELLIS.2-4B allow commercial use under the MIT license.[1][3] Both are options for a commercial project: MIT permits using, modifying and distributing the licensed software, provided its copyright and permission notices are retained.[1][3]

Facebook’s VGGT-1B uses Creative Commons Attribution-NonCommercial 4.0, so its published license does not permit commercial use.[5] NVIDIA’s dvlt uses the NVIDIA OneWay Noncommercial license and likewise is not licensed for commercial use under those terms.[15] Treat either model as requiring separate permission before commercial deployment; downloadable weights do not establish commercial rights.

Tencent’s HY-World-2.0 uses the tencent-hy-world-2.0-community license, while Tencent’s HunyuanWorld-Mirror uses tencent-hunyuanworld-mirror-community.[7][9] Those identifiers alone do not establish whether your intended commercial use is permitted. Review each agreement’s permissions, restrictions and obligations before adopting either model for a product or client workflow.

Apple’s Sharp lists apple-amlr; commercial permission should remain unconfirmed until the applicable terms have been checked.[11] Commercial-use permission for NVIDIA’s Lyra-2.0 is also unconfirmed.[13] For an initial shortlist based on explicit commercial permission, start with Microsoft’s MIT-licensed models.[1][3] Evaluate any additional dependencies and rights to input images separately before shipping.

Frequently Asked Questions

Which image-to-3D model leads this ranking?

NVIDIA’s nvidia/dvlt takes first place under the release-date rule, with a Hugging Face publication date of 2026-05-29.[15] Ranking note: newest release first, then downloads. The ranking admits only image-to-3D models from labs with a published paper or leaderboard record, using Hugging Face’s image-to-3d pipeline tag. Placement does not establish reconstruction quality or measured Apple Silicon performance.

What is the complete ranking?

The order is NVIDIA’s nvidia/dvlt,[15] Tencent’s tencent/HY-World-2.0,[7] NVIDIA’s nvidia/Lyra-2.0,[13] Apple’s Sharp (Apple),[11] Microsoft’s TRELLIS.2-4B (Microsoft),[3] Tencent’s HunyuanWorld-Mirror (Tencent),[9] Facebook’s facebook/VGGT-1B,[5] and Microsoft’s microsoft/TRELLIS-image-large.[1] VGGT-1B and TRELLIS-image-large are established picks that bypass the twelve-month recency gate.[5][1] Both follow the same ordering rule; the exemption does not move either into first place. Download popularity serves only as a tiebreak.

How much memory would quantized weights need?

For NVIDIA’s dvlt, 117 million parameters imply 58.5 MB at 4-bit, estimated (params x 0.5 bytes).[15] Tencent’s HunyuanWorld-Mirror and Facebook’s VGGT-1B each have 1.3 billion parameters, implying 650 MB per model at 4-bit, estimated (params x 0.5 bytes).[9][5] Those calculations describe parameter storage only. They do not establish quantization availability, complete application memory, or compatibility with a particular Mac.

How can I tell whether a model will run well on my Mac?

Check an explicit Apple Silicon execution path and measured application memory before choosing. NVIDIA’s Lyra-2.0 lists full-precision weights of 1.1 GB,[13] while Microsoft’s TRELLIS.2-4B lists 16.2 GB.[3] Those storage figures alone cannot establish Mac compatibility or speed. Apple’s Sharp carries the apple-amlr license,[11] but its organization and license alone do not establish hardware support either.

Which licenses should I check before using a model commercially?

Microsoft’s TRELLIS.2-4B and TRELLIS-image-large use MIT.[3][1] Facebook’s VGGT-1B uses cc-by-nc-4.0,[5] while NVIDIA’s dvlt uses nvidia-oneway-noncommercial.[15] Tencent’s HY-World-2.0 uses tencent-hy-world-2.0-community,[7] and HunyuanWorld-Mirror uses tencent-hunyuanworld-mirror-community.[9] Apple’s Sharp uses apple-amlr.[11] Review the applicable terms against your intended use before integrating a model into a product. Weight availability alone should not be treated as commercial permission.

Do benchmark scores or context limits justify the ordering?

Every candidate carries the status “no public benchmark yet” in this ranking.[1][3][5][7][9][11][13][15] Release order therefore should not be read as a quality score or measured Mac speed comparison. Any model-card benchmark must be labelled “self-reported,” and leaderboard comparisons apply only to models evaluated together. Context limits remain unspecified; no numerical context allowance can be assigned here.

Sources

  1. microsoft/TRELLIS-image-large model card (Hugging Face) — 2026-09-26
  2. Structured 3D Latents for Scalable and Versatile 3D Generation — 2024-12-02
  3. microsoft/TRELLIS.2-4B model card (Hugging Face) — 2026-09-26
  4. Native and Compact Structured Latents for 3D Generation — 2025-12-16
  5. facebook/VGGT-1B model card (Hugging Face) — 2026-09-26
  6. VGGT: Visual Geometry Grounded Transformer — 2025-03-14
  7. tencent/HY-World-2.0 model card (Hugging Face) — 2026-09-26
  8. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds — 2026-04-15
  9. tencent/HunyuanWorld-Mirror model card (Hugging Face) — 2026-09-26
  10. WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting — 2025-10-12
  11. apple/Sharp model card (Hugging Face) — 2026-09-26
  12. Sharp Monocular View Synthesis in Less Than a Second — 2025-12-11
  13. nvidia/Lyra-2.0 model card (Hugging Face) — 2026-09-26
  14. Lyra 2.0: Explorable Generative 3D Worlds — 2026-04-14
  15. nvidia/dvlt model card (Hugging Face) — 2026-09-26
  16. Déjà View: Looping Transformers for Multi-View 3D Reconstruction — 2026-05-28

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog