Quick Answer
Wan2.2-T2V-A14B-Diffusers (Wan-AI) is the top pick as of September 2026 because it is the newest listed release; too few releases meet the past-year window, so the newest available family releases are included, without a verified GPU-fit claim [1][3][4][6]. In order, the ranking is Wan2.2-T2V-A14B-Diffusers (Wan-AI), ContentV-8B (ByteDance), mochi-1-preview (genmo), and CogVideoX-5b (zai-org) [1][3].
Key Takeaways
- Ranking note: scope admits only text-to-video models from labs with a published paper or leaderboard record, using Hugging Face’s text-to-video pipeline tag. The recent-release pool is too small, so newest available entries are included. No public benchmark scores these candidates; ordering is newest release first, then downloads, with popularity used only as a tiebreak. [1][3][4][6]
- Wan-AI’s Wan2.2-T2V-A14B-Diffusers ranks first because its publication date, 2025-07-28, is the newest among these candidates. [1][3][4][6] Its 14.3B parameters imply 7.15 GB of 4-bit weights, estimated (params x 0.5 bytes); the license is Apache-2.0, with no public benchmark yet. [1]
- ByteDance’s ContentV-8B ranks second, published 2025-06-03. Its 8.1B parameters imply 4.05 GB of 4-bit weights, estimated (params x 0.5 bytes); the license is Apache-2.0, with no public benchmark yet. [6]
- genmo’s mochi-1-preview ranks third, published 2024-10-22. Its 10B parameters imply 5 GB of 4-bit weights, estimated (params x 0.5 bytes); the license is Apache-2.0, with no public benchmark yet. [3]
- zai-org’s CogVideoX-5b ranks fourth, published 2024-08-17. Its 5.6B parameters imply 2.8 GB of 4-bit weights, estimated (params x 0.5 bytes); no public benchmark yet. [4] License terms remain unverified. [4]
- Weight estimates alone do not establish GPU fit or usable generation settings. Context limits, quantized runtime memory, and generation speed remain unverified for these candidates, so the ordering does not establish which runs well on the target GPU. [1][3][4][6]
How do these text-to-video models compare on specifications and published benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Wan2.2-T2V-A14B-Diffusers [1] | Wan-AI [1] | 14.3B [1] | 4-bit weights: 7.15 GB, estimated (params x 0.5 bytes) [1]; runtime VRAM: not published | 2025-07-28 [1] | Apache-2.0 [1] | no public benchmark yet |
| ContentV-8B [6] | ByteDance [6] | 8.1B [6] | 4-bit weights: 4.05 GB, estimated (params x 0.5 bytes) [6]; runtime VRAM: not published | 2025-06-03 [6] | Apache-2.0 [6] | no public benchmark yet |
| mochi-1-preview [3] | genmo [3] | 10B [3] | 4-bit weights: 5 GB, estimated (params x 0.5 bytes) [3]; runtime VRAM: not published | 2024-10-22 [3] | Apache-2.0 [3] | no public benchmark yet |
| CogVideoX-5b [4] | zai-org [4] | 5.6B [4] | 4-bit weights: 2.8 GB, estimated (params x 0.5 bytes) [4]; runtime VRAM: not published | 2024-08-17 [4] | — | no public benchmark yet |
Which open-weight text-to-video models should you consider for local generation?
1. Wan2.2-T2V-A14B-Diffusers
Wan2.2-T2V-A14B-Diffusers by Wan-AI ranks first because its Hugging Face publication date of July 28, 2025 is the newest among the admitted candidates.[1][3][4][6] The ordering is newest release first, then downloads; no public benchmark yet establishes its quality advantage. The ranking admits only text-to-video models from labs with a published paper or leaderboard record. Too few recent releases qualify, so the newest available candidates are included despite falling outside the requested release window.[1][3][4][6]
Wan-AI lists 14.3 billion parameters, F32 weight files totaling 126.2 GB, and an Apache-2.0 license.[1] At 4-bit precision, parameter storage alone is approximately 7.15 GB, estimated (params x 0.5 bytes) from the cited parameter count.[1] That calculation is a planning estimate, not a measured inference footprint or confirmation that a compatible quantized implementation is available. The repository’s listed weight-file size should therefore not be treated as the memory requirement for a quantized run.
For local use, treat Wan as a candidate for evaluating the newest admitted generation, rather than a demonstrated benchmark winner.[1][3][4][6] Hardware selection needs a verified runtime configuration: the parameter-storage estimate alone cannot establish GPU fit, required system RAM, or usable generation settings. The practical caveat is that context limits, runtime memory and generation speed remain unestablished here, so the ranking supports an evaluation priority rather than a confirmed hardware recommendation.
2. ContentV-8B
ContentV-8B by ByteDance ranks second under the ordering rule: newest release first, then downloads.[1][3][4][6] Its Hugging Face publication date is June 3, 2025, placing it behind Wan-AI’s Wan2.2-T2V-A14B-Diffusers, published July 28, 2025.[6][1] ContentV-8B has no public benchmark yet, so its position reflects release order rather than a demonstrated quality advantage.[6][7] ByteDance describes the model’s training approach in ContentV: Efficient Training of Video Generation Models with Limited Compute.[7]
The model has 8.1 billion parameters [6], giving a 4-bit weight footprint of approximately 4.05 GB—estimated (params x 0.5 bytes).[6] The repository lists 27.8 GB of BF16 weights and an Apache-2.0 license.[6] The calculated footprint covers parameter storage only; it should not be read as a measured runtime requirement or confirmation that a complete generation workflow fits your GPU. Quantized execution and total memory use still require validation before choosing local hardware.
ContentV-8B is a candidate for local text-to-video experimentation when an Apache-2.0 license suits your project.[6] The practical caveat is that the weight estimate cannot establish GPU capacity, host RAM requirements, supported context, or generation speed. Plan an evaluation around your intended output settings before committing hardware: the published figures do not establish that the model runs comfortably within your memory budget.
3. mochi-1-preview
mochi-1-preview by genmo ranks third under the ordering rule “newest release first, then downloads,” behind Wan-AI’s Wan2.2-T2V-A14B-Diffusers and ByteDance’s ContentV-8B.[1][6][3] Its first Hugging Face publication was October 22, 2024, placing it outside the requested recent-release window; the shortlist includes older releases because too few qualify.[3] Its benchmark status is “no public benchmark yet,” so its position does not establish a quality or speed advantage.
The model has 10 billion parameters, lists 133.4 GB of F32 weights, and uses the Apache-2.0 license.[3] At 4-bit precision, weights alone would occupy approximately 5 GB—estimated (params x 0.5 bytes), using the cited 10-billion-parameter count.[3] Treat that calculation as a weight-storage estimate, not a complete GPU memory requirement. The listed repository weight size also does not establish how much memory a running, quantized pipeline needs.
Consider mochi-1-preview for local text-to-video experimentation when an Apache-2.0 license suits your project.[3] The practical caveat is hardware uncertainty: the weight calculation does not demonstrate that the complete pipeline fits your GPU or runs at an acceptable speed. Before committing to a local setup, verify quantization support, total inference memory, supported generation settings, and runtime for the implementation you intend to use.
4. CogVideoX-5b
CogVideoX-5b by zai-org ranks fourth under the ordering rule: newest release first, then downloads.[1][3][4][6] Its Hugging Face publication date is August 17, 2024, making it the oldest listed release.[1][3][4][6] CogVideoX-5b has no public benchmark yet for this comparison, so its position reflects release order rather than demonstrated video quality. The accompanying paper describes text-to-video diffusion models with an expert transformer.[5]
CogVideoX-5b has 5.6 billion parameters, with listed BF16 weights totaling 21.5 GB.[4] Quantized weight storage would be approximately 2.8 GB, estimated (params x 0.5 bytes) from that parameter count.[4] Treat that calculation as a weight-storage budget, not a measured runtime requirement. The published figures do not establish total GPU memory or system RAM needed for local generation, so they cannot confirm a complete hardware configuration.
Consider CogVideoX-5b for exploratory local text-to-video work where you can validate memory use and output quality before committing to a workflow. The practical caveat is that the quantized estimate does not demonstrate a working quantized deployment. Verify the implementation’s memory requirements, supported clip settings and license terms before choosing hardware or using generated video in production.
How do you estimate quantized weight memory from parameter counts?
Estimate quantized weight memory by multiplying the parameter count by the storage per quantized parameter: for four-bit weights, use estimated (params x 0.5 bytes), with decimal gigabytes as the output unit.[1]
Wan-AI’s Wan2.2-T2V-A14B-Diffusers has 14.3 billion parameters: 7.15 GB estimated (params x 0.5 bytes) for quantized weights.[1] ByteDance’s ContentV-8B has 8.1 billion parameters: 4.05 GB estimated (params x 0.5 bytes).[6]
genmo’s mochi-1-preview has 10 billion parameters: 5 GB estimated (params x 0.5 bytes).[3] zai-org’s CogVideoX-5b has 5.6 billion parameters: 2.8 GB estimated (params x 0.5 bytes).[4] Apply the same calculation to each candidate so the estimates remain comparable.
Treat each result as a weight-storage planning figure. A parameter-based calculation does not establish total GPU memory consumption, confirm that a compatible quantized implementation exists, or demonstrate that generation runs well on your hardware. Avoid turning a weight estimate into a GPU-fit claim.
For a deployment decision, pair the estimate with documented runtime memory requirements for the intended configuration. Keep supported context and generation performance separate from weight storage; neither follows from parameter count. Where runtime measurements are unavailable, report the weight estimate and leave actual fit unconfirmed.
Does estimated weight memory establish whether a model fits on your GPU?
No—estimated weight memory alone does not establish whether a model fits on your GPU. A parameter-based calculation estimates storage for weights at an assumed precision; it does not establish total memory during generation, supported quantization, or usable generation settings.
Wan-AI’s Wan2.2-T2V-A14B-Diffusers has 14.3 billion parameters, giving approximately 7.15 GB of weight memory at four-bit precision, estimated (params x 0.5 bytes).[1] ByteDance’s ContentV-8B has 8.1 billion parameters, giving approximately 4.05 GB at four-bit precision, estimated (params x 0.5 bytes).[6] Neither calculation establishes that a working local configuration will stay within your GPU’s memory budget.
Published weight sizes describe a different storage basis. Wan-AI lists F32 weights of 126.2 GB for Wan2.2-T2V-A14B-Diffusers; ByteDance lists BF16 weights of 27.8 GB for ContentV-8B.[1][6] Those figures should not be substituted for measured runtime memory or treated as confirmation that quantized execution is available.
Use estimated weight memory as an initial screening tool. A fit claim needs a concrete configuration and its total memory requirement, including allocations beyond weights. Without that information, the practical conclusion is “weight storage may fit; generation remains unverified.” Weight estimates also provide no evidence of generation speed, output quality, or supported context.
Which licenses do these text-to-video models use?
Apache-2.0 is the confirmed license for Wan2.2-T2V-A14B-Diffusers (Wan-AI) [1], ContentV-8B (ByteDance) [6], and mochi-1-preview (genmo) [3]. The license for CogVideoX-5b (zai-org) remains unconfirmed; do not assume it carries the same terms as the other candidates [4].
For a shortlist organized by license, group the Apache-2.0 candidates together [1][6][3]. Keep CogVideoX-5b pending a license check [4]. An unconfirmed license is a reason to investigate before adoption, rather than a basis for declaring the model either unrestricted or prohibited.
Before adopting a model, read the license attached to the exact repository and revision you intend to download. Record that revision alongside the license text in your deployment notes. If you use a converted or quantized checkpoint, check its accompanying license and notices as well.
Match that review to your intended workflow: local experimentation, an internal service, a customer-facing application, or redistribution of a checkpoint. Check separately whether the repository describes terms for generated outputs. Use the confirmed license labels to narrow your options, then resolve any missing terms before committing the model to a project.
Frequently Asked Questions
Which model should I consider first?
Wan-AI’s Wan2.2-T2V-A14B-Diffusers (Wan-AI) takes first place because its publication date is the newest among the admitted candidates.[1][6][3][4] The remaining order is ByteDance’s ContentV-8B (ByteDance),[6] genmo’s mochi-1-preview (genmo),[3] and zai-org’s CogVideoX-5b (zai-org).[4] Treat that order as a shortlist for evaluation: the leading position reflects release recency, while generation quality, runtime memory and speed remain unverified in this comparison.
Why does the ranking include older releases?
The recency gate is relaxed because the family has too few recent releases; the newest available candidates are listed, including releases published in July and June 2025 and October and August 2024.[1][6][3][4] Ranking note: no public benchmark scores any candidate, so the ordering is newest release first, then downloads. Scope: this ranking admits only text to video models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: text-to-video).
How much memory would quantized weights need?
For 4-bit weight storage, the calculations are: Wan2.2-T2V-A14B-Diffusers, 14.3B parameters, 7.15 GB estimated (params x 0.5 bytes);[1] ContentV-8B, 8.1B parameters, 4.05 GB estimated (params x 0.5 bytes);[6] mochi-1-preview, 10B parameters, 5 GB estimated (params x 0.5 bytes);[3] CogVideoX-5b, 5.6B parameters, 2.8 GB estimated (params x 0.5 bytes).[4] Each calculation describes parameter storage alone, not a measured generation workload or a verified GPU fit.
Do those weight estimates guarantee that a model will run on my GPU?
No verified GPU-fit conclusion follows from those calculations. Published repository weight sizes describe a different quantity: ContentV-8B lists 27.8 GB of BF16 weights,[6] while CogVideoX-5b lists 21.5 GB of BF16 weights.[4] Neither figure establishes runtime memory for a quantized setup. Before committing to a model, seek a reproducible configuration with measured peak memory and generation time for your intended workload.
Which models have a clearly identified license?
Wan2.2-T2V-A14B-Diffusers,[1] ContentV-8B,[6] and mochi-1-preview[3] list the Apache-2.0 license. CogVideoX-5b’s license is unspecified in this comparison, so check its current license terms before adopting it. Keep license review separate from hardware selection: an identified license does not establish that a particular local configuration will fit or perform adequately.
Where are the benchmark scores and context limits?
Every candidate is marked “no public benchmark yet” for this ranking; no comparable public score establishes a quality winner. Context limits are also unspecified in this comparison. The Wan,[2] CogVideoX,[5] and ContentV[7] papers provide technical reading, but their existence does not establish a shared benchmark result. Any model-card performance claim should be labeled self-reported and kept separate from public leaderboard comparisons.
Sources
- Wan-AI/Wan2.2-T2V-A14B-Diffusers model card (Hugging Face) — 2026-09-25
- Wan: Open and Advanced Large-Scale Video Generative Models — 2025-03-26
- genmo/mochi-1-preview model card (Hugging Face) — 2026-09-25
- zai-org/CogVideoX-5b model card (Hugging Face) — 2026-09-25
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer — 2024-08-12
- ByteDance/ContentV-8B model card (Hugging Face) — 2026-09-25
- ContentV: Efficient Training of Video Generation Models with Limited Compute — 2025-06-05