Quick Answer
Wan2.2-T2V-A14B-Diffusers (Wan-AI) is the top pick as of September 2026 because it is the newest listed release among eligible text-to-video models from labs with a published paper or leaderboard record, although its fit on a 24 GB GPU remains unverified [1][2][3][4][5][6][7]. In order, the ranking is Wan2.2-T2V-A14B-Diffusers (Wan-AI), ContentV-8B (ByteDance), stepvideo-t2v (stepfun-ai), and mochi-1-preview (genmo) [1][2].
Key Takeaways
- Ranking note: this ranking admits only text-to-video models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: text-to-video). Too few candidates meet the recency window, so the newest available releases are included. No public benchmark scores any candidate; ordering is newest release first, then downloads. [1][3][4][6]
- Wan-AI’s Wan2.2-T2V-A14B-Diffusers ranks first because its publication date, 2025-07-28, is the newest among the admitted candidates. Apache-2.0; no public benchmark yet. Estimated 4-bit weight memory: 7.15 GB, estimated (params x 0.5 bytes) from 14.3B parameters. [1]
- ByteDance’s ContentV-8B ranks second by release date. Apache-2.0; no public benchmark yet. Estimated 4-bit weight memory: 4.05 GB, estimated (params x 0.5 bytes) from 8.1B parameters. [4]
- stepfun-ai’s stepvideo-t2v ranks third by release date. MIT; no public benchmark yet. Estimated 4-bit weight memory: 14.65 GB, estimated (params x 0.5 bytes) from 29.3B parameters. [6]
- genmo’s mochi-1-preview ranks fourth by release date. Apache-2.0; no public benchmark yet. Estimated 4-bit weight memory: 5 GB, estimated (params x 0.5 bytes) from 10B parameters. [3]
- GPU fit remains unverified: parameter-based weight estimates do not establish full-pipeline memory requirements, supported context capacity or practical generation speed. Treat the ordering as a release-based shortlist, not a demonstrated performance ranking. [1][3][4][6]
How do these text-to-video models compare on parameters, estimated weight memory, licenses and public benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Wan2.2-T2V-A14B-Diffusers [1] | Wan-AI [1] | 14.3B [1] | 4-bit weights: 7.15 GB, estimated (params × 0.5 bytes) from 14.3B params [1]; runtime VRAM: not published | 2025-07-28 [1] | Apache-2.0 [1] | no public benchmark yet |
| ContentV-8B [4] | ByteDance [4] | 8.1B [4] | 4-bit weights: 4.05 GB, estimated (params × 0.5 bytes) from 8.1B params [4]; runtime VRAM: not published | 2025-06-03 [4] | Apache-2.0 [4] | no public benchmark yet |
| stepvideo-t2v [6] | stepfun-ai [6] | 29.3B [6] | 4-bit weights: 14.65 GB, estimated (params × 0.5 bytes) from 29.3B params [6]; runtime VRAM: not published | 2025-02-14 [6] | MIT [6] | no public benchmark yet |
| mochi-1-preview [3] | genmo [3] | 10B [3] | 4-bit weights: 5 GB, estimated (params × 0.5 bytes) from 10B params [3]; runtime VRAM: not published | 2024-10-22 [3] | Apache-2.0 [3] | no public benchmark yet |
Which text-to-video models should you consider for local generation?
1. Wan2.2-T2V-A14B-Diffusers
Wan2.2-T2V-A14B-Diffusers by Wan-AI ranks first because its release is the newest among the eligible candidates.[1][3][4][6] Ranking note: no public benchmark scores any candidate, so the order is newest release first, then downloads.[1][3][4][6] The release window has too few qualifying models, so the newest available candidates are included.[1][3][4][6] The ranking admits only text-to-video models from labs with a published paper or leaderboard record.
Wan-AI published the model on Hugging Face on 2025-07-28 and lists 14.3 billion parameters, F32 weights totaling 126.2 GB, and an Apache-2.0 license.[1] Weight memory at 4-bit is 7.15 GB, estimated (params x 0.5 bytes) from the listed 14.3 billion parameters.[1] The quantized estimate and the published repository weight total describe different storage assumptions; neither establishes the memory required for a complete local generation run.
Use the model as a candidate for local text-to-video experimentation when release recency and Apache-2.0 licensing matter.[1] Hardware planning remains the caveat: the weight estimate alone cannot confirm GPU fit or system-RAM requirements. A deployment decision still needs a verified runtime configuration, including its quantization support and total memory use. Benchmark status is no public benchmark yet; the ranking does not establish superior visual quality or generation speed.[1]
2. ContentV-8B
ContentV-8B by ByteDance ranks second under the ordering rule: newest release first, then downloads. Its first Hugging Face publication date is June 3, 2025 [4], behind Wan-AI’s Wan2.2-T2V-A14B-Diffusers, published July 28, 2025 [1]. ContentV-8B has no public benchmark yet, so its position does not establish a quality or speed advantage. The candidate pool has too few recent releases; the ranking therefore includes the newest available models [1][3][4][6].
ContentV-8B has 8.1 billion parameters, lists 27.8 GB of BF16 weights, and uses the Apache-2.0 license [4]. Quantized weight storage would be approximately 4.05 GB, estimated (params x 0.5 bytes) from that parameter count [4]. Treat that calculation as a weight-storage estimate, not a measured GPU-memory requirement. A local deployment still needs a compatible runtime and quantization path; the estimate does not establish that the complete generation pipeline fits your GPU.
Consider ContentV-8B for exploratory local text-to-video work where the Apache-2.0 license suits your project [4]. ByteDance also documents the model in “ContentV: Efficient Training of Video Generation Models with Limited Compute” [5]. The practical caveat is deployment uncertainty: neither a verified full-pipeline memory requirement nor a published benchmark score establishes how well this candidate will run on your hardware.
3. stepvideo-t2v
stepvideo-t2v by stepfun-ai ranks third[1][3][4][6] because its release falls between ByteDance’s ContentV-8B and genmo’s mochi-1-preview.[3][4][6] The ordering basis is newest release first, then downloads; no public benchmark scores any candidate, so stepvideo-t2v has “no public benchmark yet.” The pool has too few recent releases, so the newest available family releases are included.[1][3][4][6] Scope: this ranking admits only text-to-video models from labs with a published paper or leaderboard record, using Hugging Face’s text-to-video pipeline tag.
Published on February 14, 2025, stepvideo-t2v has 29.3 billion parameters, lists 100.9 GB of BF16 weights, and uses the MIT license.[6] Its accompanying paper is the “Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.”[7] Quantized weight memory is approximately 14.65 GB at 4-bit—estimated (params x 0.5 bytes), using the stated 29.3 billion parameters.[6] That calculation covers parameter storage; it does not establish the memory needed for a complete generation run.
Consider stepvideo-t2v for local experimentation when MIT licensing suits your project.[6] The practical caveat is hardware uncertainty: the weight estimate does not confirm GPU fit or system RAM requirements. Validate the intended quantized implementation’s total memory use before choosing hardware. Context limits and measured runtime memory are unspecified here, so neither can support a stronger deployment recommendation.
4. mochi-1-preview
mochi-1-preview by genmo ranks fourth under the ordering rule: newest release first, then downloads, because no public benchmark scores any candidate. Its first Hugging Face publication was 2024-10-22, earlier than the other ranked releases.[3][6][4][1] Its benchmark status is “no public benchmark yet”; its position reflects publication date rather than a demonstrated gap in video quality. The release falls outside the requested recency window and appears as an available-family exception.[3]
The model has 10B parameters and an Apache-2.0 license.[3] Its listed F32 weights total 133.4 GB.[3] For quantized hardware planning, the cited 10B parameters imply approximately 5 GB of 4-bit weights, estimated (params x 0.5 bytes).[3] That calculation describes parameter storage alone. A weights-only estimate does not establish the GPU capacity or system RAM needed for a complete local generation run.
Consider mochi-1-preview for local experimentation where Apache-2.0 licensing is useful.[3] The practical caveat is hardware uncertainty: the parameter calculation does not verify a working quantized configuration or establish that inference fits the target GPU. Context limits and published benchmark scores are unavailable here, so neither provides a supported reason to choose it over the higher-ranked candidates.
Why does the ranking use newest release first, then downloads?
The ranking uses newest release first, then downloads because no public benchmark scores these candidates against one another. Each therefore carries the label “no public benchmark yet”; the ordering does not establish a measured quality or speed advantage.
The ranking admits only text-to-video models from labs with a published paper or leaderboard record, using Hugging Face’s text-to-video pipeline tag. The eligible families have too few recent releases, so the newest available releases are included despite falling outside the requested release window.[1][3][4][6]
Wan2.2-T2V-A14B-Diffusers (Wan-AI) ranks first because its publication date is the newest among the admitted candidates.[1][3][4][6] ContentV-8B (ByteDance) follows, then stepvideo-t2v (stepfun-ai), and finally mochi-1-preview (genmo), following their publication dates.[4][6][3]
Downloads break ties after release recency; popularity alone cannot establish generation quality, memory efficiency or suitability for your hardware. A leaderboard can order only models it actually compares, and an older scored release cannot displace a newer unscored release under this ranking’s rules. Model-card benchmark results, when used, must be identified as self-reported.
Treat the order as a starting point for evaluation. Parameter-based weight estimates do not establish total runtime memory requirements. License terms, available context information and documented execution requirements still need separate consideration before choosing a model for local use.
How should you estimate quantized weight memory without claiming a model fits your GPU?
Estimate quantized weight memory from the cited parameter count, label the result as an estimate, and keep that calculation separate from any claim that the model fits your GPU. A weight estimate answers how much storage the parameters would occupy under an assumed representation; a fit claim requires evidence about the complete running configuration.
For Wan-AI’s Wan2.2-T2V-A14B-Diffusers (Wan-AI), the cited 14.3B parameters give 7.15 GB of weight memory at 4-bit, estimated (params × 0.5 bytes).[1] For ByteDance’s ContentV-8B (ByteDance), the cited 8.1B parameters give 4.05 GB at 4-bit, estimated (params × 0.5 bytes).[4]
For stepfun-ai’s stepvideo-t2v (stepfun-ai), the cited 29.3B parameters give 14.65 GB at 4-bit, estimated (params × 0.5 bytes).[6] For genmo’s mochi-1-preview (genmo), the cited 10B parameters give 5 GB at 4-bit, estimated (params × 0.5 bytes).[3] Each calculation assumes the cited parameter count covers the weights you intend to quantize.
Keep those estimates in a column labeled “estimated quantized weights,” rather than “required VRAM.” Before asserting GPU fit, establish which components are resident, the quantization implementation, runtime allocations, offloading settings, and the intended generation configuration. Without a documented peak-memory measurement for that configuration, leave total VRAM requirements and GPU fit unverified. Numerical headroom alone does not establish that a generation will complete.
What context and runtime memory details remain unverified?
Prompt context limits, video length limits and peak runtime GPU memory remain unverified for the listed candidates; the available specifications do not establish a complete configuration that fits the target GPU.[1][3][4][6]
Context needs separate treatment for text and video. A practical comparison would specify the prompt token limit, how longer prompts are handled, supported frame counts, output resolution and frame rate. Those details remain unverified here, so no candidate can be recommended on the basis of a demonstrated context advantage.
Parameter counts permit a narrower calculation. Wan-AI’s Wan2.2-T2V-A14B-Diffusers (Wan-AI) lists 14.3 billion parameters:[1] its weight-only footprint is 7.15 GB, estimated (params x 0.5 bytes).[1] ByteDance’s ContentV-8B (ByteDance) lists 8.1 billion parameters:[4] its weight-only footprint is 4.05 GB, estimated (params x 0.5 bytes).[4] Neither calculation establishes peak runtime memory or confirms a working quantized configuration.
Runtime verification still needs a named quantization format, software configuration, output settings, peak GPU memory, host RAM requirement and generation time. Any reliance on CPU offloading should be explicit. The listed weight sizes and licenses do not answer those questions.[1][3][4][6] Treat the candidates as options to validate locally, with GPU fit and throughput still unresolved; a parameter-based estimate is insufficient grounds for a purchase or deployment decision.
Frequently Asked Questions
Which text-to-video model should I try first?
Wan-AI’s Wan2.2-T2V-A14B-Diffusers takes first place because its July 28, 2025 release is the newest among the eligible candidates.[1][3][4][6] The recommendation follows release order; Wan2.2 has no public benchmark yet. Its 14.3B parameter count and Apache-2.0 license provide useful starting information, but neither establishes runtime memory requirements or generation speed on your GPU.[1]
How are the models ranked?
The order is Wan-AI’s Wan2.2-T2V-A14B-Diffusers, ByteDance’s ContentV-8B, stepfun-ai’s stepvideo-t2v, then genmo’s mochi-1-preview.[1][4][6][3] The ordering basis is newest release first, then downloads; no public benchmark scores any candidate. Downloads serve only as a tiebreak. Scope: this ranking admits only text to video models from labs with a published paper or leaderboard record, using Hugging Face’s text-to-video pipeline tag.
Are all the recommendations recent releases?
No. The category has too few releases within the requested window, so the newest available family releases are included.[1][3][4][6] Even Wan2.2 was published on July 28, 2025.[1] The established picks by downloads are Wan2.2, Mochi and StepVideo.[1][3][6] Established picks bypass the recency gate but follow the same ordering rule; the exemption alone never earns first place. Wan2.2 leads on release date.[1][4][6][3]
How much memory would quantized weights require?
For 4-bit weights, Wan2.2’s 14.3B parameters[1] imply 7.15 GB, estimated (params x 0.5 bytes).[1] ContentV’s 8.1B parameters[4] imply 4.05 GB, estimated (params x 0.5 bytes).[4] StepVideo’s 29.3B parameters[6] imply 14.65 GB, estimated (params x 0.5 bytes).[6] Mochi’s 10B parameters[3] imply 5 GB, estimated (params x 0.5 bytes).[3] Those calculations cover parameter storage, not total runtime memory or confirmed quantization support.
Do those weight estimates mean the models will run well on my GPU?
Weight arithmetic does not establish a working local configuration. The published weight listings describe F32 or BF16 assets, rather than measured quantized runtime configurations.[1][3][4][6] Context limits, generation settings and runtime memory measurements are unspecified here. Every candidate has no public benchmark yet. Treat the ranking as a shortlist for compatibility checks; it cannot establish comparative speed, output quality or a guaranteed hardware fit.
Which licenses do the models use?
Wan2.2, ContentV and Mochi use the Apache-2.0 license.[1][4][3] StepVideo uses the MIT license.[6] Those license declarations help evaluate a local deployment, but they do not establish hardware compatibility or generation performance. Keep licensing and deployment checks separate: select the model under its applicable terms, then verify that your intended runtime supports its weights and generation workflow.
Sources
- Wan-AI/Wan2.2-T2V-A14B-Diffusers model card (Hugging Face) — 2026-09-25
- Wan: Open and Advanced Large-Scale Video Generative Models — 2025-03-26
- genmo/mochi-1-preview model card (Hugging Face) — 2026-09-25
- ByteDance/ContentV-8B model card (Hugging Face) — 2026-09-25
- ContentV: Efficient Training of Video Generation Models with Limited Compute — 2025-06-05
- stepfun-ai/stepvideo-t2v model card (Hugging Face) — 2026-09-25
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model — 2025-02-14