Quick Answer
LTX-2 from Lightricks is the top pick to evaluate for a 24 GB GPU as of September 2026 because it is the newest eligible release, with no public benchmark yet [1][3][5]. In order, the ranking is LTX-2 (Lightricks), Video-As-Prompt-Wan2.1-14B (ByteDance), and Video-As-Prompt-CogVideoX-5B (ByteDance) [1][3].
Key Takeaways
- Ranking note: newest release first, then downloads. Scope: this ranking admits only image-to-video models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: image-to-video).[1][2][3][4][5]
- Lightricks’ LTX-2 ranks first because its Hugging Face publication date is 2026-01-03,[1] later than both ByteDance candidates’ 2025-10-22 publication date.[3][5]
- ByteDance’s Video-As-Prompt-Wan2.1-14B follows, then ByteDance’s Video-As-Prompt-CogVideoX-5B: their shared publication date makes downloads the tiebreaker, with 47 versus 32 downloads over the last 30 days.[3][5]
- Quantized weight memory is estimated (params x 0.5 bytes): LTX-2’s 18.9B parameters imply 9.45 GB;[1] Video-As-Prompt-Wan2.1-14B’s 20.7B imply 10.35 GB;[3] Video-As-Prompt-CogVideoX-5B’s 11.1B imply 5.55 GB.[5] Those estimates cover weights alone and do not establish runtime fit.
- LTX-2 uses the ltx-2-community-license-agreement;[1] both ByteDance models use Apache-2.0.[3][5] Check the applicable terms before choosing a model for deployment.
- All candidates have no public benchmark yet; context limits and measured runtime memory are unspecified, so the ordering does not establish generation quality, speed or GPU fit.[1][3][5]
How do these image-to-video models compare on memory, context, licenses and benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| Lightricks/LTX-2 [1] | Lightricks [1] | 18.9B [1] | 4-bit weights: 9.45 GB estimated (params x 0.5 bytes), from 18.9B params [1]; runtime VRAM: not published | 2026-01-03 [1] | ltx-2-community-license-agreement [1] | no public benchmark yet |
| ByteDance/Video-As-Prompt-Wan2.1-14B [3] | ByteDance [3] | 20.7B [3] | 4-bit weights: 10.35 GB estimated (params x 0.5 bytes), from 20.7B params [3]; runtime VRAM: not published | 2025-10-22 [3] | Apache-2.0 [3] | no public benchmark yet |
| ByteDance/Video-As-Prompt-CogVideoX-5B [5] | ByteDance [5] | 11.1B [5] | 4-bit weights: 5.55 GB estimated (params x 0.5 bytes), from 11.1B params [5]; runtime VRAM: not published | 2025-10-22 [5] | Apache-2.0 [5] | no public benchmark yet |
Which open-weight image-to-video models should you consider for local generation?
1. LTX-2
LTX-2 by Lightricks ranks first because its Hugging Face publication date places it ahead of both eligible ByteDance releases.[1][3][5] Ranking note: no public benchmark scores any candidate, so the ordering is newest release first, then downloads.[1][3][5] The ranking admits only image-to-video models from labs with a published paper or leaderboard record, using Hugging Face’s image-to-video pipeline tag.[1][2][3][4][5] LTX-2 has no public benchmark yet; its position does not establish superior video quality.[1]
Lightricks lists 18.9B parameters and the ltx-2-community-license-agreement license.[1] The model was first published on Hugging Face on January 3, 2026.[1] Weight storage at four-bit precision is approximately 9.45 GB, estimated (params x 0.5 bytes) from the cited 18.9B parameter count.[1] That calculation covers parameter storage only. A complete local hardware requirement cannot be established from the published figures, so treat the estimate as a starting point for evaluation rather than a confirmed GPU fit.
Consider LTX-2 for local image-to-video evaluation when release recency and a joint audio-visual model are relevant to your project; Lightricks describes that model scope in its paper.[2] The practical caveat is the gap between estimated weight storage and demonstrated operation: the listed information provides no verified quantized runtime memory, context limit, or GPU configuration.[1] Review the community license before committing to a deployment.[1]
2. Video-As-Prompt-Wan2.1-14B
Video-As-Prompt-Wan2.1-14B by ByteDance ranks second: LTX-2 by Lightricks has a later release date, while ByteDance’s Video-As-Prompt-CogVideoX-5B shares its publication date but has fewer downloads.[1][3][5] Ranking note: newest release first, then downloads. The ranking admits only image-to-video models from labs with a published paper or leaderboard record, using Hugging Face’s image-to-video pipeline tag. Video-As-Prompt-Wan2.1-14B has no public benchmark yet; its position therefore does not establish a quality or speed advantage.
The model lists 20.7B parameters, BF16 weights totaling 65.8 GB, and the Apache-2.0 license.[3] Hugging Face first publication was 2025-10-22, with 47 downloads over the preceding 30 days as of 2026-09-25.[3] Weight memory at 4-bit is approximately 10.35 GB, estimated (params x 0.5 bytes) from the cited 20.7B parameter count.[3] Treat that calculation as a weight-only planning figure, rather than a measured requirement for running the model locally.
Consider the model for local image-to-video projects exploring semantic control, the focus of ByteDance’s “Video-As-Prompt: Unified Semantic Control for Video Generation” paper.[4] Apache-2.0 provides an explicit license to review against your project’s requirements.[3] The deployment caveat is hardware uncertainty: the weight estimate cannot establish total runtime memory or guarantee GPU fit. Verify context limits, quantization support and measured memory use in your intended runtime before committing hardware.
3. Video-As-Prompt-CogVideoX-5B
Video-As-Prompt-CogVideoX-5B by ByteDance ranks third because the ordering is newest release first, then downloads.[1][3][5] The ranking admits only image-to-video models from labs with a published paper or leaderboard record, using Hugging Face’s image-to-video pipeline tag. ByteDance published this model on October 22, 2025,[5] alongside ByteDance’s Video-As-Prompt-Wan2.1-14B.[3] The download tiebreak places it behind its sibling: 32 versus 47 downloads over the last 30 days as of September 25, 2026.[3][5]
The model has 11.1 billion parameters and a listed BF16 weight footprint of 32.6 GB.[5] Parameter-based four-bit weight storage is approximately 5.55 GB, estimated (params x 0.5 bytes) from the cited 11.1 billion parameters.[5] Treat that calculation as a weight-storage budget, not a verified GPU-memory requirement. The Apache-2.0 license makes it a candidate to evaluate when that license is a project requirement.[5]
Consider it for image-to-video experimentation where licensing matters more than a demonstrated benchmark advantage. Benchmark status: no public benchmark yet. The central caveat is hardware uncertainty: a verified local GPU configuration, system-RAM requirement, and context limit remain unspecified. Before committing hardware, require a working configuration and measured runtime memory usage; the weight estimate alone does not establish that the model fits or runs well.
How should you estimate quantized weight memory before choosing a model?
Estimate quantized weight memory by multiplying the cited parameter count by the storage per parameter: for 4-bit weights, use estimated (params x 0.5 bytes), and treat the result as a weight budget rather than a GPU fit guarantee.[1][3][5]
Lightricks’ LTX-2 has 18.9 billion parameters: 9.45 GB estimated (params x 0.5 bytes) for quantized weights.[1] ByteDance’s Video-As-Prompt-Wan2.1-14B has 20.7 billion parameters: 10.35 GB estimated (params x 0.5 bytes).[3] ByteDance’s Video-As-Prompt-CogVideoX-5B has 11.1 billion parameters: 5.55 GB estimated (params x 0.5 bytes).[5] Use the parameter counts rather than interpreting the model names as complete memory specifications.
The calculation assumes 4-bit storage at 0.5 bytes per parameter, compared with BF16 storage at 2 bytes per parameter: approximately one quarter of the weight memory.[1][3][5] Keep that arithmetic separate from the listed BF16 weight totals; a repository’s weight total should not automatically become your runtime memory estimate.
Before choosing a model, ask whether the intended quantized implementation documents total GPU memory use for your intended generation settings. Leave room beyond the calculated weights, and check what the loading configuration actually keeps on the GPU. A weight estimate alone does not establish that a complete generation workload fits.
What context and runtime memory details remain unverified?
Supported context limits, peak runtime memory and practical GPU fit remain unverified for the listed checkpoints.[1][3][5] Context verification still needs explicit limits for frame count, spatial resolution, clip duration and conditioning inputs. No supported operating range for those settings is established here.[1][3][5]
LTX-2 (Lightricks) lists 18.9B parameters: 9.45 GB of 4-bit weights, estimated (params x 0.5 bytes).[1] Video-As-Prompt-Wan2.1-14B (ByteDance) lists 20.7B parameters: 10.35 GB of 4-bit weights, estimated (params x 0.5 bytes).[3] Video-As-Prompt-CogVideoX-5B (ByteDance) lists 11.1B parameters: 5.55 GB of 4-bit weights, estimated (params x 0.5 bytes).[5]
Those calculations cover parameter storage only. A usable runtime comparison still needs measured peak GPU memory and host RAM, with the quantization format, execution software, offloading settings and output configuration disclosed. Availability of a working quantized checkpoint and compatibility with a local execution path also remain unverified.[1][3][5]
Practical fit therefore remains an open question for every candidate. Verification should pair each memory measurement with its frame count, resolution, conditioning configuration and generation time. Until those details are established, treat the weight estimates as a starting point for evaluation; none establishes a supported context limit, a complete runtime memory budget or acceptable generation speed.[1][3][5]
How do the model licenses compare?
ByteDance’s Video-As-Prompt-Wan2.1-14B and Video-As-Prompt-CogVideoX-5B both list Apache-2.0, while Lightricks’ LTX-2 lists the ltx-2-community-license-agreement.[3][5][1] The practical distinction is a shared license identifier for the ByteDance models versus a separate, model-specific agreement for the Lightricks model.[3][5][1]
License choice does not distinguish the ByteDance candidates at the identifier level: both carry the same designation.[3][5] If your organization already has an Apache-2.0 review process, use that as the starting point for evaluating either repository. Check the actual license text and accompanying notices against your intended use before approving a deployment.
Lightricks’ LTX-2 requires reviewing its named community agreement separately.[1] Do not treat “community” as a statement about commercial permissions, redistribution rights or usage restrictions. For an internal tool, a customer-facing service or a redistributed package, check the agreement against the specific activity you plan to undertake.
For a deployment decision, record the applicable license text alongside the downloaded model revision. Keep permission to use and distribute the model separate from hardware suitability: a license designation does not establish whether a local workflow will run successfully. Resolve both questions before committing engineering time to integration.
Frequently Asked Questions
How are the models ranked?
Ranking note: newest release first, then downloads. The ranking admits only image-to-video models from labs with a published paper or leaderboard record, using Hugging Face’s image-to-video pipeline tag.[1][2][3][4][5] The order is LTX-2 (Lightricks),[1] Video-As-Prompt-Wan2.1-14B (ByteDance),[3] then Video-As-Prompt-CogVideoX-5B (ByteDance).[5] No public benchmark scores these candidates against each other; downloads break the tied ByteDance publication dates.[3][5]
Why does the first-ranked model lead?
LTX-2 (Lightricks) ranks first because its Hugging Face publication date of January 3, 2026 is newer than the ByteDance candidates’ October 22, 2025 publication date.[1][3][5] The placement follows release recency. Its benchmark status is “no public benchmark yet,” so the position does not establish superior image quality, generation speed, or suitability for your GPU.[1]
How much memory would quantized weights take?
For weight-only planning, LTX-2 (Lightricks) has 18.9B parameters: 9.45 GB at 4-bit, estimated (params x 0.5 bytes).[1] Video-As-Prompt-Wan2.1-14B (ByteDance) has 20.7B parameters: 10.35 GB, estimated (params x 0.5 bytes).[3] Video-As-Prompt-CogVideoX-5B (ByteDance) has 11.1B parameters: 5.55 GB, estimated (params x 0.5 bytes).[5] Those estimates describe parameter storage; they do not establish total runtime memory, an available quantized checkpoint, or successful operation on your GPU.
Do the listed weight sizes tell me how much GPU memory I need?
The listed BF16 weight totals are 314.3 GB for LTX-2 (Lightricks),[1] 65.8 GB for Video-As-Prompt-Wan2.1-14B (ByteDance),[3] and 32.6 GB for Video-As-Prompt-CogVideoX-5B (ByteDance).[5] Those figures are weight listings, not measured runtime memory requirements. Treat GPU fit as unverified until a concrete configuration has documented memory use; neither the listings nor parameter-storage estimates establish that result.
Which licenses apply if I want to build a product?
Video-As-Prompt-Wan2.1-14B (ByteDance) and Video-As-Prompt-CogVideoX-5B (ByteDance) list Apache-2.0.[3][5] LTX-2 (Lightricks) lists the ltx-2-community-license-agreement.[1] Review the applicable license text against your intended use before choosing a model for a product. The license names identify the terms to inspect; they do not, by themselves, resolve obligations for your particular deployment or distribution.
Can I compare benchmark scores, clip length, or generation speed?
All candidates have the status “no public benchmark yet.”[1][3][5] No comparable published scores, context limits, clip lengths, or measured generation speeds are specified for this comparison. Keep those fields marked as unknown when evaluating a deployment. The published release dates, parameter counts, weight listings, and licenses support a shortlist, but do not establish a performance winner.
Sources
- Lightricks/LTX-2 model card (Hugging Face) — 2026-09-25
- LTX-2: Efficient Joint Audio-Visual Foundation Model — 2026-01-06
- ByteDance/Video-As-Prompt-Wan2.1-14B model card (Hugging Face) — 2026-09-25
- Video-As-Prompt: Unified Semantic Control for Video Generation — 2025-10-23
- ByteDance/Video-As-Prompt-CogVideoX-5B model card (Hugging Face) — 2026-09-25