Quick Answer
MiniMax-Music3 from MiniMaxAI is the top pick as of September 2026 because it is the newest release among the eligible models [1][2][4][6]. In order, the ranking is MiniMax-Music3 (MiniMaxAI), magenta-realtime-2 (google), stable-audio-3-optimized (stabilityai), and stable-audio-3-medium-base (stabilityai) [1][2].
Key Takeaways
- Ranking note: Scope admits only text to audio models from labs with a published paper or leaderboard record (Hugging Face pipeline tag: text-to-audio).[1][2][3][4][5][6] Eligible releases fall within the past 12 months; no public benchmark scores any candidate, so the ordering is newest release first, then downloads.[1][2][4][6]
- MiniMaxAI’s MiniMax-Music3 ranks first because its Hugging Face publication date, 2026-08-07, is the newest among the eligible models.[1][2][4][6] Its 2.4B parameters imply 1.2 GB of 4-bit weights, estimated (params x 0.5 bytes); no public benchmark yet.[1]
- google’s magenta-realtime-2 ranks second, published on 2026-05-28, with 11.3 GB of full-precision weights and a CC-BY-4.0 license; no public benchmark yet.[2] The family’s paper is “Live Music Models.”[3]
- stabilityai’s stable-audio-3-optimized ranks third, published on 2026-05-18 under the stable-audio-community license; no public benchmark yet.[4] The family’s paper is “Stable Audio 3.”[5]
- stabilityai’s stable-audio-3-medium-base ranks fourth, published on 2026-05-17 under the stable-audio-community license.[6] Its 2.3B parameters imply 1.15 GB of 4-bit weights, estimated (params x 0.5 bytes); no public benchmark yet.[6]
How do local music and audio generation models compare on specs and benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| MiniMax-Music3 [1] | MiniMaxAI | 2.4B [1] | 4-bit weights: 1.2 GB estimated (params x 0.5 bytes), using 2.4B parameters [1]; runtime VRAM: not published | 2026-08-07 [1] | not published | no public benchmark yet |
| magenta-realtime-2 [2] | not published | Full-precision weights: 11.3 GB [2]; quantized weights/runtime VRAM: not published | 2026-05-28 [2] | CC-BY-4.0 [2] | no public benchmark yet | |
| stable-audio-3-optimized [4] | stabilityai | not published | not published | 2026-05-18 [4] | stable-audio-community [4] | no public benchmark yet |
| stable-audio-3-medium-base [6] | stabilityai | 2.3B [6] | 4-bit weights: 1.15 GB estimated (params x 0.5 bytes), using 2.3B parameters [6]; runtime VRAM: not published | 2026-05-17 [6] | stable-audio-community [6] | no public benchmark yet |
Which music and audio generation models should you consider running locally?
1. MiniMax-Music3
MiniMax-Music3 by MiniMaxAI ranks first because its Hugging Face publication date, August 7, 2026, is the newest among the admitted candidates.[1][2][4][6] Ranking note: newest release first, then downloads; no public benchmark scores any candidate, and downloads serve only as a tiebreak. The ranking admits only text to audio models from labs with a published paper or leaderboard record, using the Hugging Face text-to-audio pipeline tag.
MiniMax-Music3 has 2.4 billion parameters.[1] Weight memory at four-bit precision is 1.2 GB, estimated (params x 0.5 bytes) from that parameter count.[1] The listed F32 weights occupy 47.0 GB.[1] Those figures describe different quantities: parameter-based weight storage does not establish the memory needed to load and run the published package. A local hardware recommendation requires verified GPU and system-memory measurements, plus confirmation that the runtime supports quantized loading. The estimate alone cannot justify a GPU purchase.
The practical use case is evaluating music generation within a local workflow, with hardware compatibility and output suitability checked before committing to deployment. The central caveat is no public benchmark yet: the ranking reflects release recency, without establishing a quality advantage. License terms and context or generation-duration limits also remain unverified, so deployment planning needs those details before treating the model as a production choice.
2. magenta-realtime-2
magenta-realtime-2 by google ranks second under the ordering rule: newest release first, then downloads. Its Hugging Face publication date is May 28, 2026,[2] behind MiniMax-Music3 by MiniMaxAI, published August 7, 2026.[1] The ranking admits only text to audio models from labs with a published paper or leaderboard record, using Hugging Face’s text-to-audio pipeline tag. No public benchmark yet establishes magenta-realtime-2’s position on audio quality; its placement reflects release timing.
Full-precision weights total 11.3 GB,[2] but that figure is not a GPU memory requirement. A parameter count is needed to calculate an estimated quantized weight size; runtime memory requirements also need verification before choosing hardware. A specific GPU recommendation, minimum RAM allocation, or supported context duration would therefore be premature. For a local installation, confirm the runtime’s hardware requirements before downloading the weights or buying a GPU.
The practical use case is a local music-generation evaluation where the published CC-BY-4.0 license suits the project’s attribution requirements.[2] The associated paper is google’s “Live Music Models,” dated August 6, 2025.[3] The caveat is that the paper’s title does not establish this release’s latency or streaming performance. Evaluate responsiveness and output quality on the intended machine before committing to an interactive music workflow.
3. stable-audio-3-optimized
stable-audio-3-optimized by stabilityai ranks third because its Hugging Face publication date falls after google’s magenta-realtime-2 and before stabilityai’s stable-audio-3-medium-base.[2][4][6] The ranking admits only text-to-audio models from labs with a published paper or leaderboard record. The ordering basis is newest release first, then downloads; this position does not establish an audio-quality advantage. Benchmark status: no public benchmark yet.
The model was first published on Hugging Face on May 18, 2026, and carries the stable-audio-community license.[4] The accompanying paper, “Stable Audio 3,” was published on May 18, 2026.[5] Parameter count, context limits and weight size remain unconfirmed, so a quantized-memory estimate would be premature. The “optimized” name alone does not establish lower memory consumption, faster generation or a particular hardware target.
Consider this model for local text-to-audio evaluation when you can validate its runtime requirements on your own machine. A supported GPU configuration and measured system-memory requirement remain unconfirmed, so there is no defensible hardware purchase recommendation here. Before committing to a deployment, check the runtime’s device requirements and measure memory use with your intended workload. The practical caveat is that its ranking reflects release order rather than demonstrated generation quality or local efficiency; review the license terms for your intended use.[4]
4. stable-audio-3-medium-base
stable-audio-3-medium-base by stabilityai ranks fourth because its first Hugging Face publication, on May 17, 2026, precedes the other candidates’ publication dates [1][2][4][6]. The ordering is newest release first, then downloads; no public benchmark yet establishes its quality against the other candidates. The ranking admits only text-to-audio models from labs with a published paper or leaderboard record. The accompanying paper is “Stable Audio 3,” published on May 18, 2026 [5].
The model has 2.3B parameters [6], giving a weight-only memory figure of 1.15 GB, estimated (params x 0.5 bytes) [6]. The listed F32 weights occupy 10.4 GB [6]. Treat the parameter-based estimate as a planning input: runtime allocations and supporting components also need memory. A specific GPU or system RAM requirement cannot be established from the weight figures alone. Verify the intended runtime’s quantization support and total memory requirements before choosing hardware.
Consider this model for local text-to-audio evaluation when a published technical description and an identified license are useful selection criteria. The release uses the stable-audio-community license [6]; review its terms against your intended use. The practical caveat is incomplete deployment guidance: a confirmed context limit and measured runtime memory requirement are unavailable here. Evaluate output quality and memory use in your own workflow before committing to deployment.
How can you estimate memory requirements from parameter counts?
Estimate weight memory by multiplying the parameter count by the bytes per parameter: for 4-bit weights, use estimated (params x 0.5 bytes) alongside the cited parameter count. [1][6] Treat the result as a weight-storage estimate, not a complete memory budget for running the model.
MiniMaxAI’s MiniMax-Music3 (MiniMaxAI) has 2.4 billion parameters [1], giving approximately 1.2 GB, estimated (params x 0.5 bytes). [1] Stability AI’s stable-audio-3-medium-base (stabilityai) has 2.3 billion parameters [6], giving approximately 1.15 GB, estimated (params x 0.5 bytes). [6] Both calculations use decimal gigabytes. Neither calculation establishes that a compatible quantized release is available or that inference will fit within that amount of GPU memory.
Published checkpoint sizes are separate figures. MiniMax-Music3 (MiniMaxAI) lists F32 weights of 47.0 GB [1], while stable-audio-3-medium-base (stabilityai) lists F32 weights of 10.4 GB. [6] Keep those reported sizes separate from parameter-based estimates; do not present the estimates as measured checkpoint sizes or measured runtime memory.
For hardware planning, use the calculation as a starting point. A practical memory budget also needs to account for the implementation and its runtime allocations. Avoid turning a weight-only estimate into a claim that a particular GPU can run the model.
Which licenses apply to local music and audio generation models?
The listed licenses are CC-BY-4.0 for google’s magenta-realtime-2 (google) [2] and stable-audio-community for stabilityai’s stable-audio-3-optimized (stabilityai) [4] and stable-audio-3-medium-base (stabilityai) [6]. MiniMaxAI’s MiniMax-Music3 (MiniMaxAI) has no license specified in the available model details, so its permissions remain unconfirmed here. [1]
For Magenta, use CC-BY-4.0 as the starting point for reviewing your intended use. [2] Check the actual terms before distributing weights, incorporating the model into a product, or deciding how to provide attribution. Keep the license review tied to the exact repository you plan to download.
Both listed Stable Audio variants carry the same license identifier: stable-audio-community. [4][6] Review that agreement directly before making a deployment decision. The shared identifier establishes the listed license; it does not, by itself, explain the conditions relevant to your project.
For MiniMax-Music3, resolve the missing license information before treating the weights as cleared for your intended use. [1] Record the applicable terms alongside your deployment configuration, and distinguish permission to use or redistribute model weights from any terms governing generated audio. Local execution should be a deployment choice, not a substitute for checking permissions.
What does no public benchmark yet mean when choosing a model?
“No public benchmark yet” means no public benchmark scores these candidates against one another, so the ranking does not establish comparative audio quality.[1][2][4][6] An unscored model has an evidence gap, not a demonstrated quality problem. Treat its position as a starting point for evaluation, rather than proof that its output will suit your project.
The ranking admits only text to audio models from labs with a published paper or leaderboard record, using the Hugging Face text-to-audio pipeline tag. The ordering basis is newest release first, then downloads.[1][2][4][6] Downloads serve only as a tiebreak, not evidence of audio quality.
MiniMaxAI’s MiniMax-Music3 takes the first position because its Hugging Face publication date, August 7, 2026, is the newest among the admitted candidates.[1][2][4][6] That placement reflects release recency rather than a measured performance advantage.
For a practical choice, compare outputs using the same prompts and your intended workflow. Check whether the audio follows your instructions, whether audible artifacts make it unusable, and whether generation takes an acceptable amount of time on your hardware. Keep your observations separate from published benchmark results.
Before committing, check the license and available runtime documentation. Missing memory or context specifications should remain unknown, rather than becoming assumed capabilities. Any later model-card benchmark should be labeled self-reported; a public leaderboard should support comparisons only between models it actually evaluates.
Frequently Asked Questions
Which local music and audio generation models should I compare first?
MiniMax-Music3 (MiniMaxAI) ranks first because its August 7, 2026 publication is the newest among these candidates [1][2][4][6]. The remaining order is magenta-realtime-2 (google) [2], stable-audio-3-optimized (stabilityai) [4], then stable-audio-3-medium-base (stabilityai) [6]. Treat that order as a release-based shortlist, rather than evidence of superior sound quality.
How is the ranking decided without comparable benchmarks?
Ranking note: newest release first, then downloads. Eligibility is limited to text-to-audio models from labs with a published paper or leaderboard record, using the Hugging Face text-to-audio pipeline tag [1][2][3][4][5][6]. As of September 25, 2026, every candidate has no public benchmark yet [1][2][4][6]. Download counts serve only as a tiebreak; they do not establish audio quality.
How much memory would quantized weights need?
MiniMax-Music3 (MiniMaxAI) has 2.4B parameters [1]: 4-bit weight memory is 1.2 GB, estimated (params x 0.5 bytes) [1]. stable-audio-3-medium-base (stabilityai) has 2.3B parameters [6]: 4-bit weight memory is 1.15 GB, estimated (params x 0.5 bytes) [6]. Treat those calculations as weight-only estimates, not measured runtime requirements, confirmation of quantization support, or a guarantee that either model fits your GPU.
Can I use these models in a commercial project?
magenta-realtime-2 (google) carries the CC-BY-4.0 license [2]. stable-audio-3-optimized (stabilityai) and stable-audio-3-medium-base (stabilityai) use the stable-audio-community license [4][6]. Read the applicable license before committing to commercial use, redistribution, or a hosted service. For MiniMax-Music3 (MiniMaxAI), check the repository’s license terms directly [1]. Do not treat access to downloadable weights as permission for every intended use.
Which model should I choose for generating long tracks?
Keep duration as an unresolved selection requirement: no context-window or output-duration figures are specified for this comparison. Consult the model documentation for MiniMax-Music3 (MiniMaxAI) [1], magenta-realtime-2 (google) [2], and the stabilityai candidates [4][6]. Confirm supported generation length and continuation behavior before choosing a workflow. Avoid inferring track length from a model’s name, release date, parameter count, or weight-file size.
How large are the weight downloads?
Published weight sizes are 47.0 GB in F32 for MiniMax-Music3 (MiniMaxAI) [1], 11.3 GB in full precision for magenta-realtime-2 (google) [2], and 10.4 GB in F32 for stable-audio-3-medium-base (stabilityai) [6]. Use those figures to plan weight storage. Keep published file sizes separate from estimated quantized weight memory, and check the selected repository’s file list before downloading.
Sources
- MiniMaxAI/MiniMax-Music3 model card (Hugging Face) — 2026-09-25
- google/magenta-realtime-2 model card (Hugging Face) — 2026-09-25
- Live Music Models — 2025-08-06
- stabilityai/stable-audio-3-optimized model card (Hugging Face) — 2026-09-25
- Stable Audio 3 — 2026-05-18
- stabilityai/stable-audio-3-medium-base model card (Hugging Face) — 2026-09-25