Quick Answer
VibeVoice-ASR-HF (Microsoft) is the top pick as of September 2026 because it is the newest release in this ranking of audio text to text models from labs with a published paper or leaderboard record [1][3][5][7]. In order, the ranking is VibeVoice-ASR-HF (Microsoft), Music Flamingo (NVIDIA), Audio Flamingo 3 (NVIDIA), and Qwen2-Audio-7B-Instruct (Qwen) [1][3].
Key Takeaways
- Microsoft VibeVoice-ASR-HF leads because its Hugging Face publication date is the newest among the eligible candidates [1][3][5][7]. Its 8.3B parameters [5] give a 4-bit weight footprint of 4.15 GB, estimated (params x 0.5 bytes) [5]; context is 128K tokens, the license is MIT, and there is no public benchmark yet [5].
- NVIDIA Music Flamingo (nvidia/music-flamingo-2601-hf) has 8.3B parameters [7]: 4-bit weights are 4.15 GB, estimated (params x 0.5 bytes) [7]. Context is 1,200 tokens [7]. License status is unconfirmed; no public benchmark yet [7].
- NVIDIA Audio Flamingo 3 (nvidia/audio-flamingo-3-hf) [3] has 8.3B parameters [3]: 4-bit weights are 4.15 GB, estimated (params x 0.5 bytes) [3]. Context is 32K tokens [3]. License status is unconfirmed; no public benchmark yet [3].
- Qwen Qwen2-Audio-7B-Instruct is an established pick [1]. Its 8.4B parameters [1] give 4-bit weights of 4.2 GB, estimated (params x 0.5 bytes) [1]. Context is 8K tokens, the license is Apache-2.0, and there is no public benchmark yet [1].
- Ranking note: newest release first, then downloads; no public benchmark scores any candidate against the others [1][3][5][7]. Scope admits only audio text to text models from labs with a published paper or leaderboard record, using the Hugging Face pipeline tag
audio-text-to-text[1][2][3][4][5][6][7][8]. Qwen’s established-pick exemption bypasses the recency gate; established picks follow the same ordering and cannot take first place through that exemption [1].
How do these audio language models compare on memory, context, licenses and benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| VibeVoice-ASR-HF [5] | Microsoft | 8.3B [5] | 4-bit weights: 4.15 GB estimated (params x 0.5 bytes), based on 8.3B [5] | 2026-03-02 [5] | MIT [5] | no public benchmark yet |
| Music Flamingo (music-flamingo-2601-hf) [7] | NVIDIA | 8.3B [7] | 4-bit weights: 4.15 GB estimated (params x 0.5 bytes), based on 8.3B [7] | 2026-01-01 [7] | unconfirmed | no public benchmark yet |
| Audio Flamingo 3 (audio-flamingo-3-hf) [3] | NVIDIA | 8.3B [3] | 4-bit weights: 4.15 GB estimated (params x 0.5 bytes), based on 8.3B [3] | 2025-10-24 [3] | unconfirmed | no public benchmark yet |
| Qwen2-Audio-7B-Instruct — established pick [1] | Qwen | 8.4B [1] | 4-bit weights: 4.2 GB estimated (params x 0.5 bytes), based on 8.4B [1] | 2024-07-31 [1] | Apache-2.0 [1] | no public benchmark yet |
Which audio language models should you consider for a local GPU?
1. VibeVoice-ASR-HF
VibeVoice-ASR-HF by Microsoft ranks first because its Hugging Face release is the newest among the admitted candidates, dated March 2, 2026.[5][7][3][1] The ordering basis is newest release first, then downloads; its benchmark status is “no public benchmark yet.”[5] The ranking admits only audio text to text models from labs with a published paper or leaderboard record, using the Hugging Face pipeline tag audio-text-to-text.
Microsoft lists 8.3B parameters, a 128K-token context window, 16.7 GB of BF16 weights and an MIT license.[5] Quantized weight storage is approximately 4.15 GB at 4-bit precision, estimated (params x 0.5 bytes) from the cited 8.3B parameters.[5] Weight fit on a 24 GB GPU is therefore estimated (params x 0.5 bytes), rather than demonstrated by a published runtime measurement.[5] Treat that calculation as a starting point for hardware planning, not a complete memory budget.
Choose VibeVoice-ASR-HF as a candidate for local speech-transcription evaluation, with Microsoft’s VIBEVOICE-ASR Technical Report providing the accompanying technical reference.[6] The advertised 128K-token context makes context capacity worth evaluating against your workload.[5] The caveat is that neither the context specification nor the weight estimate establishes how well your intended workload runs locally. Validate memory use and transcription quality before committing to a deployment.
2. Music Flamingo
Music Flamingo by NVIDIA ranks second because its Hugging Face release falls between Microsoft’s VibeVoice-ASR-HF and NVIDIA’s Audio Flamingo 3 [5][7][3]. The ordering is newest release first, then downloads; Music Flamingo has no public benchmark yet. Ranking scope admits only audio text to text models from labs with a published paper or leaderboard record, using the Hugging Face audio-text-to-text pipeline tag. Qwen’s Qwen2-Audio-7B-Instruct remains an established pick under the recency exemption [1].
The nvidia/music-flamingo-2601-hf checkpoint has 8.3B parameters, a 1,200-token context and 16.5 GB of BF16 weights; its first Hugging Face publication was January 1, 2026 [7]. Weight storage at 4-bit is 4.15 GB, estimated (params x 0.5 bytes) from the cited 8.3B parameters [7]. A 24 GB GPU therefore fits the weights on that estimate (params x 0.5 bytes), but the calculation does not establish total runtime memory or verify a working local configuration [7].
Use Music Flamingo for music-understanding work, the focus of NVIDIA’s “Music Flamingo: Scaling Music Understanding in Audio Language Models” [8]. Treat its ranking as a release-order position rather than evidence of superior audio quality. The practical caveat is licensing: license terms are unconfirmed here, so check the checkpoint’s terms before adopting it [7].
3. Audio Flamingo 3
Audio Flamingo 3 by NVIDIA ranks third because its Hugging Face release falls behind the Microsoft and NVIDIA Music Flamingo candidates in release order.[3][5][7] The ordering basis is newest release first, then downloads; no public benchmark scores these candidates against one another. Audio Flamingo 3 was first published on Hugging Face on October 24, 2025, and has no public benchmark yet.[3] Its position reflects release timing, rather than demonstrated performance superiority.
The model has 8.3B parameters and a 32K-token context window.[3] Quantized weight storage is approximately 4.15 GB, estimated (params x 0.5 bytes) from the cited 8.3B parameters.[3] The model card lists BF16 weights at 33.4 GB.[3] For the target GPU, quantization is therefore the practical configuration to investigate. The weight estimate excludes runtime allocations, so it cannot establish total VRAM requirements or confirm that the full context window will fit during local inference.
Consider Audio Flamingo 3 for audio-to-text workflows where context capacity matters: its listed 32K-token window exceeds Music Flamingo’s 1,200-token window.[3][7] That specification supports considering it for longer contexts, but does not establish audio-duration limits, answer quality or speed. A deployment caveat is licensing: the license remains unconfirmed here.[3] Verify the repository’s terms before adopting the model.
4. Qwen2-Audio-7B-Instruct
Qwen2-Audio-7B-Instruct by Qwen ranks fourth as an established pick, with a Hugging Face publication date of July 31, 2024.[1] Its inclusion bypasses the recency gate because it is among its family’s three most-downloaded models.[1] The ordering follows newest release first, then downloads; the established-pick exemption does not move it ahead of newer candidates.[3][5][7] Its benchmark status is no public benchmark yet, so this position should not be read as a measured quality comparison.
The model has 8.4B parameters, an 8K-token context window and 16.8 GB of BF16 weights.[1] For local deployment, four-bit weight memory is approximately 4.2 GB, estimated (params x 0.5 bytes) from the cited 8.4B parameters.[1] Weight fit on a 24 GB GPU is therefore an estimate using that calculation, rather than a verified runtime result.[1] Treat the calculation as a starting point for hardware planning, not a complete memory budget.
Qwen2-Audio-7B-Instruct is a practical candidate for local audio-to-text projects where the stated Apache-2.0 license is a selection requirement.[1] Its 8K-token context provides a concrete limit to check against your intended workload.[1] The caveat is validation: neither the ranking nor the weight estimate establishes task accuracy or reliable operation at your chosen context length. Verify both before committing to a deployment.
How can you estimate quantized weight memory from parameter counts?
Estimate quantized weight memory by multiplying the parameter count by the assumed storage per parameter: for 4-bit weights, label the result “estimated (params x 0.5 bytes)” beside the cited parameter count [1][3][5][7].
Qwen’s Qwen2-Audio-7B-Instruct has 8.4B parameters [1], giving approximately 4.2 GB of weight memory, estimated (params x 0.5 bytes) [1]. Microsoft’s VibeVoice-ASR-HF has 8.3B parameters [5], giving approximately 4.15 GB, estimated (params x 0.5 bytes) [5]. Both calculations describe weight storage under the stated assumption; neither is a measured runtime footprint.
NVIDIA’s Music Flamingo (nvidia/music-flamingo-2601-hf) and Audio Flamingo 3 (nvidia/audio-flamingo-3-hf) each have 8.3B parameters [7][3]. Each therefore yields approximately 4.15 GB of weight memory, estimated (params x 0.5 bytes) [7][3]. Equal parameter counts produce equal estimates under this formula, but do not establish equal inference memory requirements.
Keep weight estimates separate from context specifications. VibeVoice-ASR-HF lists a context length of 128K tokens [5], while Qwen2-Audio-7B-Instruct lists 8K tokens [1]. The parameter calculation does not account for running either model at its stated context limit. Use the estimate to compare nominal weight storage, and leave GPU-fit claims unconfirmed until total runtime memory is established for the intended configuration.
Which models have confirmed licenses for local use?
Microsoft’s VibeVoice-ASR-HF has a confirmed MIT license [5], and Qwen’s Qwen2-Audio-7B-Instruct has a confirmed Apache-2.0 license [1]. Both provide an explicit license to evaluate for a local deployment. Choose between them with the applicable license terms in view; license confirmation alone does not establish suitability for your workload.
NVIDIA’s Music Flamingo (nvidia/music-flamingo-2601-hf) has an unconfirmed license status [7]. NVIDIA’s Audio Flamingo 3 (nvidia/audio-flamingo-3-hf) also has an unconfirmed license status [3]. Treat those entries as requiring verification before deployment. “Unconfirmed” does not mean prohibited, and it does not mean a license has never been published.
For a deployment decision that requires a confirmed license, keep Microsoft’s VibeVoice-ASR-HF [5] and Qwen’s Qwen2-Audio-7B-Instruct [1] on the shortlist. For either NVIDIA candidate, check the repository’s applicable license text and any terms attached to the weights before deciding whether your intended use is covered.
Record the model revision, retain the accompanying license text, and review the terms against your intended use, modification and redistribution plans. Keep that review separate from hardware evaluation: a confirmed license answers a different question from whether a model will fit and run well on your GPU.
How are models ranked when no public benchmark scores the candidates?
When no public benchmark scores the candidates, the ordering is newest release first, then downloads. [5][7][3][1] Downloads serve only as a tiebreak; they do not establish audio quality or performance. The ranking admits only audio text to text models from labs with a published paper or leaderboard record, using the Hugging Face pipeline tag audio-text-to-text.
Microsoft’s VibeVoice-ASR-HF (VibeVoice-ASR-HF (Microsoft)) takes first place because its Hugging Face publication date of March 2, 2026 is the newest among the eligible candidates. [5][7][3][1] NVIDIA’s Music Flamingo (nvidia/music-flamingo-2601-hf), published January 1, 2026, follows. [7] NVIDIA’s Audio Flamingo 3 (nvidia/audio-flamingo-3-hf), published October 24, 2025, comes next. [3] Qwen’s Qwen2-Audio-7B-Instruct (Qwen2-Audio-7B-Instruct (Qwen)), published July 31, 2024, follows as an established pick. [1]
Qwen’s established-pick status reflects its position among its family’s three most-downloaded models and permits it to bypass the twelve-month recency gate. [1] The exemption does not grant first place. Parameter count does not determine the order.
Every candidate is marked “no public benchmark yet.” A leaderboard score may rank only candidates evaluated against each other on that board, and an older scored model cannot outrank a newer unscored model. Model-card benchmark results must be labeled “self-reported.” Readers should therefore treat this order as a release-based shortlist, not a demonstrated comparison of accuracy, speed or local runtime memory.
Frequently Asked Questions
Which model should I start with?
Microsoft’s VibeVoice-ASR-HF (VibeVoice-ASR-HF (Microsoft)) ranks first because its Hugging Face publication date, 2026-03-02, is the newest among the eligible releases. [5][7][3] The remaining order is NVIDIA’s Music Flamingo (nvidia/music-flamingo-2601-hf), NVIDIA’s Audio Flamingo 3 (nvidia/audio-flamingo-3-hf), then Qwen’s Qwen2-Audio-7B-Instruct (Qwen2-Audio-7B-Instruct (Qwen)). [7][3][1] Treat that position as a release-based starting point, rather than evidence of superior audio accuracy: every candidate has no public benchmark yet.
How does the ranking work?
The ordering basis is newest release first, then downloads; no public benchmark scores any candidate. This ranking admits only audio text to text models from labs with a published paper or leaderboard record (Hugging Face pipeline tags: audio-text-to-text). [1][2][3][4][5][6][7][8] Qwen2-Audio-7B-Instruct is an established pick that bypasses the recency gate, but that exemption does not earn first place. [1] Downloads serve only as a tiebreak, never as proof of quality.
How much memory would the quantized weights need?
VibeVoice-ASR-HF, Music Flamingo and Audio Flamingo 3 each have 8.3B parameters: 4-bit weight memory is 4.15 GB, estimated (params x 0.5 bytes). [5][7][3] Qwen2-Audio-7B-Instruct has 8.4B parameters: 4-bit weight memory is 4.2 GB, estimated (params x 0.5 bytes). [1] Those calculations describe weight storage only. Use them for initial budgeting; they do not establish total runtime memory, a working quantized configuration or a verified GPU fit.
Which model offers the longest context?
VibeVoice-ASR-HF lists the longest token context among these candidates at 128K tokens. [5][3][1][7] Audio Flamingo 3 lists 32K tokens, Qwen2-Audio-7B-Instruct lists 8K tokens, and Music Flamingo lists 1,200 tokens. [3][1][7] Compare those limits when planning prompts and responses. Treat token context as a configuration limit; the listed figures do not establish supported recording duration or demonstrate memory use at the full context limit.
Which licenses are confirmed?
VibeVoice-ASR-HF lists the MIT license, while Qwen2-Audio-7B-Instruct lists Apache-2.0. [5][1] License details for Music Flamingo and Audio Flamingo 3 are unconfirmed here; neither should be described as having an unpublished license. [7][3] Before adopting either NVIDIA model, verify the applicable license terms. For a deployment decision that requires a confirmed license, begin your review with the Microsoft and Qwen candidates.
Are there benchmark scores proving which model is better?
Every candidate has the same benchmark status here: no public benchmark yet. No shared public leaderboard scores these candidates against each other, so the ranking cannot establish an accuracy winner. [5][7][3][1] Model-card benchmark results, if discussed, must be labeled self-reported. Choose a candidate using the release order, context and confirmed license, then evaluate it on your own audio before committing to deployment.
Sources
- Qwen/Qwen2-Audio-7B-Instruct model card (Hugging Face) — 2026-09-25
- Qwen2-Audio Technical Report — 2024-07-15
- nvidia/audio-flamingo-3-hf model card (Hugging Face) — 2026-09-25
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models — 2025-07-10
- microsoft/VibeVoice-ASR-HF model card (Hugging Face) — 2026-09-25
- VIBEVOICE-ASR Technical Report — 2026-03-14
- nvidia/music-flamingo-2601-hf model card (Hugging Face) — 2026-09-25
- Music Flamingo: Scaling Music Understanding in Audio Language Models — 2025-11-13