Quick Answer
moshika-rag-pytorch-bf16 (Kyutai) is the top pick as of September 2026 because it is the newest release in this ranking [12][8][10][4][6][1][3]. In order, the ranking is moshika-rag-pytorch-bf16 (Kyutai), Covo-Audio-Chat (Tencent), hibiki-zero-3b-pytorch-bf16 (Kyutai), Qwen3-TTS-Tokenizer-12Hz (Qwen), LFM2.5-Audio-1.5B (LiquidAI), bigvgan_v2_22khz_80band_256x (NVIDIA), and bigvgan_v2_44khz_128band_512x (NVIDIA) [12][8].
Key Takeaways
- Ranking note: the ordering is newest release first, then downloads; every candidate has “no public benchmark yet.” The ranking admits only audio-to-audio models from labs with a published paper or leaderboard record, using Hugging Face’s audio-to-audio pipeline tag. Established picks bypass the recency gate.[1][3][4][6][8][10][12]
- Kyutai’s moshika-rag-pytorch-bf16 takes first place because its March 19, 2026 release is the newest eligible release.[12] Its 7.7B parameters imply 3.85 GB of 4-bit weights, estimated (params x 0.5 bytes); its license is CC-BY-4.0.[12]
- Tencent’s Covo-Audio-Chat has a 128K-token context and 8.4B parameters, implying 4.2 GB of 4-bit weights, estimated (params x 0.5 bytes).[8] Kyutai’s hibiki-zero-3b-pytorch-bf16 lists 6.6 GB of full-precision weights; a parameter-based estimate is unavailable.[10]
- Qwen’s Qwen3-TTS-Tokenizer-12Hz has 171M parameters, implying 85.5 MB of 4-bit weights, estimated (params x 0.5 bytes), and uses the Apache-2.0 license.[4]
- LiquidAI’s LFM2.5-Audio-1.5B has 1.5B parameters, implying 0.75 GB of 4-bit weights, estimated (params x 0.5 bytes), and uses the lfm1.0 license.[6]
- NVIDIA’s bigvgan_v2_22khz_80band_256x and bigvgan_v2_44khz_128band_512x are MIT-licensed established picks, both published July 15, 2024.[1][3] Their tied release date puts the former ahead on downloads: 1,324,090 versus 814,676 over the last 30 days as of September 26, 2026.[1][3]
How do local audio models compare on parameters, weight memory, context, licenses and public benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| moshika-rag-pytorch-bf16 [12] | Kyutai | 7.7B [12] | 4-bit weights: 3.85 GB estimated (params x 0.5 bytes) [12] | 2026-03-19 [12] | CC-BY-4.0 [12] | no public benchmark yet |
| Covo-Audio-Chat [8] | Tencent | 8.4B [8] | 4-bit weights: 4.2 GB estimated (params x 0.5 bytes) [8] | 2026-03-16 [8] | — | no public benchmark yet |
| hibiki-zero-3b-pytorch-bf16 [10] | Kyutai | — | Full-precision weights: 6.6 GB [10] | 2026-02-09 [10] | — | no public benchmark yet |
| Qwen3-TTS-Tokenizer-12Hz [4] | Qwen | 171M [4] | 4-bit weights: 85.5 MB estimated (params x 0.5 bytes) [4] | 2026-01-21 [4] | Apache-2.0 [4] | no public benchmark yet |
| LFM2.5-Audio-1.5B [6] | LiquidAI | 1.5B [6] | 4-bit weights: 0.75 GB estimated (params x 0.5 bytes) [6] | 2025-12-18 [6] | lfm1.0 [6] | no public benchmark yet |
| bigvgan_v2_22khz_80band_256x — established pick [1] | NVIDIA | — | — | 2024-07-15 [1] | MIT [1] | no public benchmark yet |
| bigvgan_v2_44khz_128band_512x — established pick [3] | NVIDIA | — | — | 2024-07-15 [3] | MIT [3] | no public benchmark yet |
Which audio enhancement models should you consider for a CPU-only PC?
1. moshika-rag-pytorch-bf16
moshika-rag-pytorch-bf16 by Kyutai ranks first because its March 19, 2026 release is newer than the next candidate, Tencent’s Covo-Audio-Chat, released March 16, 2026.[12][8] With no public benchmark yet, the ordering is newest release first, then downloads. Its documented focus is asynchronous knowledge retrieval for full-duplex speech language models.[13]
The model has 7.7 billion parameters, a published BF16 weight size of 15.8 GB, and a CC-BY-4.0 license.[12] Weight storage at 4-bit would be 3.85 GB, estimated (params x 0.5 bytes) from the cited 7.7 billion parameters.[12] Total system RAM requirements, CPU execution speed, and a supported quantized CPU runtime remain unverified; the storage estimate alone cannot establish whether your PC can run it practically.
Choose it for evaluating retrieval-backed speech dialogue, the application described by MoshiRAG.[13] The caveat for an audio-enhancement shortlist is that its placement does not establish denoising quality or practical CPU performance. A context-length specification is also unavailable here.
2. Covo-Audio-Chat
Covo-Audio-Chat by Tencent ranks second under the newest-release-first ordering: its publication date is March 16, 2026,[8] behind Kyutai’s moshika-rag-pytorch-bf16, published March 19, 2026.[12] Covo-Audio-Chat has no public benchmark yet, so its position does not establish superior audio enhancement quality or CPU performance.
Tencent lists 8.4B parameters, a 128K-token context window and 16.8 GB of BF16 weights.[8] Weight storage at four-bit precision would be approximately 4.2 GB—estimated (params x 0.5 bytes) from the cited 8.4B parameters.[8] Treat that calculation as a weight-storage estimate, not a complete system RAM requirement or confirmation that a compatible quantized build is available.
Consider Covo-Audio-Chat for local audio-chat evaluation when a documented token context limit matters. CPU throughput, total runtime memory and a supported CPU configuration remain unverified. The practical caveat is hardware planning: the published weight size alone cannot establish whether your PC can deliver usable latency.
3. hibiki-zero-3b-pytorch-bf16
hibiki-zero-3b-pytorch-bf16 by Kyutai ranks third under the newest-release-first ordering: its Hugging Face publication date is February 9, 2026, behind the March releases occupying the preceding positions.[10][8][12] The placement reflects release timing rather than demonstrated audio quality or CPU performance; no public benchmark yet establishes its performance against the other candidates.
The published weight footprint is 6.6 GB at full precision.[10] A parameter count, context limit and license are not established here. The weight footprint alone does not establish the hardware needed to run it locally. A reliable RAM recommendation or quantized-memory estimate therefore cannot be given, and CPU compatibility and throughput still need verification.
Treat the model as a candidate for exploratory local audio evaluation. Before choosing it for an enhancement workflow, verify that its supported task matches your audio problem and measure memory use and processing time on your own PC. The central caveat is unverified suitability for CPU-only enhancement.
4. Qwen3-TTS-Tokenizer-12Hz
Qwen3-TTS-Tokenizer-12Hz by Qwen ranks here under the ordering “newest release first, then downloads.” Its publication date is January 21, 2026 [4], placing it behind Kyutai’s hibiki-zero-3b-pytorch-bf16 [10] and ahead of LiquidAI’s LFM2.5-Audio-1.5B [6]. The benchmark status is “no public benchmark yet”; its position does not establish CPU performance or enhancement quality.
Qwen lists 171M parameters, F32 weights occupying 0.7 GB, and an Apache-2.0 license [4]. Weight storage at 4-bit precision would be 85.5 MB, estimated (params x 0.5 bytes) from the cited 171M parameters [4]. Actual system RAM requirements would also need to account for runtime allocations; the weight estimate alone cannot establish hardware fit.
Consider the tokenizer for local experiments where weight storage is a selection constraint. The practical caveat is that CPU throughput, supported quantization, context length, and minimum system RAM are unspecified here. Treat it as an evaluation candidate, with CPU suitability still unverified.
5. LFM2.5-Audio-1.5B
LFM2.5-Audio-1.5B by LiquidAI[6] ranks here under the ordering rule “newest release first, then downloads,” with a Hugging Face publication date of December 18, 2025.[6] Its position reflects release timing rather than demonstrated CPU performance: no public benchmark yet.
The model has 1.5 billion parameters and published BF16 weights totaling 3.6 GB.[6] Weight storage at 4-bit is approximately 0.75 GB, estimated (params x 0.5 bytes) from the cited 1.5 billion parameters.[6] Treat that calculation as a weight-storage estimate, not a total RAM requirement or confirmation that a compatible quantized CPU runtime exists. Context length and CPU throughput are unspecified.
For engineers evaluating local audio-to-audio workloads, the practical use is an evaluation candidate whose weight budget can be estimated before runtime testing. The caveat is hardware uncertainty: a CPU-only PC recommendation needs runtime compatibility and measured memory use. The license is lfm1.0; check its terms against your intended deployment.[6]
6. bigvgan_v2_22khz_80band_256x
bigvgan_v2_22khz_80band_256x by NVIDIA ranks sixth as an established pick, with a Hugging Face publication date of July 15, 2024.[1] The ordering is newest release first, then downloads; its 1,324,090 downloads over the preceding 30 days place it ahead of the other NVIDIA entry with the same publication date.[1][3] The ranking admits only audio-to-audio models from labs with a published paper or leaderboard record. Established picks bypass the recency gate; download counts do not establish audio quality.
The model uses the MIT license.[1] Its documented role is neural vocoding, making a local audio pipeline needing a vocoder the appropriate use to evaluate.[2] Parameter count, context length and quantized weight memory are unspecified here. A weight-memory estimate cannot be calculated without a documented parameter count.
CPU runtime support, processor requirements, RAM requirements and measured latency remain unconfirmed. Treat local CPU deployment as something to validate before adoption. The benchmark status is “no public benchmark yet”; the ranking does not demonstrate enhancement quality or practical CPU speed.
7. bigvgan_v2_44khz_128band_512x
bigvgan_v2_44khz_128band_512x by NVIDIA ranks seventh as an established pick.[3] Ordering follows newest release first, then downloads, with no public benchmark yet. Its Hugging Face publication date matches NVIDIA’s bigvgan_v2_22khz_80band_256x: July 15, 2024.[1][3] Downloads break that tie: 814,676 versus 1,324,090 over the preceding 30 days, as of September 26, 2026.[1][3] Its established-pick status permits inclusion outside the recency window.[3]
The license is MIT, and the associated BigVGAN paper describes a universal neural vocoder.[2][3] Consider it for a local vocoder evaluation when that role matches your audio workflow.[2] A recommendation for general audio enhancement would require separate task-specific validation.
CPU-only deployment remains a qualification step: runtime compatibility, RAM requirements and processing speed are not established here. Parameter count, quantized weight memory and context length are also unspecified. The practical caveat is hardware uncertainty: verify that your intended runtime supports the model and measure memory use before committing to a local deployment.
How can you estimate quantized weight memory from parameter counts?
Estimate quantized weight memory by multiplying the parameter count by the storage per parameter: for 4-bit weights, use estimated (params x 0.5 bytes) alongside the cited parameter count.[4][6][8][12] Treat the result as a weight-storage calculation, rather than a prediction of total application memory.
Qwen’s Qwen3-TTS-Tokenizer-12Hz has 171M parameters: its quantized weight memory is estimated (params x 0.5 bytes) at 85.5 MB.[4] LiquidAI’s LFM2.5-Audio-1.5B has 1.5B parameters, giving 750 MB, estimated (params x 0.5 bytes).[6] Both calculations use decimal storage units.
Kyutai’s moshika-rag-pytorch-bf16 (Kyutai) has 7.7B parameters, giving 3.85 GB, estimated (params x 0.5 bytes).[12] Tencent’s Covo-Audio-Chat has 8.4B parameters, giving 4.2 GB, estimated (params x 0.5 bytes).[8] Those figures describe hypothetical quantized weights; they do not establish that a compatible quantized release or CPU execution path is available.
Use these estimates to compare weight-storage requirements before evaluating a local deployment. Keep the published checkpoint size separate from the calculated quantized size, and leave total RAM requirements unspecified without runtime measurements. A weight estimate alone does not establish CPU speed, audio quality, or whether a model will fit comfortably alongside the operating system and other applications.
Which licenses apply to local audio models?
The listed local audio models carry MIT, Apache-2.0, lfm1.0 or CC-BY-4.0 licenses, depending on the model; a license is not confirmed here for every candidate.[1][3][4][6][12]
NVIDIA’s bigvgan_v2_22khz_80band_256x (NVIDIA) uses MIT, as does NVIDIA’s bigvgan_v2_44khz_128band_512x (NVIDIA).[1][3] Both entries therefore share a license label. When choosing between them, keep the licensing decision separate from the technical evaluation.
Qwen’s Qwen3-TTS-Tokenizer-12Hz (Qwen) uses Apache-2.0.[4] LiquidAI’s LFM2.5-Audio-1.5B (LiquidAI) uses lfm1.0.[6] Treat those as separate license reviews: read the terms attached to the model you intend to deploy, and check your intended use against that license.
Kyutai’s moshika-rag-pytorch-bf16 (Kyutai) uses CC-BY-4.0.[12] Record that license alongside the downloaded model and review its terms before incorporating the weights into a distributed application.
Licensing remains unconfirmed here for Tencent’s Covo-Audio-Chat (Tencent) and Kyutai’s hibiki-zero-3b-pytorch-bf16 (Kyutai); verify each repository’s license before adoption.[8][10] Do not infer a license from another release by the same organization.
For a local deployment, keep a copy of the applicable license with your model files. Review permissions and obligations for your intended workflow, including modification or redistribution. Evaluate CPU compatibility separately: a license label does not establish runtime performance.
What remains unverified about CPU performance and audio enhancement quality?
CPU speed, total runtime memory, and audio enhancement quality remain unverified: the candidate information does not establish measured CPU performance or comparable enhancement scores. Each candidate has “no public benchmark yet” for this comparison, so the ordering cannot establish which delivers cleaner audio or lower latency on a CPU-only PC.
Weight storage provides a planning estimate, not a measured RAM requirement. Qwen’s Qwen3-TTS-Tokenizer-12Hz has 171M parameters; its 4-bit weight storage is approximately 85.5 MB, estimated (params x 0.5 bytes).[4] Actual runtime memory, quantized execution support, and any quality change from quantization remain unverified. A calculated weight footprint cannot establish whether a particular computer can run the model comfortably.
Context capacity also leaves practical questions unanswered. Tencent’s Covo-Audio-Chat lists a context length of 128K tokens.[8] That specification does not establish supported recording duration, processing speed, or memory consumption during audio processing. CPU evaluation would need a named processor, runtime, thread configuration, audio workload, and measurements of latency and peak memory.
Enhancement quality needs a separate evaluation using matched recordings and clearly defined tasks. Checks should cover noise removal, speech intelligibility, preservation of speaker characteristics, and introduced artifacts. Admission to the audio-to-audio category alone does not establish suitability for those tasks. Until comparable measurements are available, treat the ranking as a shortlist for evaluation rather than a demonstrated CPU enhancement recommendation.
Frequently Asked Questions
Which model ranks first for a CPU-only PC?
Kyutai’s moshika-rag-pytorch-bf16 (Kyutai) ranks first because its Hugging Face publication date, March 19, 2026, is the newest among the eligible candidates.[12] The model has no public benchmark yet, so its position does not establish superior enhancement quality or CPU speed. Its BF16 weights occupy 15.8 GB.[12] Treat the ranking as a release-based shortlist, with CPU runtime support and performance still requiring verification.
How is the ranking ordered?
Ranking note: this ranking admits only audio-to-audio models from labs with a published paper or leaderboard record (Hugging Face pipeline tag: audio-to-audio). No public benchmark scores any candidate, so the ordering is newest release first, then downloads: Kyutai’s moshika-rag-pytorch-bf16 (Kyutai),[12] Tencent’s Covo-Audio-Chat (Tencent),[8] Kyutai’s hibiki-zero-3b-pytorch-bf16 (Kyutai),[10] Qwen’s Qwen3-TTS-Tokenizer-12Hz (Qwen),[4] Liquid AI’s LFM2.5-Audio-1.5B (LiquidAI),[6] NVIDIA’s bigvgan_v2_22khz_80band_256x (NVIDIA),[1] and NVIDIA’s bigvgan_v2_44khz_128band_512x (NVIDIA).[3] Downloads serve only as a tiebreak; size does not determine placement.
How much memory would quantized weights need?
For Moshika, 7.7B parameters imply 3.85 GB estimated (params x 0.5 bytes).[12] For Covo-Audio-Chat, 8.4B parameters imply 4.2 GB estimated (params x 0.5 bytes).[8] For Qwen’s tokenizer, 171M parameters imply 85.5 MB estimated (params x 0.5 bytes).[4] For Liquid AI’s audio model, 1.5B parameters imply 0.75 GB estimated (params x 0.5 bytes).[6] Those calculations describe hypothetical weight storage, not verified quantized releases or total RAM requirements.
Which models have documented context limits?
Tencent’s Covo-Audio-Chat lists a context length of 128K tokens.[8] No context figure is available here for the other candidates, so a context comparison would be incomplete. Kyutai’s Hibiki checkpoint lists full-precision weights of 6.6 GB, but that storage figure does not establish a context limit or CPU memory requirement.[10] Covo-Audio-Chat and Hibiki each have no public benchmark yet; neither can be ranked by measured CPU performance.
What licenses do these models use?
Moshika uses CC-BY-4.0,[12] Qwen’s tokenizer uses Apache-2.0,[4] and Liquid AI’s audio model uses lfm1.0.[6] Both NVIDIA BigVGAN checkpoints use MIT.[1][3] License details are unspecified here for Covo-Audio-Chat and Hibiki; do not assume that another checkpoint’s terms apply. For deployment, check the exact repository’s license against your intended use. License labels alone do not establish CPU compatibility, quantization support or enhancement quality.
Why are older BigVGAN checkpoints included?
Both NVIDIA BigVGAN checkpoints are established picks: each belongs to its family’s three most-downloaded Hugging Face models and bypasses the twelve-month recency gate.[1][3] Both were first published on July 15, 2024.[1][3] Their tied release date makes downloads decisive: 1,324,090 versus 814,676 over the thirty days measured on September 26, 2026.[1][3] Established picks follow the same ordering rule, cannot take first place through the exemption, and have no public benchmark yet.
Sources
- nvidia/bigvgan_v2_22khz_80band_256x model card (Hugging Face) — 2026-09-26
- BigVGAN: A Universal Neural Vocoder with Large-Scale Training — 2022-06-09
- nvidia/bigvgan_v2_44khz_128band_512x model card (Hugging Face) — 2026-09-26
- Qwen/Qwen3-TTS-Tokenizer-12Hz model card (Hugging Face) — 2026-09-26
- Qwen3-TTS Technical Report — 2026-01-22
- LiquidAI/LFM2.5-Audio-1.5B model card (Hugging Face) — 2026-09-26
- LFM2 Technical Report — 2025-11-28
- tencent/Covo-Audio-Chat model card (Hugging Face) — 2026-09-26
- Covo-Audio Technical Report — 2026-02-10
- kyutai/hibiki-zero-3b-pytorch-bf16 model card (Hugging Face) — 2026-09-26
- Moshi: a speech-text foundation model for real-time dialogue — 2024-09-17
- kyutai/moshika-rag-pytorch-bf16 model card (Hugging Face) — 2026-09-26
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models — 2026-04-14