Quick Answer
VoxCPM2 (openbmb) is the top pick for local audiobook text-to-speech as of September 2026 because it is the newest eligible release [6]. In order, the ranking is VoxCPM2 (openbmb), s2-pro (fishaudio), Qwen3-TTS-12Hz-1.7B-CustomVoice (Qwen), Qwen3-TTS-12Hz-0.6B-CustomVoice (Qwen), VibeVoice-Realtime-0.5B (microsoft), Voxtral-4B-TTS-2603 (mistralai), Kokoro-82M (hexgrad), and XTTS-v2 (coqui) [6][10].
Key Takeaways
- openbmb’s VoxCPM2 ranks first because its Hugging Face publication date, 2026-04-03, is the newest among the eligible candidates; it has no public benchmark yet, so the placement does not establish audiobook quality superiority. [6]
- Ranking note: no public benchmark scores any candidate, so the ordering is newest release first, then downloads. Eligibility admits only text-to-speech models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days under Hugging Face’s text-to-speech pipeline tag. [1][3][5][7][9][11][14]
- fishaudio’s s2-pro has 9.1 GB of BF16 weights and uses the fish-audio-research-license; openbmb’s VoxCPM2 has 4.6 GB of BF16 weights and uses Apache-2.0. Compare the license terms alongside the weight footprint when planning local audiobook production. [10][6]
- Qwen’s Qwen3-TTS-12Hz-1.7B-CustomVoice and Qwen3-TTS-12Hz-0.6B-CustomVoice both use Apache-2.0, with BF16 weights of 4.5 GB and 2.5 GB respectively. Their shared publication date makes downloads the ordering tiebreak, rather than model size or demonstrated narration quality. [4][12]
- microsoft’s VibeVoice-Realtime-0.5B provides 2.0 GB of BF16 weights under MIT; mistralai’s Voxtral-4B-TTS-2603 provides 8.0 GB of full-precision weights under cc-by-nc-4.0. Those weight sizes do not establish a verified GPU or system RAM requirement. [8][13]
- hexgrad’s Kokoro-82M and coqui’s XTTS-v2 are established picks, each among its family’s three most-downloaded models, admitted through the exception to the 12-month recency gate; that exemption cannot earn first place. Their licenses are Apache-2.0 and coqui-public-model-license respectively, and download popularity does not establish audiobook quality. [1][3]
How do local text-to-speech models compare on specifications and audiobook benchmarks?
| Model | Org | Params | Quant/VRAM | Released (date) | License | Key benchmark (date) |
|---|---|---|---|---|---|---|
| VoxCPM2 [6] | openbmb | 2.3B [6] | BF16 weights: 4.6 GB [6]; VRAM not published | 2026-04-03 [6] | Apache-2.0 [6] | no public benchmark yet |
| s2-pro [10] | fishaudio | 4.6B [10] | BF16 weights: 9.1 GB [10]; VRAM not published | 2026-03-09 [10] | fish-audio-research-license [10] | no public benchmark yet |
| Qwen3-TTS-12Hz-1.7B-CustomVoice [4] | Qwen | 1.9B [4] | BF16 weights: 4.5 GB [4]; VRAM not published | 2026-01-21 [4] | Apache-2.0 [4] | no public benchmark yet |
| Qwen3-TTS-12Hz-0.6B-CustomVoice [12] | Qwen | 906M [12] | BF16 weights: 2.5 GB [12]; VRAM not published | 2026-01-21 [12] | Apache-2.0 [12] | no public benchmark yet |
| VibeVoice-Realtime-0.5B [8] | microsoft | 1B [8] | BF16 weights: 2.0 GB [8]; VRAM not published | 2025-12-04 [8] | MIT [8] | no public benchmark yet |
| Voxtral-4B-TTS-2603 [13] | mistralai | — | Full-precision weights: 8.0 GB [13]; VRAM not published | 2025-11-17 [13] | cc-by-nc-4.0 [13] | no public benchmark yet |
| Kokoro-82M — established pick [1] | hexgrad | — | — | 2024-12-26 [1] | Apache-2.0 [1] | no public benchmark yet |
| XTTS-v2 — established pick [3] | coqui | — | — | 2023-10-31 [3] | coqui-public-model-license [3] | no public benchmark yet |
Which local text-to-speech models should you consider for audiobooks?
1. VoxCPM2
VoxCPM2 by openbmb ranks first on release recency, with a Hugging Face publication date of April 3, 2026.[6] Its position follows the ordering rule: newest release first, then downloads.[6] For audiobook narration, the caveat is straightforward: no public benchmark yet establishes its performance against the other candidates, so placement does not demonstrate superior narration quality.
VoxCPM2 has 2.3 billion parameters and 4.6 GB of BF16 weights, with an Apache-2.0 license.[6] For local deployment, use that weight size as a storage planning figure; a verified GPU memory or system RAM requirement is unavailable. A specific graphics card recommendation would therefore require an explicitly estimated memory budget or a measured deployment.
Use VoxCPM2 as a candidate for an audiobook audition. Before committing a manuscript, generate representative passages containing dialogue, uncommon names and changes in narrative tone, then check pronunciation, voice consistency and pacing.
2. s2-pro
s2-pro by fishaudio ranks second in this release-ordered list: its Hugging Face publication date is March 9, 2026,[10] behind VoxCPM2 by openbmb, published April 3, 2026.[6] With no public benchmark yet, its position reflects release timing rather than demonstrated audiobook quality.
The model has 4.6 billion parameters and 9.1 GB of BF16 weights.[10] Treat the weight footprint as a starting point for local hardware planning, not a verified GPU memory requirement. A specific GPU recommendation or total RAM requirement cannot be established from that figure alone. Validate memory use with your intended inference setup and narration workload before buying hardware.
For audiobook work, consider s2-pro for a local narration trial: check pronunciation, pacing and voice consistency across chapter-length passages before committing to production. The licensing caveat deserves attention: s2-pro uses the fish-audio-research-license.[10] Review its terms against your intended audiobook use and distribution before adopting it.
3. Qwen3-TTS-12Hz-1.7B-CustomVoice
Qwen3-TTS-12Hz-1.7B-CustomVoice by Qwen ranks third under newest-release-first ordering, with downloads breaking same-date ties.[4][6][10][12] Its Hugging Face publication date is January 21, 2026.[4] Its 2,445,964 downloads over the preceding 30 days place it ahead of Qwen’s Qwen3-TTS-12Hz-0.6B-CustomVoice, published on the same date, in the September 25, 2026 snapshot.[4][12] Download counts determine that tie, not audiobook quality.
The model has 1.9 billion parameters and 4.5 GB of BF16 weights, with an Apache-2.0 license.[4] For local hardware planning, treat the weight size as a starting point, not a complete GPU memory requirement. A specific GPU or system RAM recommendation is not established here; budget for runtime overhead and verify memory use with your intended workload.
Consider it for local audiobook evaluation when Apache-2.0 licensing suits your project.[4] The caveat is straightforward: no public benchmark yet. Audition representative passages before committing to a full book; its ranking does not establish narration quality or long-form consistency.
4. Qwen3-TTS-12Hz-0.6B-CustomVoice
Qwen3-TTS-12Hz-0.6B-CustomVoice by Qwen ranks fourth under the ordering of newest release first, then downloads.[6][10][4][12] Its Hugging Face publication date matches Qwen’s Qwen3-TTS-12Hz-1.7B-CustomVoice: January 21, 2026.[4][12] The download tiebreak places it behind that sibling, with 973,050 versus 2,445,964 downloads over the preceding 30 days as of September 25, 2026.[4][12]
The model card lists 906 million parameters, 2.5 GB of BF16 weights, and an Apache-2.0 license.[12] For local hardware planning, use that weight size as a starting point, not a total memory requirement. A specific GPU or system RAM capacity cannot be established from the weight footprint alone.
For audiobook work, use it as a candidate for local narration evaluation when its documented weight footprint suits your hardware planning. Audition representative passages before committing to a production workflow. The caveat is no public benchmark yet: its position reflects release timing and download activity, without establishing audiobook narration quality or long-form consistency.
5. VibeVoice-Realtime-0.5B
VibeVoice-Realtime-0.5B by microsoft ranks here by release recency, with a Hugging Face publication date of December 4, 2025.[8] The newer candidates precede it under that ordering; downloads matter only as a tiebreak.[4][6][10][12] Its benchmark status is no public benchmark yet, so the position does not establish an advantage in audiobook narration quality.
The model lists 1 billion parameters and 2.0 GB of BF16 weights, with an MIT license.[8] Use the listed parameter count when planning deployment rather than interpreting “0.5B” in the model name as the total.[8] For local hardware planning, the weight footprint is a starting point, not a complete memory budget. A specific GPU or system-RAM requirement is not established.
For audiobook work, consider it for local narration trials before committing to a full manuscript. Evaluate pronunciation, pacing, and voice consistency across representative passages. The caveat is the lack of a public benchmark demonstrating sustained audiobook performance.
6. Voxtral-4B-TTS-2603
Voxtral-4B-TTS-2603 by mistralai occupies this position under the newest-release-first ordering, using its Hugging Face publication date of 2025-11-17.[13] Its status is “no public benchmark yet,” so the placement does not establish audiobook narration quality. The Voxtral TTS paper is dated 2026-03-26; that paper date should not replace the model’s publication date in this ranking.[14][13]
The full-precision weights total 8.0 GB.[13] Local hardware requirements remain unverified: a specific GPU, VRAM capacity or system RAM minimum cannot be given confidently. Treat the weight size as a download specification, and confirm runtime memory requirements before choosing hardware.
For audiobook work, consider it for local evaluation before committing to production. Check narration on your intended text rather than assuming chapter-length consistency from its placement. The practical caveat is licensing: review the cc-by-nc-4.0 terms against your intended audiobook distribution before adopting the model.[13]
7. Kokoro-82M
Kokoro-82M by hexgrad is an established pick, first published on Hugging Face on December 26, 2024.[1] Its seventh-place position follows the ranking’s placement of eligible recent releases ahead of established exceptions. Kokoro bypasses the recency gate through its established-pick status; no public benchmark yet supports an audiobook-quality comparison against the newer candidates.
Kokoro carries an Apache-2.0 license and recorded 11,695,126 downloads over the last 30 days as of September 25, 2026.[1] Consider it for audiobook production when that license suits your project, then audition representative narration before committing. Popularity serves only as a tiebreaker in this ranking, not evidence of pronunciation accuracy or consistent delivery across chapters.
Local hardware planning remains the caveat: verified weight-storage, runtime-memory and GPU requirements are unavailable for this entry. A specific RAM or graphics-card recommendation would therefore be speculative. Measure memory consumption and generation time on your intended machine before planning a full audiobook workflow.
8. XTTS-v2
XTTS-v2 by coqui ranks here as an established pick: its first Hugging Face publication was 2023-10-31, placing it outside the release window for current candidates.[3] Its position follows the release-date ordering, with no public benchmark yet to justify a higher audiobook ranking. The established-pick exemption admits the model without giving it priority over current releases.
XTTS-v2 recorded 6,850,874 downloads in the preceding 30 days as of 2026-09-25.[3] Downloads indicate adoption, not narration quality. Its license is the coqui-public-model-license.[3] A verified parameter count, weight size, and minimum RAM or VRAM requirement are unavailable for this entry, so a specific local hardware configuration cannot be recommended.
Use XTTS-v2 as an established comparison candidate when evaluating a local audiobook workflow. Before committing, verify hardware requirements and check the license against your intended audiobook use. The caveat is the lack of a public benchmark establishing its audiobook performance.
How much memory do local text-to-speech model weights require?
Reported local text-to-speech weight sizes range from 2.0 GB to 9.1 GB for the candidates with BF16 figures; those figures describe weights, not confirmed hardware memory requirements.[8][10]
Microsoft’s VibeVoice-Realtime-0.5B has 2.0 GB of BF16 weights.[8] Qwen’s Qwen3-TTS-12Hz-0.6B-CustomVoice has 2.5 GB, while Qwen’s Qwen3-TTS-12Hz-1.7B-CustomVoice has 4.5 GB, both in BF16.[12][4] Use the reported weight sizes when planning a download or comparing candidates, rather than treating the parameter labels in model names as memory specifications.
openbmb’s VoxCPM2 has 4.6 GB of BF16 weights, and fishaudio’s s2-pro has 9.1 GB.[6][10] mistralai’s Voxtral-4B-TTS-2603 lists 8.0 GB of full-precision weights; keep that precision label attached when comparing its footprint with the BF16 entries.[13]
For quantization planning, VoxCPM2’s cited 2.3 billion parameters imply 1.15 GB for weights alone at four-bit precision, estimated (params x 0.5 bytes).[6] s2-pro’s cited 4.6 billion parameters similarly imply 2.3 GB, estimated (params x 0.5 bytes).[10] Neither calculation establishes an available quantized release or a tested hardware requirement.
Weight-size figures are unspecified here for hexgrad’s Kokoro-82M and coqui’s XTTS-v2.[1][3] For an audiobook workstation, leave total RAM, GPU memory and hardware fit unconfirmed until the intended runtime provides measured requirements. Avoid turning a weight download size into a promise that a particular GPU can run the model.
What licenses do these text-to-speech models use?
The models use Apache-2.0 [1][4][6][12], MIT [8], coqui-public-model-license [3], fish-audio-research-license [10], and cc-by-nc-4.0 [13], so your audiobook workflow needs a license check for the specific model you choose.
Apache-2.0 covers Kokoro-82M from hexgrad [1], VoxCPM2 from openbmb [6], and both Qwen3-TTS-12Hz-1.7B-CustomVoice [4] and Qwen3-TTS-12Hz-0.6B-CustomVoice from Qwen [12]. For an audiobook project considering those models, record the license alongside the exact checkpoint name in your production documentation.
VibeVoice-Realtime-0.5B from Microsoft uses MIT [8]. Keep that license entry separate from the Apache-licensed candidates when documenting your chosen model and reviewing its terms.
XTTS-v2 from coqui uses coqui-public-model-license [3]. Meanwhile, s2-pro from fishaudio uses fish-audio-research-license [10]. Review each named license directly before committing to a publishing workflow; avoid treating a downloadable model as automatic permission for your intended use.
Voxtral-4B-TTS-2603 from mistralai uses cc-by-nc-4.0 [13]. For paid audiobook production, check the applicable terms before investing time in narration generation. Make the review specific: describe whether you will sell recordings, distribute model files, or offer narration as a service, then check whether the chosen license permits that activity.
How should you choose an audiobook model without public benchmark scores?
Choose an audiobook model by auditioning representative passages on your own hardware, checking its license, and treating release order as a shortlist rather than proof of narration quality.
Compare dialogue, unfamiliar names, abbreviations, and chapter transitions. Listen for pronunciation errors, unstable character voices, awkward pauses, and changes in pacing. Record generation time and memory use locally, then check how easily you can regenerate a sentence without disrupting the surrounding narration.
Ranking note: every candidate has “no public benchmark yet”; the ordering is newest release first, then downloads, with popularity used only to break ties. hexgrad’s Kokoro-82M (hexgrad) and coqui’s XTTS-v2 (coqui) are established picks that bypass the recency gate but cannot take the first position through that exemption.[1][3]
The ranking admits only text-to-speech models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days under Hugging Face’s text-to-speech pipeline tag.[1][3]
openbmb’s VoxCPM2 (openbmb) takes the first position because its publication date, 2026-04-03, makes it the newest eligible release.[4][6][8][10][12][13]
For hardware planning, VoxCPM2 (openbmb) lists BF16 weights of 4.6 GB, while fishaudio’s s2-pro (fishaudio) lists 9.1 GB; validate actual memory requirements before choosing hardware.[6][10] License review belongs alongside listening tests: VoxCPM2 (openbmb) uses Apache-2.0, while s2-pro (fishaudio) uses fish-audio-research-license.[6][10] Confirm that the terms cover your intended audiobook distribution before committing to production.
Frequently Asked Questions
Which text-to-speech model ranks first for local audiobooks?
openbmb’s VoxCPM2 ranks first because its April 3, 2026 publication date places it first under release-recency ordering [6]. The ranking continues with fishaudio’s s2-pro [10], Qwen’s Qwen3-TTS-12Hz-1.7B-CustomVoice [4], Qwen’s Qwen3-TTS-12Hz-0.6B-CustomVoice [12], microsoft’s VibeVoice-Realtime-0.5B [8], mistralai’s Voxtral-4B-TTS-2603 [13], hexgrad’s Kokoro-82M (established pick) [1], and coqui’s XTTS-v2 (established pick) [3]. Placement does not establish audiobook narration quality.
How does the ranking choose and order models?
Ranking note: newest release first, then downloads. The ranking admits only text to speech models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days, using Hugging Face’s text-to-speech pipeline tag [1][3][5][7][9][11][14]. Established picks bypass the 12-month gate as members of their family’s three most-downloaded models, follow the same ordering rule, and cannot take first place through that exemption [1][3].
How much GPU memory do I need to run these models locally?
Weight files provide a starting point: VoxCPM2 has 4.6 GB of BF16 weights [6]; s2-pro has 9.1 GB [10]; Qwen3-TTS-12Hz-1.7B-CustomVoice has 4.5 GB [4]; Qwen3-TTS-12Hz-0.6B-CustomVoice has 2.5 GB [12]; and VibeVoice-Realtime-0.5B has 2.0 GB [8]. Voxtral-4B-TTS-2603 has 8.0 GB of full-precision weights [13]. A GPU recommendation would require a separate runtime memory measurement or an explicitly labeled estimate; weight-file sizes alone do not establish memory requirements.
What licenses do the audiobook candidates use?
Apache-2.0 covers VoxCPM2 [6], Qwen3-TTS-12Hz-1.7B-CustomVoice [4], Qwen3-TTS-12Hz-0.6B-CustomVoice [12], and Kokoro-82M [1]. VibeVoice-Realtime-0.5B uses MIT [8]. The remaining licenses are fish-audio-research-license for s2-pro [10], cc-by-nc-4.0 for Voxtral-4B-TTS-2603 [13], and coqui-public-model-license for XTTS-v2 [3]. Check the applicable license against your intended audiobook production and distribution before committing to a model.
Do benchmarks show which model produces better audiobook narration?
Every candidate is marked “no public benchmark yet”; no public leaderboard scores these candidates against one another [1][3][4][6][8][10][12][13]. A benchmark-based audiobook winner therefore cannot be named here. Model-card benchmarks must be labeled self-reported whenever used. For a practical evaluation, compare chapter-length narration, pronunciation, pauses, and voice consistency using the same manuscript passages.
Are Kokoro-82M and XTTS-v2 still worth considering?
Kokoro-82M and XTTS-v2 remain established picks under the ranking’s exemption [1][3]. Their Hugging Face publication dates are December 26, 2024 [1], and October 31, 2023 [3], respectively. As of September 25, 2026, their last-30-day downloads were 11,695,126 [1] and 6,850,874 [3]. Download counts indicate adoption rather than audiobook quality; use them only as a ranking tiebreak.
Sources
- hexgrad/Kokoro-82M model card (Hugging Face) — 2026-09-25
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models — 2023-06-13
- coqui/XTTS-v2 model card (Hugging Face) — 2026-09-25
- Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice model card (Hugging Face) — 2026-09-25
- Qwen3-TTS Technical Report — 2026-01-22
- openbmb/VoxCPM2 model card (Hugging Face) — 2026-09-25
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning — 2025-09-29
- microsoft/VibeVoice-Realtime-0.5B model card (Hugging Face) — 2026-09-25
- VibeVoice Technical Report — 2025-08-26
- fishaudio/s2-pro model card (Hugging Face) — 2026-09-25
- Fish Audio S2 Technical Report — 2026-03-09
- Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice model card (Hugging Face) — 2026-09-25
- mistralai/Voxtral-4B-TTS-2603 model card (Hugging Face) — 2026-09-25
- Voxtral TTS — 2026-03-26