Best Text-to-Speech Models for Multilingual Speech in 2026: VoxCPM2 and s2-pro

Rankings 2026-09-28 Last updated 2026-09-28 13 min read By Q4KM

Quick Answer

VoxCPM2 (openbmb) is the top pick as of September 2026 because it has the newest release date among the eligible candidates [6]. In order, the ranking is VoxCPM2 (openbmb), s2-pro (fishaudio), Qwen3-TTS-12Hz-1.7B-CustomVoice (Qwen), Qwen3-TTS-12Hz-0.6B-CustomVoice (Qwen), VibeVoice-Realtime-0.5B (microsoft), Voxtral-4B-TTS-2603 (mistralai), Kokoro-82M (hexgrad), and XTTS-v2 (coqui) [6][10].

Key Takeaways

How do local text-to-speech models compare on specifications and public benchmarks?

Model Org Params Quant/VRAM Released (date) License Key benchmark (date)
VoxCPM2 [6] openbmb [6] 2.3B [6] BF16 weights: 4.6 GB [6] 2026-04-03 [6] Apache-2.0 [6] no public benchmark yet
s2-pro [10] fishaudio [10] 4.6B [10] BF16 weights: 9.1 GB [10] 2026-03-09 [10] fish-audio-research-license [10] no public benchmark yet
Qwen3-TTS-12Hz-1.7B-CustomVoice [4] Qwen [4] 1.9B [4] BF16 weights: 4.5 GB [4] 2026-01-21 [4] Apache-2.0 [4] no public benchmark yet
Qwen3-TTS-12Hz-0.6B-CustomVoice [12] Qwen [12] 906M [12] BF16 weights: 2.5 GB [12] 2026-01-21 [12] Apache-2.0 [12] no public benchmark yet
VibeVoice-Realtime-0.5B [8] microsoft [8] 1B [8] BF16 weights: 2.0 GB [8] 2025-12-04 [8] MIT [8] no public benchmark yet
Voxtral-4B-TTS-2603 [13] mistralai [13] — Full-precision weights: 8.0 GB [13] 2025-11-17 [13] cc-by-nc-4.0 [13] no public benchmark yet
Kokoro-82M — established pick [1] hexgrad [1] — — 2024-12-26 [1] Apache-2.0 [1] no public benchmark yet
XTTS-v2 — established pick [3] coqui [3] — — 2023-10-31 [3] coqui-public-model-license [3] no public benchmark yet

Which local text-to-speech models should you consider for multilingual speech?

1. VoxCPM2

VoxCPM2 by openbmb ranks first under the newest-release-first rule, with a Hugging Face publication date of April 3, 2026.[6] VoxCPM2 has no public benchmark yet, so its position reflects release recency rather than demonstrated superiority in multilingual speech. Treat it as a candidate for local evaluation, with language quality still to verify against your own scripts.

VoxCPM2 has 2.3 billion parameters and 4.6 GB of BF16 weights, distributed under the Apache-2.0 license.[6] For local deployment, distinguish weight storage from total runtime memory. A specific GPU or system-RAM recommendation cannot be made from the published weight size alone; measure memory use with your chosen runtime before committing hardware.

The practical use case is an evaluation workflow: generate representative passages in your target languages, check pronunciation and consistency, and confirm acceptable performance on your machine. The caveat is the absence of a public benchmark establishing multilingual quality.

2. s2-pro

s2-pro by fishaudio ranks second by release recency: its Hugging Face publication date is March 9, 2026,[10] behind VoxCPM2 by openbmb, published April 3, 2026.[6] Its benchmark status is “no public benchmark yet,” so the position does not establish superior multilingual speech quality. The accompanying paper is fishaudio’s Fish Audio S2 Technical Report.[11]

The model has 4.6 billion parameters and 9.1 GB of BF16 weights.[10] For local hardware planning, use that weight footprint as a starting point, not a complete GPU-memory requirement. A specific GPU or RAM recommendation remains unverified; measure runtime memory with your intended inference setup before committing to hardware.

For multilingual work, use s2-pro as an evaluation candidate and check pronunciation and intelligibility in your target languages. The deployment caveat is its fish-audio-research-license: review the terms against your intended use before adoption.[10]

3. Qwen3-TTS-12Hz-1.7B-CustomVoice

Qwen3-TTS-12Hz-1.7B-CustomVoice by Qwen ranks third under the ordering rule of newest release first, then downloads.[4][6][10][12] Published on January 21, 2026,[4] it follows openbmb’s VoxCPM2 and fishaudio’s s2-pro by release date.[6][10] Its 2,445,964 downloads over the preceding 30 days, measured on September 25, 2026, break the release-date tie with Qwen’s Qwen3-TTS-12Hz-0.6B-CustomVoice.[4][12]

The model has 1.9 billion parameters and 4.5 GB of BF16 weights, with an Apache-2.0 license.[4] For local hardware planning, treat that weight size as a storage figure, not a complete runtime memory requirement. A specific GPU or system-RAM recommendation cannot be established from the published weight size alone.

Use it as a candidate for local multilingual speech evaluation when Apache-2.0 licensing fits your project.[4] The caveat is comparison quality: no public benchmark yet establishes its position against the other ranked candidates. Its placement reflects release timing and the download tiebreak, rather than demonstrated multilingual speech superiority.

4. Qwen3-TTS-12Hz-0.6B-CustomVoice

Qwen3-TTS-12Hz-0.6B-CustomVoice by Qwen ranks here under newest release first, then downloads.[12] Its publication date matches its larger sibling’s: January 21, 2026.[4][12] The download tiebreak places it behind that sibling, with 973,050 versus 2,445,964 downloads over the preceding 30 days as of September 25, 2026.[4][12] Multilingual quality remains unranked: no public benchmark yet.

The model has 906 million parameters and 2.5 GB of BF16 weights, with an Apache-2.0 license.[12] For local hardware planning, treat that weight size as a starting point, not a complete memory requirement. A supported minimum RAM capacity or GPU configuration is not specified, so a particular hardware fit cannot be promised.

A practical use is local multilingual evaluation when Apache-2.0 licensing and a smaller weight footprint matter: the larger Qwen variant lists 4.5 GB of BF16 weights.[4][12] The caveat is the absent public benchmark; evaluate pronunciation and intelligibility in your target languages before choosing it for deployment.

5. VibeVoice-Realtime-0.5B

VibeVoice-Realtime-0.5B by microsoft ranks here because ordering follows newest release first, then downloads; its Hugging Face publication date is December 4, 2025.[8] The model has no public benchmark yet, so this position does not establish multilingual speech quality or latency relative to the other candidates.

The listed parameter count is 1B, with 2.0 GB of BF16 weights.[8] For local hardware planning, treat that weight size as a starting point, not a complete memory requirement. A specific GPU or system RAM recommendation is not established, so confirm runtime memory use before committing hardware.

Consider VibeVoice-Realtime-0.5B for a local evaluation where the MIT license suits your deployment requirements.[8] Evaluate speech in your intended languages before selecting it for production. The practical caveat is the benchmark gap: multilingual pronunciation, intelligibility and responsiveness remain evaluation questions rather than demonstrated comparative strengths.

6. Voxtral-4B-TTS-2603

Voxtral-4B-TTS-2603 by mistralai ranks sixth under the release-date ordering, with a Hugging Face publication date of November 17, 2025.[13] Its position reflects release timing rather than demonstrated multilingual speech quality: no public benchmark yet. The accompanying paper, “Voxtral TTS,” was published on March 26, 2026.[14]

The full-precision weights occupy 8.0 GB, and the model uses the cc-by-nc-4.0 license.[13] Treat that weight size as a starting point for local hardware planning, not a total memory requirement. A specific GPU or RAM recommendation remains unverified; measure runtime memory with your intended workload before committing hardware.

Consider Voxtral for a local multilingual speech evaluation where you can check pronunciation and intelligibility in your target languages. The practical caveat is the lack of a public benchmark: its ranking does not establish a quality advantage for your language mix.

7. Kokoro-82M

Kokoro-82M by hexgrad ranks seventh as an established pick, with a Hugging Face publication date of December 26, 2024.[1] Its established status permits inclusion outside the recency window; its position follows publication recency, with downloads used only to break ties. Kokoro has no public benchmark yet, so its placement does not establish multilingual speech quality.

Kokoro uses the Apache-2.0 license and recorded 11,695,126 downloads over the preceding 30 days as of September 25, 2026.[1] Download volume indicates adoption, not pronunciation accuracy or consistency across languages. Treat Kokoro as an established candidate for local evaluation when the license suits your project.

For local hardware planning, verify the runtime’s weight precision and peak memory requirements before choosing a GPU or allocating RAM. The model name alone does not establish either requirement. Its practical use here is an evaluation baseline: test your target languages, names and technical vocabulary before selecting it for production.

8. XTTS-v2

XTTS-v2 by coqui ranks here as an established pick, with a Hugging Face publication date of October 31, 2023.[3] Its exemption keeps it eligible, but does not move it ahead of newer releases. The ordering follows newest release first, then downloads; XTTS-v2 has no public benchmark yet to establish a comparative multilingual quality advantage.

The model uses the coqui-public-model-license and recorded 6,850,874 Hugging Face downloads over the preceding 30 days as of September 25, 2026.[3] Those downloads indicate adoption, not speech quality. Treat XTTS-v2 as an established comparison candidate for a local multilingual evaluation, checking pronunciation and consistency in your target languages before selecting it.

Hardware requirements remain unverified for this comparison: a parameter count, weight size and measured runtime memory footprint are unavailable. A specific GPU or RAM recommendation would therefore be speculative. Check memory consumption on your intended hardware, and review the model’s license before committing to deployment.[3]

What do model weight sizes tell you about local hardware requirements?

Model weight sizes tell you how much data the weights contain, giving you a starting point for local hardware planning rather than a complete memory requirement. Use the published weight size to compare storage needs; do not treat it as a guarantee that a GPU with matching memory can run the model.

Microsoft’s VibeVoice-Realtime-0.5B has BF16 weights of 2.0 GB.[8] Qwen’s Qwen3-TTS-12Hz-0.6B-CustomVoice has BF16 weights of 2.5 GB, while its Qwen3-TTS-12Hz-1.7B-CustomVoice has BF16 weights of 4.5 GB.[12][4] Those figures describe the distributed weights, not measured peak memory during speech generation.

openbmb’s VoxCPM2 has BF16 weights of 4.6 GB, and fishaudio’s s2-pro has BF16 weights of 9.1 GB.[6][10] mistralai’s Voxtral-4B-TTS-2603 lists full-precision weights of 8.0 GB.[13] Preserve those precision labels when comparing downloads: a weight size belongs to a particular representation, not simply to a model name.

For a hardware purchase, look for a measured memory requirement or documented configuration for the exact model and runtime you intend to use. Account for runtime memory beyond the weights before choosing a GPU or allocating system RAM. Weight sizes alone also cannot establish generation speed, multilingual pronunciation quality, or whether a particular quantized version is available. Treat them as a planning input, then verify the complete inference setup.

Which licenses apply to these text-to-speech models?

The listed text-to-speech models use Apache-2.0 [1][4][6][12], MIT [8], fish-audio-research-license [10], cc-by-nc-4.0 [13], and coqui-public-model-license [3], depending on the model.

Apache-2.0 applies to hexgrad’s Kokoro-82M (hexgrad) [1], openbmb’s VoxCPM2 (openbmb) [6], and Qwen’s Qwen3-TTS-12Hz-1.7B-CustomVoice (Qwen) [4] and Qwen3-TTS-12Hz-0.6B-CustomVoice (Qwen) [12]. The Qwen variants share the same license, so choosing between them does not change the listed license. [4][12]

Microsoft’s VibeVoice-Realtime-0.5B (microsoft) uses MIT. [8] Its license differs from the Apache-2.0 license attached to the Qwen models, even though all appear in the same local speech shortlist. [4][8][12]

fishaudio’s s2-pro (fishaudio) uses fish-audio-research-license. [10] mistralai’s Voxtral-4B-TTS-2603 (mistralai) uses cc-by-nc-4.0. [13] coqui’s XTTS-v2 (coqui) uses coqui-public-model-license. [3] Each of those entries needs its own license review; the license attached to another candidate does not establish permission for your deployment.

For deployment planning, record the exact model repository and its license alongside the downloaded weights. Check the linked license terms against your intended use, including commercial deployment, redistribution and modification. Treat local execution as a hardware choice, and license review as a separate decision about how you may use and distribute the model.

How are models ranked when no public benchmark scores the candidates?

When no public benchmark scores the candidates, the ordering is newest release first, then downloads.[6][10][4][12][8][13][1][3] Each candidate is marked “no public benchmark yet”; the order does not establish multilingual speech quality.

The ranking admits only text-to-speech models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days under Hugging Face’s text-to-speech pipeline tag.[1][3] Eligible releases normally fall within the last 12 months.[6][10][4][12][8][13]

openbmb’s VoxCPM2 (openbmb) ranks first because its publication date is the newest among the eligible candidates.[6][10][4][12][8][13] The remaining order is fishaudio’s s2-pro (fishaudio),[10] Qwen’s Qwen3-TTS-12Hz-1.7B-CustomVoice (Qwen),[4] Qwen’s Qwen3-TTS-12Hz-0.6B-CustomVoice (Qwen),[12] microsoft’s VibeVoice-Realtime-0.5B (microsoft),[8] mistralai’s Voxtral-4B-TTS-2603 (mistralai),[13] hexgrad’s Kokoro-82M (hexgrad),[1] and coqui’s XTTS-v2 (coqui).[3]

The Qwen releases share a publication date, so downloads break their tie.[4][12] Popularity serves only as a tiebreaker; parameter count and weight size do not determine placement. Kokoro-82M (hexgrad) and XTTS-v2 (coqui) are established picks: each belongs to its family’s three most-downloaded models and bypasses the recency gate.[1][3] Established picks follow the same ordering rule and cannot take first place through that exemption.

A public leaderboard may compare only candidates it actually scores, and an older scored release cannot displace a newer unscored release. Model-card results must be labeled “self-reported.” For multilingual deployment, treat this order as an evaluation queue: the absence of comparable scores leaves language quality, pronunciation and speaker consistency unresolved.

Frequently Asked Questions

Which multilingual text-to-speech model should I try first?

VoxCPM2 (openbmb) ranks first because its Hugging Face publication date, 2026-04-03, puts it first under the release-date ordering.[6] Its status is “no public benchmark yet,” so that position does not establish superior multilingual speech quality. Treat it as a starting point for evaluation with your target languages.

How does the ranking decide which models qualify?

The ordering basis is: newest release first, then downloads. No public benchmark scores any candidate. Scope: this ranking admits only text to speech models from labs with a published paper or leaderboard record, or other publishers above 10,000 downloads in the last 30 days (Hugging Face pipeline tags: text-to-speech).[1][3] Established picks bypass the 12-month recency gate; that exemption cannot give them first place.[1][3]

What is the full model order?

The order is VoxCPM2 (openbmb);[6] s2-pro (fishaudio);[10] Qwen3-TTS-12Hz-1.7B-CustomVoice (Qwen);[4] Qwen3-TTS-12Hz-0.6B-CustomVoice (Qwen);[12] VibeVoice-Realtime-0.5B (microsoft);[8] Voxtral-4B-TTS-2603 (mistralai);[13] Kokoro-82M (hexgrad), an established pick;[1] and XTTS-v2 (coqui), an established pick.[3] Read this as an evaluation order based on publication dates and downloads, rather than a measured comparison of multilingual speech quality.

How much memory do I need to run these models locally?

Published weight sizes provide a starting point: VibeVoice-Realtime-0.5B has 2.0 GB of BF16 weights;[8] Qwen3-TTS-12Hz-0.6B-CustomVoice has 2.5 GB;[12] Qwen3-TTS-12Hz-1.7B-CustomVoice has 4.5 GB;[4] VoxCPM2 has 4.6 GB;[6] and s2-pro has 9.1 GB.[10] Voxtral-4B-TTS-2603 lists 8.0 GB of full-precision weights.[13] Treat those as weight sizes, not verified GPU or system RAM requirements. Check runtime requirements before buying hardware.

Which licenses should I check before deploying a model?

VoxCPM2, Qwen3-TTS-12Hz-1.7B-CustomVoice, Qwen3-TTS-12Hz-0.6B-CustomVoice and Kokoro-82M list Apache-2.0.[6][4][12][1] VibeVoice-Realtime-0.5B lists MIT.[8] Different terms apply to s2-pro under fish-audio-research-license,[10] XTTS-v2 under coqui-public-model-license,[3] and Voxtral-4B-TTS-2603 under cc-by-nc-4.0.[13] Read the applicable license against your intended deployment before committing to a model, particularly when planning a commercial product.

How should I compare pronunciation and speech quality across languages?

Every candidate has “no public benchmark yet” under this ranking. Qwen publishes the Qwen3-TTS Technical Report,[5] and fishaudio publishes the Fish Audio S2 Technical Report.[11] Treat any model-card benchmark as self-reported. For your evaluation, use matching passages in each target language, including names, abbreviations and language switches. Listen for pronunciation errors and omitted words before choosing a deployment candidate.

Sources

  1. hexgrad/Kokoro-82M model card (Hugging Face) — 2026-09-25
  2. StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models — 2023-06-13
  3. coqui/XTTS-v2 model card (Hugging Face) — 2026-09-25
  4. Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice model card (Hugging Face) — 2026-09-25
  5. Qwen3-TTS Technical Report — 2026-01-22
  6. openbmb/VoxCPM2 model card (Hugging Face) — 2026-09-25
  7. VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning — 2025-09-29
  8. microsoft/VibeVoice-Realtime-0.5B model card (Hugging Face) — 2026-09-25
  9. VibeVoice Technical Report — 2025-08-26
  10. fishaudio/s2-pro model card (Hugging Face) — 2026-09-25
  11. Fish Audio S2 Technical Report — 2026-03-09
  12. Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice model card (Hugging Face) — 2026-09-25
  13. mistralai/Voxtral-4B-TTS-2603 model card (Hugging Face) — 2026-09-25
  14. Voxtral TTS — 2026-03-26

Get these models on a hard drive

Skip the downloads. Browse our catalog of 985+ commercially-licensed AI models, available pre-loaded on high-speed drives.

Browse Model Catalog