Selecting an Open-Source TTS Model for Content Automation
Choose a TTS model by script difficulty, expression, hardware and license
Definition
Open-source TTS selection is a workload-fit decision across pronunciation, expression, language coverage, compute and license constraints. The source compared Qwen3-TTS, VoxCPM2, Higgs TTS 3, Supertonic 3 and Audio8 on an RTX 5090 with 32 GB VRAM; the presented samples were played at 1.1× speed.
Perspectives
Sam Hottman (2026-08-29, YouTube)
For straightforward Korean narration, all five finalists were usable enough that the content format matters more than choosing an overall winner.
| Need observed in this test | Starting point |
|---|---|
| CPU-only or first local setup | Supertonic 3 |
| Technical scripts containing terminal commands | VoxCPM2 performed best on the tested command sample |
| Story content needing anger, joy or whispering | Higgs TTS 3 was the preferred expressive option when compute allowed |
| Balanced general use | Compare Qwen3-TTS, VoxCPM2 and Audio8 against the actual production script |
Do not treat model size as a guarantee of correct reading. Every tested model made mistakes on some combination of dates, money, percentages, version numbers, hardware notation or identifiers. Check model-weight and code licenses separately before turning a model into a paid app, API or service; the source specifically flags Higgs TTS 3 for additional commercial-license review.
How to apply
- Use this comparison when local inference cost, Korean narration or expressive delivery is part of the product or content workflow.
- Test finalists with the hardest real script, not a clean narration sentence; route number- and command-heavy copy through Normalizing TTS Scripts Before Synthesis.
- Start from Voice when deciding whether the bottleneck is model selection, script preparation or downstream review.
- Reject a candidate before quality testing if its compute or commercial-license terms do not fit the intended deployment.
Limits
-
This is one creator's qualitative test, not a blinded listening study or reproducible benchmark.
-
The video does not provide the exact repository URL, model revision, inference settings, measured latency table or audio files in the public description.
-
Automatic Korean captions contain recognition errors; model names and test categories were taken from the video description where possible.
-
License statements and behavior on GPUs other than the reported RTX 5090 were not independently verified.