Internal · TTS test bench
TTS Lab
One sentence, every model. Tune each voice individually and compare by ear — costs shown per clip as you type.
Kokoro-82M via OpenRouter
Open-source budget voice. Chinese is an added language — fine, but not native prosody. No tuning: what you hear is what you get.
This is everything OpenRouter exposes for Kokoro — voice choice only (it even hangs on a speed param). Output fixed to mp3.
Qwen3-TTS (open-source) via fal.ai
The native-Mandarin family, open-sourced by Alibaba and hosted on fal. Takes an optional free-text prompt plus full sampling controls. Voice cloning possible later via fal's clone endpoint.
Sampling & clone settings (their defaults shown)
ElevenLabs Multilingual v2 direct API
Premium voices, very natural — but English-first doing Chinese, so listen for accent. Tune with the dials, or paste any voice ID from the voice library to try it.
v3 only honours stability (its API ignores the other dials — they apply to the v2-family models). v3 also reads inline tags in the sentence, e.g. [slowly] [warmly]. Language code applies to Turbo/Flash/v3 only. Not exposed (plumbing, not voice): output format (fixed mp3), streaming, previous/next-text continuity, pronunciation dictionaries.
Other options — can wire in on request
- Alibaba's commercial Qwen3-TTS-Flash — the tuned flagship of the family (~$12 / 1M, free-text instructions). Blocked for now: Alibaba Cloud signup wouldn't go through; the open-source version above is its sibling.
- OpenAI gpt-4o-mini-tts — takes free-text voice instructions; ~$12 / 1M chars, pay-per-use. Needs a direct OpenAI key (OpenRouter doesn't host it).
- CosyVoice 2 (Alibaba, open model) — native Mandarin, very strong tones; ~$1–3 / 1M via Replicate or Fal, pay-per-use.
- MiniMax Speech-02 — Chinese provider, excellent Mandarin; pay-per-use API.
If Kokoro / Qwen / ElevenLabs don't satisfy, say the word and any of these gets a card here.