Qwen/Qwen/Qwen3-TTS-VoiceDesign
模型介绍
本条目的官方说明为英文,以下内容保留原文,关键章节标题已中文标注。
Qwen/Qwen3-TTS-VoiceDesign
● Qwen3-TTS-VoiceDesign is a voice design variant of Qwen3-TTS by Alibaba’s Qwen team. Instead of selecting from preset voices, you describe the voice you want in natural language — and the model generates speech in that voice. Key capabilities: – Natural language voice control — describe any voice with free text (e.g. “a deep male voice with a calm, authoritative presence”, “a young cheerful female with a warm and friendly tone”) – 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese – Streaming support — real-time PCM streaming – Multiple output formats — WAV, MP3, FLAC, PCM Built on the same 1.7B parameter architecture as Qwen3-TTS, using discrete multi-codebook language modeling and a custom 12Hz acoustic tokenizer for high-quality end-to-end speech synthesis.
- 上下文长度:未知
- 最大输出:未知
- 标签:tts
- 提供方:Qwen
- 来源:DeepInfra
适合场景
暂无可靠数据。
已知限制
- 本站未登记确切发布日期,版本时间线可能不完整。
- 本站未登记上下文长度,长文本能力请以官方文档为准。
- 本页评测与硬件数据为区间估算,未做统一环境实测,不能替代官方评测报告。
评测
暂无可靠数据。本站只收录标注了「评测来源、模型版本、是否官方数据、测试日期」的成绩, 不混合不同设置下的分数,也不把单一 Benchmark 解释为整体能力。