Creator · M.
Last updated · Sep 5, 2026
文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh)
Do not auto-install
Creator · M.
Last updated · Sep 5, 2026
文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh)
Do not auto-install
Creator · M.
Last updated · Sep 5, 2026
文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh)
Do not auto-install
Creator · M.
Last updated · Sep 5, 2026
文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh)
Do not auto-install
Install targets
Codex install prompt
Install the "text-to-speech" agent skill from https://github.com/wlzh/skills/tree/main/text-to-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wlzh-text-to-speech","task":"Install text-to-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wlzh/skills --skill text-to-speech
Maintenance
fresh
9d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
612
75/100 Quality · 66/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
612 GitHub stars
Repo activity
612 stars, 75 forks
Maintenance
9d since push
License
MIT
Install
npx skills add wlzh/skills --skill text-to-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wlzh/skills --skill text-to-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wlzh-text-to-speech/install
Agent should check
Copy prompt
Task: Use text-to-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install
Install command: npx skills add wlzh/skills --skill text-to-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wlzh-text-to-speech/install
LLM text format
/api/skills/wlzh-text-to-speech/install?format=text
Find alternatives
/api/skills/search?q=text-to-speech&limit=3
Agent prompt
Use text-to-speech for this task. Review https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install, then install with: npx skills add wlzh/skills --skill text-to-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wlzh-text-to-speech
LLM text
/api/registry/manifest/wlzh-text-to-speech?format=text
Install alias
/api/registry/install/wlzh-text-to-speech
Recommend
/api/registry/recommend?task=Use%20text-to-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
GitHub automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
GitHub automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO612 GitHub stars
Stars/forks activity
INFO612 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS9d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Manage repositories
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: text-to-speech description: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) version: 3.6.0 changelog: - 2026-08-01: v3.6.0 新增并默认启用 friendly_tutorial 整期表达档,面向普通大众的免费教程保持正常语速、自然口语和固定声线 - 2026-08-01: v3.5.0 新增 MiniMax 词级时间戳 sidecar;只保存校验后的时间段,不保存短期签名下载 URL;Edge/Kokoro 显式拒绝该参数 - 2026-08-01: v3.4.0 新增 MiniMax 声音克隆上传/创建/激活费用门禁、本机 0600 克隆音色档,以及整期固定 commercial_narration 表达档;克隆音色或表达档变化会使下游缓存失效 - 2026-07-18: v3.3.1 MiniMax 语境适配移除逐 beat emotion 注入,避免同一场景语气跳变;保留 speed/volume/pitch 轻量调整并同步测试文档 - 2026-07-18: v3.3.0 MiniMax 默认音色改为 Chinese (Mandarin)_Reliable_Executive;新增 MiniMax 专属语境适配层,按开场、总结、解释、步骤、提醒、资源、结论和关注引导自动微调表达;Edge/Kokoro 行为不变 - 2026-07-17: v3.2.0 新增 MiniMax TTS 引擎并设为默认,默认音色 male-qn-jingying(精英青年)、语速 1.0;API Key 只读取 MINIMAX_API_KEY 环境变量;保留 Kokoro/Edge 可配置切换 - 2026-05-17: v3.1.0 强化 localhost Kokoro 代理绕过规则——curl/requests 直连本地服务默认必须 NO_PROXY,不允许先走代理失败后重试
author: M. ---
# Text-to-Speech Skill
将文本转换为语音。默认使用 MiniMax TTS(在线高质量中文配音),并保留 Kokoro TTS v1.1-zh(本地 Docker,102 个中文音色)和 Edge TTS(在线)作为可配置后备。
## 引擎对比
| 特性 | MiniMax TTS | Kokoro TTS v1.1-zh | Edge TTS | |------|-------------|-------------------|----------| | 质量 | 默认推荐,中文短视频旁白更自然 | 本地可用、接近真人 | 标准 Neural 语音 | | 网络 | 需要 MiniMax API | 不需要(本地 Docker) | 需要网络连接 | | 默认音色 | `Chinese (Mandarin)_Reliable_Executive`(可靠高管) | `zm_009` | `zh-CN-YunyangNeural` | | 语速调节 | speed 参数,默认 1.0 | speed 参数 | rate/pitch/volume | | 前提 | `MINIMAX_API_KEY` 环境变量 | Docker 容器需运行 | 安装 `edge-tts` | | 配置值 | `minimax` | `kokoro` | `edge` |
## 使用说明
```bash # 默认使用 MiniMax TTS(当前配置) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py <文本文件>
# 指定引擎 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine minimax python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine kokoro python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine edge
# 指定声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -v zf_094
# 指定输出文件 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3
# MiniMax 专属:同时保存经过校验的词级时间戳 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt \ -o output.mp3 --subtitle-output output.subtitles.json
# 调整语速(MiniMax/Kokoro) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --speed 1.2
# 列出所有可用声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py --list-voices ```
## MiniMax TTS
默认配置:
- 引擎:`tts_engine = "minimax"` - 模型:`speech-2.8-hd` - 音色:`Chinese (Mandarin)_Reliable_Executive`(可靠高管) - 语速:`1.0` - 输出:MP3
### MiniMax 词级时间戳
`--subtitle-output <path>` 会为 MiniMax 请求启用词级时间戳,下载官方 sidecar 后校验文本、字符范围和时间单调性,再写入本地 JSON。sidecar 只保留 `provider/model/voice/subtitle_type/segments`,不会记录官方返回的短期签名 URL。该参数属于 MiniMax 专属能力;Edge/Kokoro 会失败关闭,避免下游误把估算时间当成官方时间戳。
### 整期表达一致性(默认)
`minimax_tts.delivery_consistency` 默认启用 `friendly_tutorial`。它用于面向普通大众的免费教程:像真人耐心讲解,允许正文自然使用“好、其实、这里呢、别急”等连接词,但不靠逐句变调制造表演感。同一期视频内所有句子固定使用相同的音色、速度、音量和音调,避免逐 beat 割裂。
- 默认档:`friendly_tutorial`,speed `1.0`、volume `1.0`、pitch `0` - 显式选择:`--delivery-profile friendly_tutorial` - 保留档:`commercial_narration`,仅用于公告、品牌声明等确需正式表达的内容 - 克隆音色选择顺序:`--voice` > `MINIMAX_VOICE_ID` > 已激活的本机克隆档 > 仓库系统音色 - 本机档:`~/.config/duanku/minimax-voice.json`,必须为 `0600`,不提交仓库
### MiniMax 专属语境适配(兼容回退)
`minimax_tts.context_adaptation` 只在 MiniMax 分支生效。Edge 和 Kokoro 不读取此配置,也不会改变原请求参数。
- 仅当 `delivery_consistency` 关闭时,根据文本自动识别 `opening`、`summary`、`explanation`、`instruction`、`warning`、`resource`、`conclusion`、`call_to_action`、`neutral` - 语境档只轻量调整 MiniMax 的语速、音量和音高,不修改原文、不注入 emotion、不自动插入声音标签 - 未命中规则时使用 `explanation` - 可用 `--context warning` 等参数显式覆盖自动识别;选择 Edge/Kokoro 时该参数被忽略 - 调用方无需理解语境规则。视频流水线只需继续提交文本并使用返回音频,音频时长仍由下游实测
```bash # 自动识别语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt
# MiniMax 显式指定风险提醒语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --context warning ```
### MiniMax 声音克隆
声音克隆分为不收费的创建阶段和首次 TTS 激活阶段。首次使用新克隆音色合成会产生官方克隆费及试听字符费,脚本要求 quote ID 和精确金额确认,不能用布尔参数绕过。
**引擎边界**:克隆 profile、克隆 `voice_id`、`--delivery-profile` 和克隆样本门禁只允许在实际 `tts_engine=minimax` 时读取。切换到 Edge 或 Kokoro 后必须完全跳过这些配置,即使本机仍保留已激活的 MiniMax profile,也不得影响非 MiniMax 的音色、请求参数或预检结果。
```bash # 样本检查,不联网、不收费 python3 scripts/minimax_voice_clone.py inspect --sample /path/to/source.m4a
# 输出激活报价,不联网、不收费 python3 scripts/minimax_voice_clone.py quote \ --voice-id DuankuNarrator20260801 \ --text "试听文案"
# 上传并创建克隆,不执行 TTS;必须确认拥有声音授权 python3 scripts/minimax_voice_clone.py clone \ --sample /path/to/source.m4a \ --voice-id DuankuNarrator20260801 \ --rights-confirmed
# 首次付费激活并生成试听,必须使用上一步 quote 的 ID 和金额 python3 scripts/minimax_voice_clone.py activate \ --text "试听文案" \ --output /path/to/preview.mp3 \ --confirm-quote-id '<quote_id>' \ --confirm-amount-usd '<estimated_total_usd>' ```
样本要求:`mp3/m4a/wav`、10 秒至 5 分钟、最大 20 MB。克隆脚本默认请求 MiniMax 降噪和音量归一化。
密钥规则:
- API Key 只从环境变量读取,默认变量名 `MINIMAX_API_KEY` - 克隆 `voice_id`、远程 file ID 和激活记录只保存在本机 `0600` profile - 不要把真实 Key 写进 `config/tts_config.json`、README、命令行参数或提交历史 - 如果要换变量名,只改 `config/tts_config.json` 中的 `minimax_tts.api_key_env`
```bash export MINIMAX_API_KEY="你的本机 key" python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3 ```
## Kokoro TTS v1.1-zh 声音
使用 `--list-voices` 查看完整列表(102 个)。
### 推荐声音 - `zm_009` - 男声(默认) - `zf_094` - 女声(自然温柔) - `zf_001` - 女声 - `zm_050` - 男声
### 英文声音 - `af_maple` - 女声(Maple) - `af_sol` - 女声(Sol) - `bf_vale` - 男声(Vale)
### 声音命名规则 - `zf_XXX` - 中文女声(55 个) - `zm_XXX` - 中文男声(44 个) - `af_`/`bf_` - 英文声音(3 个)
## 启动 Kokoro 服务
Kokoro TTS 需要 Docker 容器运行:
```bash # 启动 cd /Users/m/document/QNSZ/project/kokoro-tts && ./start.sh
# 停止 cd /Users/m/document/QNSZ/project/kokoro-tts && ./stop.sh
# Web UI 试听 # http://localhost:8880/web/ ```
## 核心功能
### 1. 脚本解析 自动识别并移除播客脚本中的注释和标记: - 时间戳:`(00:00)` - BGM 注释:`[BGM渐入:...]` - 舞台指示:`(主播声音:...)` `(停顿 1秒)` - Markdown 标记:`**文本**`
### 2. 中英文混合朗读 v1.1-zh 模型支持中英文混合文本的自然朗读。
### 3. 后处理集成 可选集成 voice-changer skill 进行变声处理。
## 配置文件
配置文件位于:`~/.claude/skills/text-to-speech/config/tts_config.json`
关键配置项: - `tts_engine`: `"minimax"`、`"kokoro"` 或 `"edge"`(默认引擎) - `minimax_tts`: MiniMax 引擎配置(API URL、模型、默认音色、语速;Key 仅从环境变量读取) - `minimax_tts.context_adaptation`: MiniMax 专属语境档和自动识别规则;不影响其他引擎 - `kokoro_tts`: Kokoro 引擎配置(API URL、默认声音、语速) - `edge_tts`: Edge 引擎配置(声音、语速、音调、音量) - `available_voices`: 按引擎分组的可用声音列表
## 工作流程
``` 输入文本/文件 ↓ 脚本解析(移除注释和标记) ↓ MiniMax / Kokoro TTS / Edge TTS 语音合成 ↓ 后处理(voice-changer,可选) ↓ 输出 MP3 文件 ```
## 代理绕过(重要)
Kokoro TTS 运行在 `localhost:8880`。如果系统配置了 HTTP 代理(`http_proxy`/`https_proxy`),请求 localhost 会被代理拦截导致连接失败(curl 返回 HTTP 000)。
**规则**: - Python 脚本已内置 `os.environ.setdefault("no_proxy", "localhost,127.0.0.1")`,通过脚本调用无需额外处理 - 如果 AI 需要直接用 `curl` 测试或调用 Kokoro API,**必须**加 `--noproxy localhost,127.0.0.1` 或设置 `no_proxy=localhost,127.0.0.1` - 如果 AI 直接写 Python `requests.post("http://localhost:8880/...")`,必须设置 `proxies={"http": None, "https": None}`,或使用 `requests.Session(); session.trust_env = False`,并设置 `NO_PROXY/no_proxy=localhost,127.0.0.1,::1` - 禁止不加代理绕过直接 curl/requests localhost;不要先走代理失败再重试,localhost Kokoro 请求默认就必须绕过代理
```bash # 正确:绕过代理 curl --noproxy localhost,127.0.0.1 -X POST http://localhost:8880/v1/audio/speech ...
# 错误:走了代理,返回 HTTP 000 curl -X POST http://localhost:8880/v1/audio/speech ... ```
## 依赖
- MiniMax TTS: `MINIMAX_API_KEY` 环境变量 - Kokoro TTS: Docker(容器运行在 localhost:8880) - Edge TTS: `pip install edge-tts`
## 性能参考
- Kokoro TTS: 1000字约 3-5 秒(本地 Docker CPU) - Edge TTS: 1000字约 10-20 秒(受网络影响)
Source provenance
Decision snapshot
612 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for text-to-speech, ready for a manual X post.
For a repeatable workflow, this is a skill worth shortlisting before another blank prompt. text-to-speech: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) 612 stars https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x
Listing + install path for text-to-speech: https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x Install: npx skills add wlzh/skills --skill text-to-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to M. but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech/audit)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)M.
@m.
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsInstall targets
Codex install prompt
Install the "text-to-speech" agent skill from https://github.com/wlzh/skills/tree/main/text-to-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wlzh-text-to-speech","task":"Install text-to-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wlzh/skills --skill text-to-speech
Maintenance
fresh
9d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
612
75/100 Quality · 66/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
612 GitHub stars
Repo activity
612 stars, 75 forks
Maintenance
9d since push
License
MIT
Install
npx skills add wlzh/skills --skill text-to-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wlzh/skills --skill text-to-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wlzh-text-to-speech/install
Agent should check
Copy prompt
Task: Use text-to-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install
Install command: npx skills add wlzh/skills --skill text-to-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wlzh-text-to-speech/install
LLM text format
/api/skills/wlzh-text-to-speech/install?format=text
Find alternatives
/api/skills/search?q=text-to-speech&limit=3
Agent prompt
Use text-to-speech for this task. Review https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install, then install with: npx skills add wlzh/skills --skill text-to-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wlzh-text-to-speech
LLM text
/api/registry/manifest/wlzh-text-to-speech?format=text
Install alias
/api/registry/install/wlzh-text-to-speech
Recommend
/api/registry/recommend?task=Use%20text-to-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
GitHub automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
GitHub automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO612 GitHub stars
Stars/forks activity
INFO612 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS9d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Manage repositories
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: text-to-speech description: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) version: 3.6.0 changelog: - 2026-08-01: v3.6.0 新增并默认启用 friendly_tutorial 整期表达档,面向普通大众的免费教程保持正常语速、自然口语和固定声线 - 2026-08-01: v3.5.0 新增 MiniMax 词级时间戳 sidecar;只保存校验后的时间段,不保存短期签名下载 URL;Edge/Kokoro 显式拒绝该参数 - 2026-08-01: v3.4.0 新增 MiniMax 声音克隆上传/创建/激活费用门禁、本机 0600 克隆音色档,以及整期固定 commercial_narration 表达档;克隆音色或表达档变化会使下游缓存失效 - 2026-07-18: v3.3.1 MiniMax 语境适配移除逐 beat emotion 注入,避免同一场景语气跳变;保留 speed/volume/pitch 轻量调整并同步测试文档 - 2026-07-18: v3.3.0 MiniMax 默认音色改为 Chinese (Mandarin)_Reliable_Executive;新增 MiniMax 专属语境适配层,按开场、总结、解释、步骤、提醒、资源、结论和关注引导自动微调表达;Edge/Kokoro 行为不变 - 2026-07-17: v3.2.0 新增 MiniMax TTS 引擎并设为默认,默认音色 male-qn-jingying(精英青年)、语速 1.0;API Key 只读取 MINIMAX_API_KEY 环境变量;保留 Kokoro/Edge 可配置切换 - 2026-05-17: v3.1.0 强化 localhost Kokoro 代理绕过规则——curl/requests 直连本地服务默认必须 NO_PROXY,不允许先走代理失败后重试
author: M. ---
# Text-to-Speech Skill
将文本转换为语音。默认使用 MiniMax TTS(在线高质量中文配音),并保留 Kokoro TTS v1.1-zh(本地 Docker,102 个中文音色)和 Edge TTS(在线)作为可配置后备。
## 引擎对比
| 特性 | MiniMax TTS | Kokoro TTS v1.1-zh | Edge TTS | |------|-------------|-------------------|----------| | 质量 | 默认推荐,中文短视频旁白更自然 | 本地可用、接近真人 | 标准 Neural 语音 | | 网络 | 需要 MiniMax API | 不需要(本地 Docker) | 需要网络连接 | | 默认音色 | `Chinese (Mandarin)_Reliable_Executive`(可靠高管) | `zm_009` | `zh-CN-YunyangNeural` | | 语速调节 | speed 参数,默认 1.0 | speed 参数 | rate/pitch/volume | | 前提 | `MINIMAX_API_KEY` 环境变量 | Docker 容器需运行 | 安装 `edge-tts` | | 配置值 | `minimax` | `kokoro` | `edge` |
## 使用说明
```bash # 默认使用 MiniMax TTS(当前配置) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py <文本文件>
# 指定引擎 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine minimax python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine kokoro python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine edge
# 指定声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -v zf_094
# 指定输出文件 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3
# MiniMax 专属:同时保存经过校验的词级时间戳 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt \ -o output.mp3 --subtitle-output output.subtitles.json
# 调整语速(MiniMax/Kokoro) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --speed 1.2
# 列出所有可用声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py --list-voices ```
## MiniMax TTS
默认配置:
- 引擎:`tts_engine = "minimax"` - 模型:`speech-2.8-hd` - 音色:`Chinese (Mandarin)_Reliable_Executive`(可靠高管) - 语速:`1.0` - 输出:MP3
### MiniMax 词级时间戳
`--subtitle-output <path>` 会为 MiniMax 请求启用词级时间戳,下载官方 sidecar 后校验文本、字符范围和时间单调性,再写入本地 JSON。sidecar 只保留 `provider/model/voice/subtitle_type/segments`,不会记录官方返回的短期签名 URL。该参数属于 MiniMax 专属能力;Edge/Kokoro 会失败关闭,避免下游误把估算时间当成官方时间戳。
### 整期表达一致性(默认)
`minimax_tts.delivery_consistency` 默认启用 `friendly_tutorial`。它用于面向普通大众的免费教程:像真人耐心讲解,允许正文自然使用“好、其实、这里呢、别急”等连接词,但不靠逐句变调制造表演感。同一期视频内所有句子固定使用相同的音色、速度、音量和音调,避免逐 beat 割裂。
- 默认档:`friendly_tutorial`,speed `1.0`、volume `1.0`、pitch `0` - 显式选择:`--delivery-profile friendly_tutorial` - 保留档:`commercial_narration`,仅用于公告、品牌声明等确需正式表达的内容 - 克隆音色选择顺序:`--voice` > `MINIMAX_VOICE_ID` > 已激活的本机克隆档 > 仓库系统音色 - 本机档:`~/.config/duanku/minimax-voice.json`,必须为 `0600`,不提交仓库
### MiniMax 专属语境适配(兼容回退)
`minimax_tts.context_adaptation` 只在 MiniMax 分支生效。Edge 和 Kokoro 不读取此配置,也不会改变原请求参数。
- 仅当 `delivery_consistency` 关闭时,根据文本自动识别 `opening`、`summary`、`explanation`、`instruction`、`warning`、`resource`、`conclusion`、`call_to_action`、`neutral` - 语境档只轻量调整 MiniMax 的语速、音量和音高,不修改原文、不注入 emotion、不自动插入声音标签 - 未命中规则时使用 `explanation` - 可用 `--context warning` 等参数显式覆盖自动识别;选择 Edge/Kokoro 时该参数被忽略 - 调用方无需理解语境规则。视频流水线只需继续提交文本并使用返回音频,音频时长仍由下游实测
```bash # 自动识别语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt
# MiniMax 显式指定风险提醒语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --context warning ```
### MiniMax 声音克隆
声音克隆分为不收费的创建阶段和首次 TTS 激活阶段。首次使用新克隆音色合成会产生官方克隆费及试听字符费,脚本要求 quote ID 和精确金额确认,不能用布尔参数绕过。
**引擎边界**:克隆 profile、克隆 `voice_id`、`--delivery-profile` 和克隆样本门禁只允许在实际 `tts_engine=minimax` 时读取。切换到 Edge 或 Kokoro 后必须完全跳过这些配置,即使本机仍保留已激活的 MiniMax profile,也不得影响非 MiniMax 的音色、请求参数或预检结果。
```bash # 样本检查,不联网、不收费 python3 scripts/minimax_voice_clone.py inspect --sample /path/to/source.m4a
# 输出激活报价,不联网、不收费 python3 scripts/minimax_voice_clone.py quote \ --voice-id DuankuNarrator20260801 \ --text "试听文案"
# 上传并创建克隆,不执行 TTS;必须确认拥有声音授权 python3 scripts/minimax_voice_clone.py clone \ --sample /path/to/source.m4a \ --voice-id DuankuNarrator20260801 \ --rights-confirmed
# 首次付费激活并生成试听,必须使用上一步 quote 的 ID 和金额 python3 scripts/minimax_voice_clone.py activate \ --text "试听文案" \ --output /path/to/preview.mp3 \ --confirm-quote-id '<quote_id>' \ --confirm-amount-usd '<estimated_total_usd>' ```
样本要求:`mp3/m4a/wav`、10 秒至 5 分钟、最大 20 MB。克隆脚本默认请求 MiniMax 降噪和音量归一化。
密钥规则:
- API Key 只从环境变量读取,默认变量名 `MINIMAX_API_KEY` - 克隆 `voice_id`、远程 file ID 和激活记录只保存在本机 `0600` profile - 不要把真实 Key 写进 `config/tts_config.json`、README、命令行参数或提交历史 - 如果要换变量名,只改 `config/tts_config.json` 中的 `minimax_tts.api_key_env`
```bash export MINIMAX_API_KEY="你的本机 key" python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3 ```
## Kokoro TTS v1.1-zh 声音
使用 `--list-voices` 查看完整列表(102 个)。
### 推荐声音 - `zm_009` - 男声(默认) - `zf_094` - 女声(自然温柔) - `zf_001` - 女声 - `zm_050` - 男声
### 英文声音 - `af_maple` - 女声(Maple) - `af_sol` - 女声(Sol) - `bf_vale` - 男声(Vale)
### 声音命名规则 - `zf_XXX` - 中文女声(55 个) - `zm_XXX` - 中文男声(44 个) - `af_`/`bf_` - 英文声音(3 个)
## 启动 Kokoro 服务
Kokoro TTS 需要 Docker 容器运行:
```bash # 启动 cd /Users/m/document/QNSZ/project/kokoro-tts && ./start.sh
# 停止 cd /Users/m/document/QNSZ/project/kokoro-tts && ./stop.sh
# Web UI 试听 # http://localhost:8880/web/ ```
## 核心功能
### 1. 脚本解析 自动识别并移除播客脚本中的注释和标记: - 时间戳:`(00:00)` - BGM 注释:`[BGM渐入:...]` - 舞台指示:`(主播声音:...)` `(停顿 1秒)` - Markdown 标记:`**文本**`
### 2. 中英文混合朗读 v1.1-zh 模型支持中英文混合文本的自然朗读。
### 3. 后处理集成 可选集成 voice-changer skill 进行变声处理。
## 配置文件
配置文件位于:`~/.claude/skills/text-to-speech/config/tts_config.json`
关键配置项: - `tts_engine`: `"minimax"`、`"kokoro"` 或 `"edge"`(默认引擎) - `minimax_tts`: MiniMax 引擎配置(API URL、模型、默认音色、语速;Key 仅从环境变量读取) - `minimax_tts.context_adaptation`: MiniMax 专属语境档和自动识别规则;不影响其他引擎 - `kokoro_tts`: Kokoro 引擎配置(API URL、默认声音、语速) - `edge_tts`: Edge 引擎配置(声音、语速、音调、音量) - `available_voices`: 按引擎分组的可用声音列表
## 工作流程
``` 输入文本/文件 ↓ 脚本解析(移除注释和标记) ↓ MiniMax / Kokoro TTS / Edge TTS 语音合成 ↓ 后处理(voice-changer,可选) ↓ 输出 MP3 文件 ```
## 代理绕过(重要)
Kokoro TTS 运行在 `localhost:8880`。如果系统配置了 HTTP 代理(`http_proxy`/`https_proxy`),请求 localhost 会被代理拦截导致连接失败(curl 返回 HTTP 000)。
**规则**: - Python 脚本已内置 `os.environ.setdefault("no_proxy", "localhost,127.0.0.1")`,通过脚本调用无需额外处理 - 如果 AI 需要直接用 `curl` 测试或调用 Kokoro API,**必须**加 `--noproxy localhost,127.0.0.1` 或设置 `no_proxy=localhost,127.0.0.1` - 如果 AI 直接写 Python `requests.post("http://localhost:8880/...")`,必须设置 `proxies={"http": None, "https": None}`,或使用 `requests.Session(); session.trust_env = False`,并设置 `NO_PROXY/no_proxy=localhost,127.0.0.1,::1` - 禁止不加代理绕过直接 curl/requests localhost;不要先走代理失败再重试,localhost Kokoro 请求默认就必须绕过代理
```bash # 正确:绕过代理 curl --noproxy localhost,127.0.0.1 -X POST http://localhost:8880/v1/audio/speech ...
# 错误:走了代理,返回 HTTP 000 curl -X POST http://localhost:8880/v1/audio/speech ... ```
## 依赖
- MiniMax TTS: `MINIMAX_API_KEY` 环境变量 - Kokoro TTS: Docker(容器运行在 localhost:8880) - Edge TTS: `pip install edge-tts`
## 性能参考
- Kokoro TTS: 1000字约 3-5 秒(本地 Docker CPU) - Edge TTS: 1000字约 10-20 秒(受网络影响)
Source provenance
Decision snapshot
612 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for text-to-speech, ready for a manual X post.
For a repeatable workflow, this is a skill worth shortlisting before another blank prompt. text-to-speech: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) 612 stars https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x
Listing + install path for text-to-speech: https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x Install: npx skills add wlzh/skills --skill text-to-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to M. but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech/audit)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)M.
@m.
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsInstall targets
Codex install prompt
Install the "text-to-speech" agent skill from https://github.com/wlzh/skills/tree/main/text-to-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wlzh-text-to-speech","task":"Install text-to-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wlzh/skills --skill text-to-speech
Maintenance
fresh
9d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
612
75/100 Quality · 66/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
612 GitHub stars
Repo activity
612 stars, 75 forks
Maintenance
9d since push
License
MIT
Install
npx skills add wlzh/skills --skill text-to-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wlzh/skills --skill text-to-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wlzh-text-to-speech/install
Agent should check
Copy prompt
Task: Use text-to-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install
Install command: npx skills add wlzh/skills --skill text-to-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wlzh-text-to-speech/install
LLM text format
/api/skills/wlzh-text-to-speech/install?format=text
Find alternatives
/api/skills/search?q=text-to-speech&limit=3
Agent prompt
Use text-to-speech for this task. Review https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install, then install with: npx skills add wlzh/skills --skill text-to-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wlzh-text-to-speech
LLM text
/api/registry/manifest/wlzh-text-to-speech?format=text
Install alias
/api/registry/install/wlzh-text-to-speech
Recommend
/api/registry/recommend?task=Use%20text-to-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
GitHub automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
GitHub automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO612 GitHub stars
Stars/forks activity
INFO612 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS9d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Manage repositories
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: text-to-speech description: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) version: 3.6.0 changelog: - 2026-08-01: v3.6.0 新增并默认启用 friendly_tutorial 整期表达档,面向普通大众的免费教程保持正常语速、自然口语和固定声线 - 2026-08-01: v3.5.0 新增 MiniMax 词级时间戳 sidecar;只保存校验后的时间段,不保存短期签名下载 URL;Edge/Kokoro 显式拒绝该参数 - 2026-08-01: v3.4.0 新增 MiniMax 声音克隆上传/创建/激活费用门禁、本机 0600 克隆音色档,以及整期固定 commercial_narration 表达档;克隆音色或表达档变化会使下游缓存失效 - 2026-07-18: v3.3.1 MiniMax 语境适配移除逐 beat emotion 注入,避免同一场景语气跳变;保留 speed/volume/pitch 轻量调整并同步测试文档 - 2026-07-18: v3.3.0 MiniMax 默认音色改为 Chinese (Mandarin)_Reliable_Executive;新增 MiniMax 专属语境适配层,按开场、总结、解释、步骤、提醒、资源、结论和关注引导自动微调表达;Edge/Kokoro 行为不变 - 2026-07-17: v3.2.0 新增 MiniMax TTS 引擎并设为默认,默认音色 male-qn-jingying(精英青年)、语速 1.0;API Key 只读取 MINIMAX_API_KEY 环境变量;保留 Kokoro/Edge 可配置切换 - 2026-05-17: v3.1.0 强化 localhost Kokoro 代理绕过规则——curl/requests 直连本地服务默认必须 NO_PROXY,不允许先走代理失败后重试
author: M. ---
# Text-to-Speech Skill
将文本转换为语音。默认使用 MiniMax TTS(在线高质量中文配音),并保留 Kokoro TTS v1.1-zh(本地 Docker,102 个中文音色)和 Edge TTS(在线)作为可配置后备。
## 引擎对比
| 特性 | MiniMax TTS | Kokoro TTS v1.1-zh | Edge TTS | |------|-------------|-------------------|----------| | 质量 | 默认推荐,中文短视频旁白更自然 | 本地可用、接近真人 | 标准 Neural 语音 | | 网络 | 需要 MiniMax API | 不需要(本地 Docker) | 需要网络连接 | | 默认音色 | `Chinese (Mandarin)_Reliable_Executive`(可靠高管) | `zm_009` | `zh-CN-YunyangNeural` | | 语速调节 | speed 参数,默认 1.0 | speed 参数 | rate/pitch/volume | | 前提 | `MINIMAX_API_KEY` 环境变量 | Docker 容器需运行 | 安装 `edge-tts` | | 配置值 | `minimax` | `kokoro` | `edge` |
## 使用说明
```bash # 默认使用 MiniMax TTS(当前配置) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py <文本文件>
# 指定引擎 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine minimax python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine kokoro python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine edge
# 指定声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -v zf_094
# 指定输出文件 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3
# MiniMax 专属:同时保存经过校验的词级时间戳 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt \ -o output.mp3 --subtitle-output output.subtitles.json
# 调整语速(MiniMax/Kokoro) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --speed 1.2
# 列出所有可用声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py --list-voices ```
## MiniMax TTS
默认配置:
- 引擎:`tts_engine = "minimax"` - 模型:`speech-2.8-hd` - 音色:`Chinese (Mandarin)_Reliable_Executive`(可靠高管) - 语速:`1.0` - 输出:MP3
### MiniMax 词级时间戳
`--subtitle-output <path>` 会为 MiniMax 请求启用词级时间戳,下载官方 sidecar 后校验文本、字符范围和时间单调性,再写入本地 JSON。sidecar 只保留 `provider/model/voice/subtitle_type/segments`,不会记录官方返回的短期签名 URL。该参数属于 MiniMax 专属能力;Edge/Kokoro 会失败关闭,避免下游误把估算时间当成官方时间戳。
### 整期表达一致性(默认)
`minimax_tts.delivery_consistency` 默认启用 `friendly_tutorial`。它用于面向普通大众的免费教程:像真人耐心讲解,允许正文自然使用“好、其实、这里呢、别急”等连接词,但不靠逐句变调制造表演感。同一期视频内所有句子固定使用相同的音色、速度、音量和音调,避免逐 beat 割裂。
- 默认档:`friendly_tutorial`,speed `1.0`、volume `1.0`、pitch `0` - 显式选择:`--delivery-profile friendly_tutorial` - 保留档:`commercial_narration`,仅用于公告、品牌声明等确需正式表达的内容 - 克隆音色选择顺序:`--voice` > `MINIMAX_VOICE_ID` > 已激活的本机克隆档 > 仓库系统音色 - 本机档:`~/.config/duanku/minimax-voice.json`,必须为 `0600`,不提交仓库
### MiniMax 专属语境适配(兼容回退)
`minimax_tts.context_adaptation` 只在 MiniMax 分支生效。Edge 和 Kokoro 不读取此配置,也不会改变原请求参数。
- 仅当 `delivery_consistency` 关闭时,根据文本自动识别 `opening`、`summary`、`explanation`、`instruction`、`warning`、`resource`、`conclusion`、`call_to_action`、`neutral` - 语境档只轻量调整 MiniMax 的语速、音量和音高,不修改原文、不注入 emotion、不自动插入声音标签 - 未命中规则时使用 `explanation` - 可用 `--context warning` 等参数显式覆盖自动识别;选择 Edge/Kokoro 时该参数被忽略 - 调用方无需理解语境规则。视频流水线只需继续提交文本并使用返回音频,音频时长仍由下游实测
```bash # 自动识别语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt
# MiniMax 显式指定风险提醒语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --context warning ```
### MiniMax 声音克隆
声音克隆分为不收费的创建阶段和首次 TTS 激活阶段。首次使用新克隆音色合成会产生官方克隆费及试听字符费,脚本要求 quote ID 和精确金额确认,不能用布尔参数绕过。
**引擎边界**:克隆 profile、克隆 `voice_id`、`--delivery-profile` 和克隆样本门禁只允许在实际 `tts_engine=minimax` 时读取。切换到 Edge 或 Kokoro 后必须完全跳过这些配置,即使本机仍保留已激活的 MiniMax profile,也不得影响非 MiniMax 的音色、请求参数或预检结果。
```bash # 样本检查,不联网、不收费 python3 scripts/minimax_voice_clone.py inspect --sample /path/to/source.m4a
# 输出激活报价,不联网、不收费 python3 scripts/minimax_voice_clone.py quote \ --voice-id DuankuNarrator20260801 \ --text "试听文案"
# 上传并创建克隆,不执行 TTS;必须确认拥有声音授权 python3 scripts/minimax_voice_clone.py clone \ --sample /path/to/source.m4a \ --voice-id DuankuNarrator20260801 \ --rights-confirmed
# 首次付费激活并生成试听,必须使用上一步 quote 的 ID 和金额 python3 scripts/minimax_voice_clone.py activate \ --text "试听文案" \ --output /path/to/preview.mp3 \ --confirm-quote-id '<quote_id>' \ --confirm-amount-usd '<estimated_total_usd>' ```
样本要求:`mp3/m4a/wav`、10 秒至 5 分钟、最大 20 MB。克隆脚本默认请求 MiniMax 降噪和音量归一化。
密钥规则:
- API Key 只从环境变量读取,默认变量名 `MINIMAX_API_KEY` - 克隆 `voice_id`、远程 file ID 和激活记录只保存在本机 `0600` profile - 不要把真实 Key 写进 `config/tts_config.json`、README、命令行参数或提交历史 - 如果要换变量名,只改 `config/tts_config.json` 中的 `minimax_tts.api_key_env`
```bash export MINIMAX_API_KEY="你的本机 key" python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3 ```
## Kokoro TTS v1.1-zh 声音
使用 `--list-voices` 查看完整列表(102 个)。
### 推荐声音 - `zm_009` - 男声(默认) - `zf_094` - 女声(自然温柔) - `zf_001` - 女声 - `zm_050` - 男声
### 英文声音 - `af_maple` - 女声(Maple) - `af_sol` - 女声(Sol) - `bf_vale` - 男声(Vale)
### 声音命名规则 - `zf_XXX` - 中文女声(55 个) - `zm_XXX` - 中文男声(44 个) - `af_`/`bf_` - 英文声音(3 个)
## 启动 Kokoro 服务
Kokoro TTS 需要 Docker 容器运行:
```bash # 启动 cd /Users/m/document/QNSZ/project/kokoro-tts && ./start.sh
# 停止 cd /Users/m/document/QNSZ/project/kokoro-tts && ./stop.sh
# Web UI 试听 # http://localhost:8880/web/ ```
## 核心功能
### 1. 脚本解析 自动识别并移除播客脚本中的注释和标记: - 时间戳:`(00:00)` - BGM 注释:`[BGM渐入:...]` - 舞台指示:`(主播声音:...)` `(停顿 1秒)` - Markdown 标记:`**文本**`
### 2. 中英文混合朗读 v1.1-zh 模型支持中英文混合文本的自然朗读。
### 3. 后处理集成 可选集成 voice-changer skill 进行变声处理。
## 配置文件
配置文件位于:`~/.claude/skills/text-to-speech/config/tts_config.json`
关键配置项: - `tts_engine`: `"minimax"`、`"kokoro"` 或 `"edge"`(默认引擎) - `minimax_tts`: MiniMax 引擎配置(API URL、模型、默认音色、语速;Key 仅从环境变量读取) - `minimax_tts.context_adaptation`: MiniMax 专属语境档和自动识别规则;不影响其他引擎 - `kokoro_tts`: Kokoro 引擎配置(API URL、默认声音、语速) - `edge_tts`: Edge 引擎配置(声音、语速、音调、音量) - `available_voices`: 按引擎分组的可用声音列表
## 工作流程
``` 输入文本/文件 ↓ 脚本解析(移除注释和标记) ↓ MiniMax / Kokoro TTS / Edge TTS 语音合成 ↓ 后处理(voice-changer,可选) ↓ 输出 MP3 文件 ```
## 代理绕过(重要)
Kokoro TTS 运行在 `localhost:8880`。如果系统配置了 HTTP 代理(`http_proxy`/`https_proxy`),请求 localhost 会被代理拦截导致连接失败(curl 返回 HTTP 000)。
**规则**: - Python 脚本已内置 `os.environ.setdefault("no_proxy", "localhost,127.0.0.1")`,通过脚本调用无需额外处理 - 如果 AI 需要直接用 `curl` 测试或调用 Kokoro API,**必须**加 `--noproxy localhost,127.0.0.1` 或设置 `no_proxy=localhost,127.0.0.1` - 如果 AI 直接写 Python `requests.post("http://localhost:8880/...")`,必须设置 `proxies={"http": None, "https": None}`,或使用 `requests.Session(); session.trust_env = False`,并设置 `NO_PROXY/no_proxy=localhost,127.0.0.1,::1` - 禁止不加代理绕过直接 curl/requests localhost;不要先走代理失败再重试,localhost Kokoro 请求默认就必须绕过代理
```bash # 正确:绕过代理 curl --noproxy localhost,127.0.0.1 -X POST http://localhost:8880/v1/audio/speech ...
# 错误:走了代理,返回 HTTP 000 curl -X POST http://localhost:8880/v1/audio/speech ... ```
## 依赖
- MiniMax TTS: `MINIMAX_API_KEY` 环境变量 - Kokoro TTS: Docker(容器运行在 localhost:8880) - Edge TTS: `pip install edge-tts`
## 性能参考
- Kokoro TTS: 1000字约 3-5 秒(本地 Docker CPU) - Edge TTS: 1000字约 10-20 秒(受网络影响)
Source provenance
Decision snapshot
612 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for text-to-speech, ready for a manual X post.
For a repeatable workflow, this is a skill worth shortlisting before another blank prompt. text-to-speech: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) 612 stars https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x
Listing + install path for text-to-speech: https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x Install: npx skills add wlzh/skills --skill text-to-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to M. but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech/audit)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)M.
@m.
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsInstall targets
Codex install prompt
Install the "text-to-speech" agent skill from https://github.com/wlzh/skills/tree/main/text-to-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wlzh-text-to-speech","task":"Install text-to-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wlzh/skills --skill text-to-speech
Maintenance
fresh
9d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
612
75/100 Quality · 66/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
612 GitHub stars
Repo activity
612 stars, 75 forks
Maintenance
9d since push
License
MIT
Install
npx skills add wlzh/skills --skill text-to-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wlzh/skills --skill text-to-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wlzh-text-to-speech/install
Agent should check
Copy prompt
Task: Use text-to-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20text-to-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install
Install command: npx skills add wlzh/skills --skill text-to-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wlzh-text-to-speech/install
LLM text format
/api/skills/wlzh-text-to-speech/install?format=text
Find alternatives
/api/skills/search?q=text-to-speech&limit=3
Agent prompt
Use text-to-speech for this task. Review https://www.openagentskill.com/api/skills/wlzh-text-to-speech/install, then install with: npx skills add wlzh/skills --skill text-to-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wlzh-text-to-speech
LLM text
/api/registry/manifest/wlzh-text-to-speech?format=text
Install alias
/api/registry/install/wlzh-text-to-speech
Recommend
/api/registry/recommend?task=Use%20text-to-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
GitHub automation
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
GitHub automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO612 GitHub stars
Stars/forks activity
INFO612 stars, 75 forks; issue activity unavailable in current metadata
Recent maintenance
PASS9d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Manage repositories
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: text-to-speech description: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) version: 3.6.0 changelog: - 2026-08-01: v3.6.0 新增并默认启用 friendly_tutorial 整期表达档,面向普通大众的免费教程保持正常语速、自然口语和固定声线 - 2026-08-01: v3.5.0 新增 MiniMax 词级时间戳 sidecar;只保存校验后的时间段,不保存短期签名下载 URL;Edge/Kokoro 显式拒绝该参数 - 2026-08-01: v3.4.0 新增 MiniMax 声音克隆上传/创建/激活费用门禁、本机 0600 克隆音色档,以及整期固定 commercial_narration 表达档;克隆音色或表达档变化会使下游缓存失效 - 2026-07-18: v3.3.1 MiniMax 语境适配移除逐 beat emotion 注入,避免同一场景语气跳变;保留 speed/volume/pitch 轻量调整并同步测试文档 - 2026-07-18: v3.3.0 MiniMax 默认音色改为 Chinese (Mandarin)_Reliable_Executive;新增 MiniMax 专属语境适配层,按开场、总结、解释、步骤、提醒、资源、结论和关注引导自动微调表达;Edge/Kokoro 行为不变 - 2026-07-17: v3.2.0 新增 MiniMax TTS 引擎并设为默认,默认音色 male-qn-jingying(精英青年)、语速 1.0;API Key 只读取 MINIMAX_API_KEY 环境变量;保留 Kokoro/Edge 可配置切换 - 2026-05-17: v3.1.0 强化 localhost Kokoro 代理绕过规则——curl/requests 直连本地服务默认必须 NO_PROXY,不允许先走代理失败后重试
author: M. ---
# Text-to-Speech Skill
将文本转换为语音。默认使用 MiniMax TTS(在线高质量中文配音),并保留 Kokoro TTS v1.1-zh(本地 Docker,102 个中文音色)和 Edge TTS(在线)作为可配置后备。
## 引擎对比
| 特性 | MiniMax TTS | Kokoro TTS v1.1-zh | Edge TTS | |------|-------------|-------------------|----------| | 质量 | 默认推荐,中文短视频旁白更自然 | 本地可用、接近真人 | 标准 Neural 语音 | | 网络 | 需要 MiniMax API | 不需要(本地 Docker) | 需要网络连接 | | 默认音色 | `Chinese (Mandarin)_Reliable_Executive`(可靠高管) | `zm_009` | `zh-CN-YunyangNeural` | | 语速调节 | speed 参数,默认 1.0 | speed 参数 | rate/pitch/volume | | 前提 | `MINIMAX_API_KEY` 环境变量 | Docker 容器需运行 | 安装 `edge-tts` | | 配置值 | `minimax` | `kokoro` | `edge` |
## 使用说明
```bash # 默认使用 MiniMax TTS(当前配置) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py <文本文件>
# 指定引擎 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine minimax python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine kokoro python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --engine edge
# 指定声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -v zf_094
# 指定输出文件 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3
# MiniMax 专属:同时保存经过校验的词级时间戳 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt \ -o output.mp3 --subtitle-output output.subtitles.json
# 调整语速(MiniMax/Kokoro) python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --speed 1.2
# 列出所有可用声音 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py --list-voices ```
## MiniMax TTS
默认配置:
- 引擎:`tts_engine = "minimax"` - 模型:`speech-2.8-hd` - 音色:`Chinese (Mandarin)_Reliable_Executive`(可靠高管) - 语速:`1.0` - 输出:MP3
### MiniMax 词级时间戳
`--subtitle-output <path>` 会为 MiniMax 请求启用词级时间戳,下载官方 sidecar 后校验文本、字符范围和时间单调性,再写入本地 JSON。sidecar 只保留 `provider/model/voice/subtitle_type/segments`,不会记录官方返回的短期签名 URL。该参数属于 MiniMax 专属能力;Edge/Kokoro 会失败关闭,避免下游误把估算时间当成官方时间戳。
### 整期表达一致性(默认)
`minimax_tts.delivery_consistency` 默认启用 `friendly_tutorial`。它用于面向普通大众的免费教程:像真人耐心讲解,允许正文自然使用“好、其实、这里呢、别急”等连接词,但不靠逐句变调制造表演感。同一期视频内所有句子固定使用相同的音色、速度、音量和音调,避免逐 beat 割裂。
- 默认档:`friendly_tutorial`,speed `1.0`、volume `1.0`、pitch `0` - 显式选择:`--delivery-profile friendly_tutorial` - 保留档:`commercial_narration`,仅用于公告、品牌声明等确需正式表达的内容 - 克隆音色选择顺序:`--voice` > `MINIMAX_VOICE_ID` > 已激活的本机克隆档 > 仓库系统音色 - 本机档:`~/.config/duanku/minimax-voice.json`,必须为 `0600`,不提交仓库
### MiniMax 专属语境适配(兼容回退)
`minimax_tts.context_adaptation` 只在 MiniMax 分支生效。Edge 和 Kokoro 不读取此配置,也不会改变原请求参数。
- 仅当 `delivery_consistency` 关闭时,根据文本自动识别 `opening`、`summary`、`explanation`、`instruction`、`warning`、`resource`、`conclusion`、`call_to_action`、`neutral` - 语境档只轻量调整 MiniMax 的语速、音量和音高,不修改原文、不注入 emotion、不自动插入声音标签 - 未命中规则时使用 `explanation` - 可用 `--context warning` 等参数显式覆盖自动识别;选择 Edge/Kokoro 时该参数被忽略 - 调用方无需理解语境规则。视频流水线只需继续提交文本并使用返回音频,音频时长仍由下游实测
```bash # 自动识别语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt
# MiniMax 显式指定风险提醒语境 python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt --context warning ```
### MiniMax 声音克隆
声音克隆分为不收费的创建阶段和首次 TTS 激活阶段。首次使用新克隆音色合成会产生官方克隆费及试听字符费,脚本要求 quote ID 和精确金额确认,不能用布尔参数绕过。
**引擎边界**:克隆 profile、克隆 `voice_id`、`--delivery-profile` 和克隆样本门禁只允许在实际 `tts_engine=minimax` 时读取。切换到 Edge 或 Kokoro 后必须完全跳过这些配置,即使本机仍保留已激活的 MiniMax profile,也不得影响非 MiniMax 的音色、请求参数或预检结果。
```bash # 样本检查,不联网、不收费 python3 scripts/minimax_voice_clone.py inspect --sample /path/to/source.m4a
# 输出激活报价,不联网、不收费 python3 scripts/minimax_voice_clone.py quote \ --voice-id DuankuNarrator20260801 \ --text "试听文案"
# 上传并创建克隆,不执行 TTS;必须确认拥有声音授权 python3 scripts/minimax_voice_clone.py clone \ --sample /path/to/source.m4a \ --voice-id DuankuNarrator20260801 \ --rights-confirmed
# 首次付费激活并生成试听,必须使用上一步 quote 的 ID 和金额 python3 scripts/minimax_voice_clone.py activate \ --text "试听文案" \ --output /path/to/preview.mp3 \ --confirm-quote-id '<quote_id>' \ --confirm-amount-usd '<estimated_total_usd>' ```
样本要求:`mp3/m4a/wav`、10 秒至 5 分钟、最大 20 MB。克隆脚本默认请求 MiniMax 降噪和音量归一化。
密钥规则:
- API Key 只从环境变量读取,默认变量名 `MINIMAX_API_KEY` - 克隆 `voice_id`、远程 file ID 和激活记录只保存在本机 `0600` profile - 不要把真实 Key 写进 `config/tts_config.json`、README、命令行参数或提交历史 - 如果要换变量名,只改 `config/tts_config.json` 中的 `minimax_tts.api_key_env`
```bash export MINIMAX_API_KEY="你的本机 key" python3 ~/.claude/skills/text-to-speech/scripts/text_to_speech.py script.txt -o output.mp3 ```
## Kokoro TTS v1.1-zh 声音
使用 `--list-voices` 查看完整列表(102 个)。
### 推荐声音 - `zm_009` - 男声(默认) - `zf_094` - 女声(自然温柔) - `zf_001` - 女声 - `zm_050` - 男声
### 英文声音 - `af_maple` - 女声(Maple) - `af_sol` - 女声(Sol) - `bf_vale` - 男声(Vale)
### 声音命名规则 - `zf_XXX` - 中文女声(55 个) - `zm_XXX` - 中文男声(44 个) - `af_`/`bf_` - 英文声音(3 个)
## 启动 Kokoro 服务
Kokoro TTS 需要 Docker 容器运行:
```bash # 启动 cd /Users/m/document/QNSZ/project/kokoro-tts && ./start.sh
# 停止 cd /Users/m/document/QNSZ/project/kokoro-tts && ./stop.sh
# Web UI 试听 # http://localhost:8880/web/ ```
## 核心功能
### 1. 脚本解析 自动识别并移除播客脚本中的注释和标记: - 时间戳:`(00:00)` - BGM 注释:`[BGM渐入:...]` - 舞台指示:`(主播声音:...)` `(停顿 1秒)` - Markdown 标记:`**文本**`
### 2. 中英文混合朗读 v1.1-zh 模型支持中英文混合文本的自然朗读。
### 3. 后处理集成 可选集成 voice-changer skill 进行变声处理。
## 配置文件
配置文件位于:`~/.claude/skills/text-to-speech/config/tts_config.json`
关键配置项: - `tts_engine`: `"minimax"`、`"kokoro"` 或 `"edge"`(默认引擎) - `minimax_tts`: MiniMax 引擎配置(API URL、模型、默认音色、语速;Key 仅从环境变量读取) - `minimax_tts.context_adaptation`: MiniMax 专属语境档和自动识别规则;不影响其他引擎 - `kokoro_tts`: Kokoro 引擎配置(API URL、默认声音、语速) - `edge_tts`: Edge 引擎配置(声音、语速、音调、音量) - `available_voices`: 按引擎分组的可用声音列表
## 工作流程
``` 输入文本/文件 ↓ 脚本解析(移除注释和标记) ↓ MiniMax / Kokoro TTS / Edge TTS 语音合成 ↓ 后处理(voice-changer,可选) ↓ 输出 MP3 文件 ```
## 代理绕过(重要)
Kokoro TTS 运行在 `localhost:8880`。如果系统配置了 HTTP 代理(`http_proxy`/`https_proxy`),请求 localhost 会被代理拦截导致连接失败(curl 返回 HTTP 000)。
**规则**: - Python 脚本已内置 `os.environ.setdefault("no_proxy", "localhost,127.0.0.1")`,通过脚本调用无需额外处理 - 如果 AI 需要直接用 `curl` 测试或调用 Kokoro API,**必须**加 `--noproxy localhost,127.0.0.1` 或设置 `no_proxy=localhost,127.0.0.1` - 如果 AI 直接写 Python `requests.post("http://localhost:8880/...")`,必须设置 `proxies={"http": None, "https": None}`,或使用 `requests.Session(); session.trust_env = False`,并设置 `NO_PROXY/no_proxy=localhost,127.0.0.1,::1` - 禁止不加代理绕过直接 curl/requests localhost;不要先走代理失败再重试,localhost Kokoro 请求默认就必须绕过代理
```bash # 正确:绕过代理 curl --noproxy localhost,127.0.0.1 -X POST http://localhost:8880/v1/audio/speech ...
# 错误:走了代理,返回 HTTP 000 curl -X POST http://localhost:8880/v1/audio/speech ... ```
## 依赖
- MiniMax TTS: `MINIMAX_API_KEY` 环境变量 - Kokoro TTS: Docker(容器运行在 localhost:8880) - Edge TTS: `pip install edge-tts`
## 性能参考
- Kokoro TTS: 1000字约 3-5 秒(本地 Docker CPU) - Edge TTS: 1000字约 10-20 秒(受网络影响)
Source provenance
Decision snapshot
612 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for text-to-speech, ready for a manual X post.
For a repeatable workflow, this is a skill worth shortlisting before another blank prompt. text-to-speech: 文本转语音工具 - 默认 MiniMax TTS,支持切换 Edge TTS 和 Kokoro TTS (v1.1-zh) 612 stars https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x
Listing + install path for text-to-speech: https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=x Install: npx skills add wlzh/skills --skill text-to-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to M. but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech/audit)
[](https://www.openagentskill.com/skills/wlzh-text-to-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)M.
@m.
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsPermission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness