Creator · jianshuo
Last updated · Sep 6, 2026
Use when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub. Routes by voice ID — Volcano (豆包) TTS for Chinese, edge-tts neural for any language. Defaults to one voice (single-speaker); opt-in
Do not auto-install
Install targets
Codex install prompt
Install the "wjs-dubbing-video" agent skill from https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-video. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Use when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub. Routes by voice ID — Volcano (豆包) TTS for Chinese, edge-tts neural for any language. Defaults to one voice (single-speaker); opt-in multi-speaker via visual diarization. Outputs `*_<lang>_dub.mp4` with the dub audio in place of the original. Final mixing (audio bed + burn-in) is handed off to `/wjs-burning-subtitles`. Triggers — "配音", "中文配音", "Chinese dub", "voice over this", "dub the video", "TTS this SRT", "different voice for each speaker". After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"jianshuo-wjs-dubbing-video","task":"Install wjs-dubbing-video","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Scenario
Design and creative
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add jianshuo/claude-skills --skill wjs-dubbing-video
Maintenance
fresh
18d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
129
68/100 Quality · 66/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Do not auto-installTrust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Choose a stronger alternative or inspect the source manually before any install attempt.
Stars
129 GitHub stars
Repo activity
129 stars, 20 forks
Maintenance
18d since push
License
MIT
Install
npx skills add jianshuo/claude-skills --skill wjs-dubbing-video
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add jianshuo/claude-skills --skill wjs-dubbing-videoDo not use when
Alternative
175.1K Stars
npx skills add anthropics/skills --skill frontend-design
Alternative
85.2K Stars
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Alternative
1.8K Stars
npx skills add Alisa0808/vox-director --skill vox-director
Alternative
175.1K Stars
npx skills add anthropics/skills --skill canvas-design
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20wjs-dubbing-video%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20wjs-dubbing-video%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/jianshuo-wjs-dubbing-video/install
Agent should check
Copy prompt
Task: Use wjs-dubbing-video in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20wjs-dubbing-video%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/jianshuo-wjs-dubbing-video/install
Install command: npx skills add jianshuo/claude-skills --skill wjs-dubbing-video
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/jianshuo-wjs-dubbing-video/install
LLM text format
/api/skills/jianshuo-wjs-dubbing-video/install?format=text
Find alternatives
/api/skills/search?q=wjs-dubbing-video&limit=3
Agent prompt
Use wjs-dubbing-video for this task. Review https://www.openagentskill.com/api/skills/jianshuo-wjs-dubbing-video/install, then install with: npx skills add jianshuo/claude-skills --skill wjs-dubbing-videoRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/jianshuo-wjs-dubbing-video
LLM text
/api/registry/manifest/jianshuo-wjs-dubbing-video?format=text
Install alias
/api/registry/install/jianshuo-wjs-dubbing-video
Recommend
/api/registry/recommend?task=Use%20wjs-dubbing-video%20in%20an%20agent%20workflow&limit=3
Agent fit
Testing and QA
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Testing and QA
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.
GitHub adoption
INFO129 GitHub stars
Stars/forks activity
CHECK129 stars, 20 forks; issue activity unavailable in current metadata
Recent maintenance
PASS18d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Choose a stronger alternative or inspect the source manually before any install attempt.
Quality profile
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Verify behavior
I need my agent to test a web app, reproduce bugs, and verify fixes.
Build and ship code
I need a coding agent that can understand a repository, edit code, and review pull requests.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Workflow fit
Design, build, test, and ship interfaces
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
--- name: wjs-dubbing-video description: Use when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub. Routes by voice ID — Volcano (豆包) TTS for Chinese, edge-tts neural for any language. Defaults to one voice (single-speaker); opt-in multi-speaker via visual diarization. Outputs `*_<lang>_dub.mp4` with the dub audio in place of the original. Final mixing (audio bed + burn-in) is handed off to `/wjs-burning-subtitles`. Triggers — "配音", "中文配音", "Chinese dub", "voice over this", "dub the video", "TTS this SRT", "different voice for each speaker". ---
# wjs-dubbing-video
Video + target-language SRT → `*_<lang>_dub.mp4` with a time-aligned TTS voice. **This skill stops at the dub track.** Burn-in + audio bed mixing is the next skill (`/wjs-burning-subtitles/render.py` composites everything in one final encode).
## When to use
- User has a target-language SRT (e.g., `entrevista.zh-CN.srt`) and wants the video to speak that language. - User says "中文配音 / 配音 / 帮我做配音 / dub it / voice over". - User has multiple speakers on camera and wants different voices per speaker.
## When NOT to use
- No SRT yet → run `/wjs-transcribing-audio` then `/wjs-translating-subtitles` first. - Source-language only TTS (rare; usually you translate first) → still use this skill, but pass the source SRT. - Burn-in only, no audio change → skip to `/wjs-burning-subtitles`.
## Number of speakers — default to one
**Default: assume one speaker.** Use a single voice for the entire dub. This is the right answer for monologues, vlogs, recorded talks, narrator-only clips, and the overwhelming majority of videos people ask about. Don't run diarization, don't tag the SRT with `[A]`/`[B]`, don't bring up multi-speaker complexity.
**Switch to multi-speaker only when the user explicitly says so** — phrasings like "two people", "interview", "dialogue", "conversation between", "separate the speakers", "different voice for each", or a direct request to do diarization. When triggered, follow the "Multi-speaker dubbing" section below.
If you're unsure whether a video is one speaker or many, ship the single-voice version first. Adding speaker separation later is cheap (just regenerate the dub); shipping confused multi-speaker output by default wastes the user's time.
## Engine routing — by voice ID
`scripts/dub.py` auto-routes by voice-ID prefix:
| Voice ID pattern | Engine | Auth | |---|---|---| | `zh_..._bigtts` | **Volcano (字节跳动豆包) TTS** | `VOLC_TTS_APPID` + `VOLC_TTS_ACCESS_TOKEN` | | `zh-CN-...Neural` / `en-US-...Neural` / etc. | **edge-tts** (Microsoft Edge neural) | none (free) |
For Mandarin, Volcano is markedly more natural than edge-tts, especially for emotional/contemplative content. Use edge-tts when Volcano credentials aren't available or as a debugging fallback.
## Volcano TTS (Chinese only)
Endpoint: `https://openspeech.bytedance.com/api/v3/tts/unidirectional` (used for both TTS 1.0 and 2.0; the Resource-Id header picks the backend).
Headers:
``` X-Api-App-Id: (env: VOLC_TTS_APPID) # 10-digit speech App ID X-Api-Access-Key: (env: VOLC_TTS_ACCESS_TOKEN) # 32-char token from speech console X-Api-Resource-Id: volc.service_type.10029 # see resource ID note below Content-Type: application/json ```
Loading credentials: most users keep them in `~/code/.env`. Read them at the top of any session via:
```bash set -a; source ~/code/.env; set +a ```
### Resource ID — important quirk
The doc lists `seed-tts-2.0` as the "TTS 2.0 (recommended)" resource, but a typical TTS-SeedTTS2.0 console instance does **not** include the popular `*_bigtts` speaker catalog (爽快斯斯, 高冷御姐, 开朗姐姐, etc.). Trying those speakers against `seed-tts-2.0` returns `200 code=55000000 "resource ID is mismatched with speaker related resource"`. The fix is to use `volc.service_type.10029` (the TTS 1.0 V3 endpoint) — the audio quality of the bigtts speakers is identical, and they all work against this resource. The bundled `dub.py` defaults to `volc.service_type.10029`; override with `VOLC_TTS_RESOURCE` env if you have a different instance.
Other 401/403 errors:
- `401 code=45000010 "load grant: requested grant not found in SaaS storage"` — the App ID + key combo is valid against the gateway, but the user has not activated this resource. They must go to 火山引擎 → 语音技术 → 语音合成大模型 → 实例管理 and 开通 the service. No workaround. - `403 code=45000030` — the speaker isn't included in the user's instance bundle.
### Response format
Despite the doc's casual language, the response is **streaming NDJSON**, not a single JSON object and not raw audio bytes. Each line is a separate JSON event with a base64-encoded MP3 chunk in `data`. The terminal event has `code: 20000000` (which means OK in this API's success codes — different from `code: 0`). Concatenate the decoded chunks for the full MP3.
```python import base64, json, requests audio = b"" r = requests.post(url, headers=h, json=payload, timeout=60, stream=True) for line in r.iter_lines(): if not line: continue evt = json.loads(line) if evt.get("code") not in (0, None, 20000000): raise RuntimeError(f"code={evt.get('code')} {evt.get('message')}") if evt.get("data"): audio += base64.b64decode(evt["data"]) ```
### Speaker catalog (verified working under `volc.service_type.10029`)
Full list at volcengine.com/docs/6561/1257544 — but availability depends on your instance bundle. Confirmed-working female voices for the typical SeedTTS-2.0 starter instance:
| Speaker ID | 中文名 | Feel | | --- | --- | --- | | `zh_female_gaolengyujie_moon_bigtts` | 高冷御姐 | **Best for contemplative/spiritual content.** Mature, restrained, calm. | | `zh_female_kailangjiejie_moon_bigtts` | 开朗姐姐 | Warm older-sister storytelling. | | `zh_female_shuangkuaisisi_moon_bigtts` | 爽快斯斯 | Versatile, conversational baseline. | | `zh_female_linjianvhai_moon_bigtts` | 邻家女孩 | Casual, lifestyle-vlog. | | `zh_female_yuanqinvyou_moon_bigtts` | 元气女友 | Lively, upbeat. | | `zh_female_meilinvyou_moon_bigtts` | 美丽女友 | Soft, intimate. | | `zh_female_shuangkuaisisi_emo_v2_mars_bigtts` | 斯斯情感版 | Full emotional range — pair with explicit emotion + scale. |
These voices return 55000000 against the typical instance even though the doc lists them: `vv_uranus_bigtts`, `wenroushunv_moon_bigtts`, `qingxin_moon_bigtts`, `yingmaoxiaoyuan_moon_bigtts`, `tianxinxiaoling_moon_bigtts`, `shaoergushi_moon_bigtts`. Don't promise them without testing.
### Audio params
`speech_rate` is Volcano's native scale [-50, +100] where the value is a percentage delta (so `-8` means 8% slower). The script passes `--rate -8%` through as `-8`.
Useful emotion presets:
- `emotion="calm"`, `emotion_scale=4` — contemplative, default for this skill's spiritual-content niche. - `emotion="gentle"` — softer / more intimate. - `emotion="neutral"` — flat / informational. - `emotion="sad"` — melancholic. Use sparingly.
Override `dub.py` defaults with `VOLC_TTS_EMOTION` and `VOLC_TTS_EMOTION_SCALE` env vars without editing code.
**No English Volcano voices** are wired up in this skill — for English use edge-tts (next section). Volcano does have English speakers (`en_male_*_bigtts`, `en_female_*_bigtts`) but they aren't typically included in TTS-SeedTTS-2.0 starter instances. Add them by extending the voice routing in `dub.py` once verified.
## edge-tts (Microsoft Edge neural TTS)
Free, no API key, high-quality but less expressive than Volcano. Install into a project venv — **do not** call it via `uvx` once per segment. Each `uvx` invocation spawns a fresh Python process and the bing endpoint will rate-limit or RST the connection after a handful of rapid hits, breaking mid-render.
```bash uv venv .venv uv pip install --python .venv/bin/python edge-tts ```
Then drive it from a single long-lived Python process using `edge_tts.Communicate(...)` directly, with retry-on-failure logic. The bundled `scripts/dub.py` does this.
## Voice selection — match the original speaker
There is no perfect cross-language match — choose gender, age feel, and tone deliberately, then bend with rate/pitch.
### Chinese voices (Volcano preferred, edge-tts fallback)
Volcano's `zh_female_gaolengyujie_moon_bigtts` (高冷御姐, calm, `speech_rate=-8`) is the validated baseline for mature contemplative female speakers — equivalent to or better than any edge-tts option for that profile. See the Volcano speaker table above for the rest.
edge-tts catalog (Chinese):
| Voice | Gender | Default feel | | --- | --- | --- | | `zh-CN-XiaoxiaoNeural` | F | Warm, news/novel | | `zh-CN-XiaoyiNeural` | F | Lively, young | | `zh-CN-YunjianNeural` | M | Passionate, sports | | `zh-CN-YunxiNeural` | M | Sunshine, lively | | `zh-CN-YunyangNeural` | M | Professional newsreader | | `zh-HK-HiuMaanNeural` | F | Friendly, slightly mature | | `zh-TW-HsiaoChenNeural` | F | Friendly |
### English voices (edge-tts neural, all multilingual)
All voices below speak fluent American/British/Australian English; the `*Multilingual*` ones also handle Spanish names, French/Italian loanwords, etc. without mispronunciation.
| Voice | Gender | Default feel | | --- | --- | --- | | `en-US-AvaMultilingualNeural` | F | **Best for warm/mature/caring** — natural for spiritual or coaching content | | `en-US-EmmaMultilingualNeural` | F | Cheerful, conversational, younger | | `en-US-AndrewMultilingualNeural` | M | Warm, confident, sincere | | `en-US-BrianMultilingualNeural` | M | Approachable, casual | | `en-US-AriaNeural` | F | Crisp newsreader | | `en-US-GuyNeural` | M | Steady male newsreader | | `en-GB-SoniaNeural` | F | British female (RP) | | `en-GB-RyanNeural` | M | British male (RP) | | `en-AU-WilliamMultilingualNeural` | M | Australian male | | `fr-FR-VivienneMultilingualNeural` | F | Mature European female who also reads English |
For matching a mature contemplative Spanish female (this skill's canonical use case), start with `en-US-AvaMultilingualNeural` at `--rate -5% --pitch -3Hz`. Do **not** use the news-style `Aria` or `Guy` for spiritual content — they sound clinical.
### Picking heuristics
- **Mature contemplative female speaker (yoga/spirituality/coaching):** `zh-CN-XiaoxiaoNeural` with `--rate=-8% --pitch=-10Hz` (or Volcano `gaolengyujie`). - **Mature professional male:** `zh-CN-YunyangNeural` with `--rate=-5%`. Avoid Yunjian/Yunxi (too energetic). - **Young casual speaker:** Defaults; no pitch shift. - **Western-mouth feel:** one of the `*MultilingualNeural` voices.
## Always sample before committing
🛑 **Checkpoint — sample before full dub.** A full-video dub is the most expensive step (TTS API calls + atempo + ffmpeg mux). Before running `dub.py` over the whole SRT:
1. Pick the longest-text cue (worst stretch case) and one short/casual cue (timbre check). 2. Synthesize 3–4 voice/rate/pitch combos at 3–8s each. 3. Show the user the audio panel and ask: "选哪个 voice?rate/pitch 要调吗?确认后我再跑全片。" Wait for explicit pick.
Skip the checkpoint
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for wjs-dubbing-video, ready for a manual X post.
A practical pick for design or creative work: wjs-dubbing-video: Use when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a tim... 129 stars https://www.openagentskill.com/skills/jianshuo-wjs-dubbing-video?ref=x
Listing + install path for wjs-dubbing-video: https://www.openagentskill.com/skills/jianshuo-wjs-dubbing-video?ref=x Install: npx skills add jianshuo/claude-skills --skill wjs-dubbing-video
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to jianshuo but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/jianshuo-wjs-dubbing-video?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jianshuo-wjs-dubbing-video?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/jianshuo-wjs-dubbing-video/audit)
[](https://www.openagentskill.com/skills/jianshuo-wjs-dubbing-video?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)jianshuo
@jianshuo
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Do not auto-install
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
175.1K StarsTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
85.2K StarsVox Director
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K StarsCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
175.1K StarsPermission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness