Creator · calesthio
Last updated · Sep 1, 2026
Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice clo
Creator · calesthio
Last updated · Sep 1, 2026
Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice clo
Creator · calesthio
Last updated · Sep 1, 2026
Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice clo
Creator · calesthio
Last updated · Sep 1, 2026
Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice clo
Sandbox only
Install targets
Codex install prompt
Install the "elevenlabs" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/elevenlabs. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"calesthio-elevenlabs","task":"Install elevenlabs","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Scenario
Design and creative
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add calesthio/OpenMontage --skill elevenlabs
Maintenance
fresh
16d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
55K
93/100 Quality · 80/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
55K GitHub stars
Repo activity
55K stars, 6.9K forks
Maintenance
16d since push
License
AGPL-3.0
Install
npx skills add calesthio/OpenMontage --skill elevenlabs
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add calesthio/OpenMontage --skill elevenlabsDo not use when
Alternative
174.9K Stars
npx skills add anthropics/skills --skill frontend-design
Alternative
84.9K Stars
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Alternative
1.8K Stars
npx skills add Alisa0808/vox-director --skill vox-director
Alternative
174.9K Stars
npx skills add anthropics/skills --skill canvas-design
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/calesthio-elevenlabs/install
Agent should check
Copy prompt
Task: Use elevenlabs in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install
Install command: npx skills add calesthio/OpenMontage --skill elevenlabs
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/calesthio-elevenlabs/install
LLM text format
/api/skills/calesthio-elevenlabs/install?format=text
Find alternatives
/api/skills/search?q=elevenlabs&limit=3
Agent prompt
Use elevenlabs for this task. Review https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install, then install with: npx skills add calesthio/OpenMontage --skill elevenlabsRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/calesthio-elevenlabs
LLM text
/api/registry/manifest/calesthio-elevenlabs?format=text
Install alias
/api/registry/install/calesthio-elevenlabs
Recommend
/api/registry/recommend?task=Use%20elevenlabs%20in%20an%20agent%20workflow&limit=3
Agent fit
Workflow automation
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Workflow automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS55K GitHub stars
Stars/forks activity
PASS55K stars, 6.9K forks; issue activity unavailable in current metadata
Recent maintenance
PASS16d since push
License clarity
PASSAGPL-3.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Design, build, test, and ship interfaces
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
--- name: elevenlabs description: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. ---
# ElevenLabs Audio Generation
## OpenMontage provider routing
Inspect the OpenMontage registry before choosing an authentication path.
- Prefer `fal_elevenlabs_tts` when it is available. It provides Eleven v3, Multilingual v2, and Turbo v2.5 through the centrally managed fal.ai connection; no separate ElevenLabs credential is needed. - Use `elevenlabs_tts` only when that direct provider is already reported as available by the registry. - In a shared installation, never tell the user to create a `.env`, export a key, or paste a credential. Report missing direct-provider access as an administrator setup request.
The direct API examples below require a centrally configured `ELEVENLABS_API_KEY`; they are not the default path when the fal.ai provider is available.
## Text-to-Speech
```python from elevenlabs.client import ElevenLabs from elevenlabs import save, VoiceSettings import os
client = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
audio = client.text_to_speech.convert( text="Welcome to my video!", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2", voice_settings=VoiceSettings( stability=0.5, similarity_boost=0.75, style=0.5, speed=1.0 ) ) save(audio, "voiceover.mp3") ```
### Models
| Model | Quality | SSML Support | Notes | |-------|---------|--------------|-------| | `eleven_multilingual_v2` | Highest consistency | None | Stable, production-ready, 29 languages | | `eleven_flash_v2_5` | Good | `<break>`, `<phoneme>` | Fast, supports pause/pronunciation tags | | `eleven_turbo_v2_5` | Good | `<break>`, `<phoneme>` | Fastest latency | | `eleven_v3` | Most expressive | None | Alpha — unreliable, needs prompt engineering |
**Choose:** multilingual_v2 for reliability, flash/turbo for SSML control, v3 for maximum expressiveness (expect retakes).
### Voice Settings by Style
| Style | stability | similarity | style | speed | |-------|-----------|------------|-------|-------| | Natural/professional | 0.75-0.85 | 0.9 | 0.0-0.1 | 1.0 | | Conversational | 0.5-0.6 | 0.85 | 0.3-0.4 | 0.9-1.0 | | Energetic/YouTuber | 0.3-0.5 | 0.75 | 0.5-0.7 | 1.0-1.1 |
### Pauses Between Sections
**With flash/turbo models:** Use SSML break tags inline: ``` ...end of section. <break time="1.5s" /> Start of next... ``` Max 3 seconds per break. Excessive breaks can cause speed artifacts.
**With multilingual_v2 / v3:** No SSML support. Options: - Paragraph breaks (blank lines) — creates ~0.3-0.5s natural pause - Post-process with ffmpeg: split audio and insert silence
**WARNING:** `...` (ellipsis) is NOT a reliable pause — it can be vocalized as a word/sound. Do not use ellipsis as a pause mechanism.
### Pronunciation Control
**Phonetic spelling (any model):** Write words as you want them pronounced: - `Janus` → `Jan-us` - `nginx` → `engine-x` - Use dashes, capitals, apostrophes to guide pronunciation
**SSML phoneme tags (flash/turbo only):** ``` <phoneme alphabet="ipa" ph="ˈdʒeɪnəs">Janus</phoneme> ```
### Iterative Workflow
1. Generate → listen → identify pronunciation/pacing issues 2. Adjust: phonetic spellings, break tags, voice settings 3. Regenerate. If pauses aren't precise enough, add silence in post with ffmpeg rather than fighting the TTS engine.
## Voice Cloning
### Instant Voice Clone
```python with open("sample.mp3", "rb") as f: voice = client.voices.ivc.create( name="My Voice", files=[f], remove_background_noise=True ) print(f"Voice ID: {voice.voice_id}") ```
- Use `client.voices.ivc.create()` (not `client.voices.clone()`) - Pass file handles in binary mode (`"rb"`), not paths - Convert m4a first: `ffmpeg -i input.m4a -codec:a libmp3lame -qscale:a 2 output.mp3` - Multiple samples (2-3 clips) improve accuracy - Save voice ID for reuse
**Professional Voice Clone:** Requires Creator plan+, 30+ min audio. See [reference.md](reference.md).
## Sound Effects
Max 22 seconds per generation.
```python result = client.text_to_sound_effects.convert( text="Thunder rumbling followed by heavy rain", duration_seconds=10, prompt_influence=0.3 ) with open("thunder.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt tips:** Be specific — "Heavy footsteps on wooden floorboards, slow and deliberate, with creaking"
## Music Generation
10 seconds to 5 minutes. Use `client.music.compose()` (not `.generate()`).
```python result = client.music.compose( prompt="Upbeat indie rock, catchy guitar riff, energetic drums, travel vlog", music_length_ms=60000, force_instrumental=True ) with open("music.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt structure:** Genre, mood, instruments, tempo, use case. Add "no vocals" or use `force_instrumental=True` for background music.
## Remotion Integration
### Complete Workflow: Script to Synchronized Scene
``` VOICEOVER-SCRIPT.md → voiceover.py → public/audio/ → Remotion composition ↓ ↓ ↓ ↓ Scene narration Generate MP3 Audio files <Audio> component with durations per scene with timing synced to scenes ```
### Step 1: Generate Per-Scene Audio
Use the toolkit's voiceover tool to generate audio for each scene:
```bash # Generate voiceover files for each scene python tools/voiceover.py --scene-dir public/audio/scenes --json
# Output: # public/audio/scenes/ # ├── scene-01-title.mp3 # ├── scene-02-problem.mp3 # ├── scene-03-solution.mp3 # └── manifest.json (durations for each file) ```
The `manifest.json` contains timing info: ```json { "scenes": [ { "file": "scene-01-title.mp3", "duration": 4.2 }, { "file": "scene-02-problem.mp3", "duration": 12.8 }, { "file": "scene-03-solution.mp3", "duration": 15.3 } ], "totalDuration": 32.3 } ```
### Step 2: Use Audio in Remotion Composition
```tsx // src/Composition.tsx import { Audio, staticFile, Series, useVideoConfig } from 'remotion';
// Import scene components import { TitleSlide } from './scenes/TitleSlide'; import { ProblemSlide } from './scenes/ProblemSlide'; import { SolutionSlide } from './scenes/SolutionSlide';
// Scene durations (from manifest.json, converted to frames at 30fps) const SCENE_DURATIONS = { title: Math.ceil(4.2 * 30), // 126 frames problem: Math.ceil(12.8 * 30), // 384 frames solution: Math.ceil(15.3 * 30), // 459 frames };
export const MainComposition: React.FC = () => { return ( <> {/* Scene sequence */} <Series> <Series.Sequence durationInFrames={SCENE_DURATIONS.title}> <TitleSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.problem}> <ProblemSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.solution}> <SolutionSlide /> </Series.Sequence> </Series>
{/* Audio track - plays continuously across all scenes */} <Audio src={staticFile('audio/voiceover.mp3')} volume={1} />
{/* Optional: Background music at lower volume */} <Audio src={staticFile('audio/music.mp3')} volume={0.15} /> </> ); }; ```
### Step 3: Per-Scene Audio (Alternative)
For more control, add audio to each scene individually:
```tsx // src/scenes/ProblemSlide.tsx import { Audio, staticFile, useCurrentFrame } from 'remotion';
export const ProblemSlide: React.FC = () => { const frame = useCurrentFrame();
return ( <div style={{ /* slide styles */ }}> <h1>The Problem</h1> {/* Scene content */}
{/* Audio starts when this scene starts (frame 0 of this sequence) */} <Audio src={staticFile('audio/scenes/scene-02-problem.mp3')} /> </div> ); }; ```
### Syncing Visuals to Voiceover
Calculate scene duration from audio, not the other way around:
```tsx // src/config/timing.ts import manifest from '../../public/audio/scenes/manifest.json';
const FPS = 30;
// Convert audio durations to frame counts export const sceneDurations = manifest.scenes.reduce((acc, scene) => { const name = scene.file.replace(/^scene-\d+-/, '').replace('.mp3', ''); acc[name] = Math.ceil(scene.duration * FPS); return acc; }, {} as Record<string, number>);
// Usage in composition: // <Series.Sequence durationInFrames={sceneDurations.title}> ```
### Audio Timing Patterns
```tsx import { Audio, Sequence, interpolate, useCurrentFrame } from 'remotion';
// Fade in audio export const FadeInAudio: React.FC<{ src: string; fadeFrames?: number }> = ({ src, fadeFrames = 30 }) => { const frame = useCurrentFrame(); const volume = interpolate(frame, [0, fadeFrames], [0, 1], { extrapolateRight: 'clamp', }); return <Audio src={src} volume={volume} />; };
// Delayed audio start export const DelayedAudio: React.FC<{ src: string; delayFrames: number }> = ({ src, delayFrames }) => ( <Sequence from={delayFrames}> <Audio src={src} /> </Sequence> );
// Usage: // <FadeInAudio src={staticFile('audio/music.mp3')} fadeFrames={60} /> // <DelayedAudio src={staticFile('audio/sfx/whoosh.mp3')} delayFrames={45} /> ```
### Voiceover + Demo Video Sync
When a scene has both voiceover and demo video:
```tsx import { Audio, OffthreadVideo, staticFile, useVideoConfig } from 'remotion';
export const DemoScene: React.FC = () => { const { durationInFrames, fps } = useVideoConfig();
// Calculate playback rate to fit demo into voiceover duration const demoDuration = 45; // seconds (original demo length) const sceneDuration = durationInFrames / fps; // seconds (from voiceover) const playbackRate = demoDuration / sceneDuration;
return ( <> <OffthreadVideo src={staticFile('demos/feature-demo.mp4')} playbackRate={playbackRate} /> <Audio src={staticFile('audio/scenes/scene-04-demo.mp3')} /> </> ); }; ```
### Error Handling
```tsx import { Audio, staticFile, delayRender, continueRender } from 'remotion'; import { useEffect, useState } from 'react';
export const SafeAudio: React.FC<{ src: string }> = ({ src }) => { const [handle] = useState(() => delayRender()); const [audioReady, setAudioReady] = useState(false);
useEffect(() => { const audio = new window.Audio(src); audio.oncanplaythrough = () => { setAudioReady(true); continueRender(handle); }; audio.onerror = () => { console.error(`Failed to load audio: ${src}`); continueRender(handle); // Continue without audio rather than hang }; }, [src, handle]);
if (!audioReady) return null; return <Audio src={src} />; }; ```
### Toolkit Command: /generate-voiceover
The `/generate-voiceover` command handles the full workflow:
``` /generate-voiceover
1. Reads VOICEOVER-SCRIPT.md 2. Extracts narration for each scene 3. Generates audio via ElevenLabs API 4. Saves to public/audio/scenes/ 5. Creates manifest.json with durations 6. Updates project.json with timing info ```
## Popular Voices
- George: `JBFqnCBsd6RMkjVDRZzb` (warm narrator) - Rachel: `21m00Tcm4TlvDq8ikWAM` (clear female) - Adam: `pNInz6obpgDQGcFmaJgB` (professional male)
List all: `client.voices.get_all()`
For full API docs, see [reference.md](reference.md).
Source provenance
Decision snapshot
55,318 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for elevenlabs, ready for a manual X post.
elevenlabs: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating aud... 55.3K stars https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x
Listing + install path for elevenlabs: https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x Install: npx skills add calesthio/OpenMontage --skill elevenlabs
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to calesthio but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs/audit)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)calesthio
@calesthio
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
174.9K StarsTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
84.9K StarsVox Director
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K StarsCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
174.9K StarsSandbox only
Install targets
Codex install prompt
Install the "elevenlabs" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/elevenlabs. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"calesthio-elevenlabs","task":"Install elevenlabs","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Scenario
Design and creative
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add calesthio/OpenMontage --skill elevenlabs
Maintenance
fresh
16d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
55K
93/100 Quality · 80/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
55K GitHub stars
Repo activity
55K stars, 6.9K forks
Maintenance
16d since push
License
AGPL-3.0
Install
npx skills add calesthio/OpenMontage --skill elevenlabs
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add calesthio/OpenMontage --skill elevenlabsDo not use when
Alternative
174.9K Stars
npx skills add anthropics/skills --skill frontend-design
Alternative
84.9K Stars
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Alternative
1.8K Stars
npx skills add Alisa0808/vox-director --skill vox-director
Alternative
174.9K Stars
npx skills add anthropics/skills --skill canvas-design
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/calesthio-elevenlabs/install
Agent should check
Copy prompt
Task: Use elevenlabs in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install
Install command: npx skills add calesthio/OpenMontage --skill elevenlabs
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/calesthio-elevenlabs/install
LLM text format
/api/skills/calesthio-elevenlabs/install?format=text
Find alternatives
/api/skills/search?q=elevenlabs&limit=3
Agent prompt
Use elevenlabs for this task. Review https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install, then install with: npx skills add calesthio/OpenMontage --skill elevenlabsRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/calesthio-elevenlabs
LLM text
/api/registry/manifest/calesthio-elevenlabs?format=text
Install alias
/api/registry/install/calesthio-elevenlabs
Recommend
/api/registry/recommend?task=Use%20elevenlabs%20in%20an%20agent%20workflow&limit=3
Agent fit
Workflow automation
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Workflow automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS55K GitHub stars
Stars/forks activity
PASS55K stars, 6.9K forks; issue activity unavailable in current metadata
Recent maintenance
PASS16d since push
License clarity
PASSAGPL-3.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Design, build, test, and ship interfaces
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
--- name: elevenlabs description: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. ---
# ElevenLabs Audio Generation
## OpenMontage provider routing
Inspect the OpenMontage registry before choosing an authentication path.
- Prefer `fal_elevenlabs_tts` when it is available. It provides Eleven v3, Multilingual v2, and Turbo v2.5 through the centrally managed fal.ai connection; no separate ElevenLabs credential is needed. - Use `elevenlabs_tts` only when that direct provider is already reported as available by the registry. - In a shared installation, never tell the user to create a `.env`, export a key, or paste a credential. Report missing direct-provider access as an administrator setup request.
The direct API examples below require a centrally configured `ELEVENLABS_API_KEY`; they are not the default path when the fal.ai provider is available.
## Text-to-Speech
```python from elevenlabs.client import ElevenLabs from elevenlabs import save, VoiceSettings import os
client = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
audio = client.text_to_speech.convert( text="Welcome to my video!", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2", voice_settings=VoiceSettings( stability=0.5, similarity_boost=0.75, style=0.5, speed=1.0 ) ) save(audio, "voiceover.mp3") ```
### Models
| Model | Quality | SSML Support | Notes | |-------|---------|--------------|-------| | `eleven_multilingual_v2` | Highest consistency | None | Stable, production-ready, 29 languages | | `eleven_flash_v2_5` | Good | `<break>`, `<phoneme>` | Fast, supports pause/pronunciation tags | | `eleven_turbo_v2_5` | Good | `<break>`, `<phoneme>` | Fastest latency | | `eleven_v3` | Most expressive | None | Alpha — unreliable, needs prompt engineering |
**Choose:** multilingual_v2 for reliability, flash/turbo for SSML control, v3 for maximum expressiveness (expect retakes).
### Voice Settings by Style
| Style | stability | similarity | style | speed | |-------|-----------|------------|-------|-------| | Natural/professional | 0.75-0.85 | 0.9 | 0.0-0.1 | 1.0 | | Conversational | 0.5-0.6 | 0.85 | 0.3-0.4 | 0.9-1.0 | | Energetic/YouTuber | 0.3-0.5 | 0.75 | 0.5-0.7 | 1.0-1.1 |
### Pauses Between Sections
**With flash/turbo models:** Use SSML break tags inline: ``` ...end of section. <break time="1.5s" /> Start of next... ``` Max 3 seconds per break. Excessive breaks can cause speed artifacts.
**With multilingual_v2 / v3:** No SSML support. Options: - Paragraph breaks (blank lines) — creates ~0.3-0.5s natural pause - Post-process with ffmpeg: split audio and insert silence
**WARNING:** `...` (ellipsis) is NOT a reliable pause — it can be vocalized as a word/sound. Do not use ellipsis as a pause mechanism.
### Pronunciation Control
**Phonetic spelling (any model):** Write words as you want them pronounced: - `Janus` → `Jan-us` - `nginx` → `engine-x` - Use dashes, capitals, apostrophes to guide pronunciation
**SSML phoneme tags (flash/turbo only):** ``` <phoneme alphabet="ipa" ph="ˈdʒeɪnəs">Janus</phoneme> ```
### Iterative Workflow
1. Generate → listen → identify pronunciation/pacing issues 2. Adjust: phonetic spellings, break tags, voice settings 3. Regenerate. If pauses aren't precise enough, add silence in post with ffmpeg rather than fighting the TTS engine.
## Voice Cloning
### Instant Voice Clone
```python with open("sample.mp3", "rb") as f: voice = client.voices.ivc.create( name="My Voice", files=[f], remove_background_noise=True ) print(f"Voice ID: {voice.voice_id}") ```
- Use `client.voices.ivc.create()` (not `client.voices.clone()`) - Pass file handles in binary mode (`"rb"`), not paths - Convert m4a first: `ffmpeg -i input.m4a -codec:a libmp3lame -qscale:a 2 output.mp3` - Multiple samples (2-3 clips) improve accuracy - Save voice ID for reuse
**Professional Voice Clone:** Requires Creator plan+, 30+ min audio. See [reference.md](reference.md).
## Sound Effects
Max 22 seconds per generation.
```python result = client.text_to_sound_effects.convert( text="Thunder rumbling followed by heavy rain", duration_seconds=10, prompt_influence=0.3 ) with open("thunder.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt tips:** Be specific — "Heavy footsteps on wooden floorboards, slow and deliberate, with creaking"
## Music Generation
10 seconds to 5 minutes. Use `client.music.compose()` (not `.generate()`).
```python result = client.music.compose( prompt="Upbeat indie rock, catchy guitar riff, energetic drums, travel vlog", music_length_ms=60000, force_instrumental=True ) with open("music.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt structure:** Genre, mood, instruments, tempo, use case. Add "no vocals" or use `force_instrumental=True` for background music.
## Remotion Integration
### Complete Workflow: Script to Synchronized Scene
``` VOICEOVER-SCRIPT.md → voiceover.py → public/audio/ → Remotion composition ↓ ↓ ↓ ↓ Scene narration Generate MP3 Audio files <Audio> component with durations per scene with timing synced to scenes ```
### Step 1: Generate Per-Scene Audio
Use the toolkit's voiceover tool to generate audio for each scene:
```bash # Generate voiceover files for each scene python tools/voiceover.py --scene-dir public/audio/scenes --json
# Output: # public/audio/scenes/ # ├── scene-01-title.mp3 # ├── scene-02-problem.mp3 # ├── scene-03-solution.mp3 # └── manifest.json (durations for each file) ```
The `manifest.json` contains timing info: ```json { "scenes": [ { "file": "scene-01-title.mp3", "duration": 4.2 }, { "file": "scene-02-problem.mp3", "duration": 12.8 }, { "file": "scene-03-solution.mp3", "duration": 15.3 } ], "totalDuration": 32.3 } ```
### Step 2: Use Audio in Remotion Composition
```tsx // src/Composition.tsx import { Audio, staticFile, Series, useVideoConfig } from 'remotion';
// Import scene components import { TitleSlide } from './scenes/TitleSlide'; import { ProblemSlide } from './scenes/ProblemSlide'; import { SolutionSlide } from './scenes/SolutionSlide';
// Scene durations (from manifest.json, converted to frames at 30fps) const SCENE_DURATIONS = { title: Math.ceil(4.2 * 30), // 126 frames problem: Math.ceil(12.8 * 30), // 384 frames solution: Math.ceil(15.3 * 30), // 459 frames };
export const MainComposition: React.FC = () => { return ( <> {/* Scene sequence */} <Series> <Series.Sequence durationInFrames={SCENE_DURATIONS.title}> <TitleSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.problem}> <ProblemSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.solution}> <SolutionSlide /> </Series.Sequence> </Series>
{/* Audio track - plays continuously across all scenes */} <Audio src={staticFile('audio/voiceover.mp3')} volume={1} />
{/* Optional: Background music at lower volume */} <Audio src={staticFile('audio/music.mp3')} volume={0.15} /> </> ); }; ```
### Step 3: Per-Scene Audio (Alternative)
For more control, add audio to each scene individually:
```tsx // src/scenes/ProblemSlide.tsx import { Audio, staticFile, useCurrentFrame } from 'remotion';
export const ProblemSlide: React.FC = () => { const frame = useCurrentFrame();
return ( <div style={{ /* slide styles */ }}> <h1>The Problem</h1> {/* Scene content */}
{/* Audio starts when this scene starts (frame 0 of this sequence) */} <Audio src={staticFile('audio/scenes/scene-02-problem.mp3')} /> </div> ); }; ```
### Syncing Visuals to Voiceover
Calculate scene duration from audio, not the other way around:
```tsx // src/config/timing.ts import manifest from '../../public/audio/scenes/manifest.json';
const FPS = 30;
// Convert audio durations to frame counts export const sceneDurations = manifest.scenes.reduce((acc, scene) => { const name = scene.file.replace(/^scene-\d+-/, '').replace('.mp3', ''); acc[name] = Math.ceil(scene.duration * FPS); return acc; }, {} as Record<string, number>);
// Usage in composition: // <Series.Sequence durationInFrames={sceneDurations.title}> ```
### Audio Timing Patterns
```tsx import { Audio, Sequence, interpolate, useCurrentFrame } from 'remotion';
// Fade in audio export const FadeInAudio: React.FC<{ src: string; fadeFrames?: number }> = ({ src, fadeFrames = 30 }) => { const frame = useCurrentFrame(); const volume = interpolate(frame, [0, fadeFrames], [0, 1], { extrapolateRight: 'clamp', }); return <Audio src={src} volume={volume} />; };
// Delayed audio start export const DelayedAudio: React.FC<{ src: string; delayFrames: number }> = ({ src, delayFrames }) => ( <Sequence from={delayFrames}> <Audio src={src} /> </Sequence> );
// Usage: // <FadeInAudio src={staticFile('audio/music.mp3')} fadeFrames={60} /> // <DelayedAudio src={staticFile('audio/sfx/whoosh.mp3')} delayFrames={45} /> ```
### Voiceover + Demo Video Sync
When a scene has both voiceover and demo video:
```tsx import { Audio, OffthreadVideo, staticFile, useVideoConfig } from 'remotion';
export const DemoScene: React.FC = () => { const { durationInFrames, fps } = useVideoConfig();
// Calculate playback rate to fit demo into voiceover duration const demoDuration = 45; // seconds (original demo length) const sceneDuration = durationInFrames / fps; // seconds (from voiceover) const playbackRate = demoDuration / sceneDuration;
return ( <> <OffthreadVideo src={staticFile('demos/feature-demo.mp4')} playbackRate={playbackRate} /> <Audio src={staticFile('audio/scenes/scene-04-demo.mp3')} /> </> ); }; ```
### Error Handling
```tsx import { Audio, staticFile, delayRender, continueRender } from 'remotion'; import { useEffect, useState } from 'react';
export const SafeAudio: React.FC<{ src: string }> = ({ src }) => { const [handle] = useState(() => delayRender()); const [audioReady, setAudioReady] = useState(false);
useEffect(() => { const audio = new window.Audio(src); audio.oncanplaythrough = () => { setAudioReady(true); continueRender(handle); }; audio.onerror = () => { console.error(`Failed to load audio: ${src}`); continueRender(handle); // Continue without audio rather than hang }; }, [src, handle]);
if (!audioReady) return null; return <Audio src={src} />; }; ```
### Toolkit Command: /generate-voiceover
The `/generate-voiceover` command handles the full workflow:
``` /generate-voiceover
1. Reads VOICEOVER-SCRIPT.md 2. Extracts narration for each scene 3. Generates audio via ElevenLabs API 4. Saves to public/audio/scenes/ 5. Creates manifest.json with durations 6. Updates project.json with timing info ```
## Popular Voices
- George: `JBFqnCBsd6RMkjVDRZzb` (warm narrator) - Rachel: `21m00Tcm4TlvDq8ikWAM` (clear female) - Adam: `pNInz6obpgDQGcFmaJgB` (professional male)
List all: `client.voices.get_all()`
For full API docs, see [reference.md](reference.md).
Source provenance
Decision snapshot
55,318 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for elevenlabs, ready for a manual X post.
elevenlabs: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating aud... 55.3K stars https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x
Listing + install path for elevenlabs: https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x Install: npx skills add calesthio/OpenMontage --skill elevenlabs
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to calesthio but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs/audit)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)calesthio
@calesthio
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
174.9K StarsTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
84.9K StarsVox Director
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K StarsCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
174.9K StarsSandbox only
Install targets
Codex install prompt
Install the "elevenlabs" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/elevenlabs. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"calesthio-elevenlabs","task":"Install elevenlabs","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Scenario
Design and creative
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add calesthio/OpenMontage --skill elevenlabs
Maintenance
fresh
16d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
55K
93/100 Quality · 80/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
55K GitHub stars
Repo activity
55K stars, 6.9K forks
Maintenance
16d since push
License
AGPL-3.0
Install
npx skills add calesthio/OpenMontage --skill elevenlabs
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add calesthio/OpenMontage --skill elevenlabsDo not use when
Alternative
174.9K Stars
npx skills add anthropics/skills --skill frontend-design
Alternative
84.9K Stars
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Alternative
1.8K Stars
npx skills add Alisa0808/vox-director --skill vox-director
Alternative
174.9K Stars
npx skills add anthropics/skills --skill canvas-design
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/calesthio-elevenlabs/install
Agent should check
Copy prompt
Task: Use elevenlabs in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install
Install command: npx skills add calesthio/OpenMontage --skill elevenlabs
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/calesthio-elevenlabs/install
LLM text format
/api/skills/calesthio-elevenlabs/install?format=text
Find alternatives
/api/skills/search?q=elevenlabs&limit=3
Agent prompt
Use elevenlabs for this task. Review https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install, then install with: npx skills add calesthio/OpenMontage --skill elevenlabsRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/calesthio-elevenlabs
LLM text
/api/registry/manifest/calesthio-elevenlabs?format=text
Install alias
/api/registry/install/calesthio-elevenlabs
Recommend
/api/registry/recommend?task=Use%20elevenlabs%20in%20an%20agent%20workflow&limit=3
Agent fit
Workflow automation
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Workflow automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS55K GitHub stars
Stars/forks activity
PASS55K stars, 6.9K forks; issue activity unavailable in current metadata
Recent maintenance
PASS16d since push
License clarity
PASSAGPL-3.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Design, build, test, and ship interfaces
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
--- name: elevenlabs description: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. ---
# ElevenLabs Audio Generation
## OpenMontage provider routing
Inspect the OpenMontage registry before choosing an authentication path.
- Prefer `fal_elevenlabs_tts` when it is available. It provides Eleven v3, Multilingual v2, and Turbo v2.5 through the centrally managed fal.ai connection; no separate ElevenLabs credential is needed. - Use `elevenlabs_tts` only when that direct provider is already reported as available by the registry. - In a shared installation, never tell the user to create a `.env`, export a key, or paste a credential. Report missing direct-provider access as an administrator setup request.
The direct API examples below require a centrally configured `ELEVENLABS_API_KEY`; they are not the default path when the fal.ai provider is available.
## Text-to-Speech
```python from elevenlabs.client import ElevenLabs from elevenlabs import save, VoiceSettings import os
client = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
audio = client.text_to_speech.convert( text="Welcome to my video!", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2", voice_settings=VoiceSettings( stability=0.5, similarity_boost=0.75, style=0.5, speed=1.0 ) ) save(audio, "voiceover.mp3") ```
### Models
| Model | Quality | SSML Support | Notes | |-------|---------|--------------|-------| | `eleven_multilingual_v2` | Highest consistency | None | Stable, production-ready, 29 languages | | `eleven_flash_v2_5` | Good | `<break>`, `<phoneme>` | Fast, supports pause/pronunciation tags | | `eleven_turbo_v2_5` | Good | `<break>`, `<phoneme>` | Fastest latency | | `eleven_v3` | Most expressive | None | Alpha — unreliable, needs prompt engineering |
**Choose:** multilingual_v2 for reliability, flash/turbo for SSML control, v3 for maximum expressiveness (expect retakes).
### Voice Settings by Style
| Style | stability | similarity | style | speed | |-------|-----------|------------|-------|-------| | Natural/professional | 0.75-0.85 | 0.9 | 0.0-0.1 | 1.0 | | Conversational | 0.5-0.6 | 0.85 | 0.3-0.4 | 0.9-1.0 | | Energetic/YouTuber | 0.3-0.5 | 0.75 | 0.5-0.7 | 1.0-1.1 |
### Pauses Between Sections
**With flash/turbo models:** Use SSML break tags inline: ``` ...end of section. <break time="1.5s" /> Start of next... ``` Max 3 seconds per break. Excessive breaks can cause speed artifacts.
**With multilingual_v2 / v3:** No SSML support. Options: - Paragraph breaks (blank lines) — creates ~0.3-0.5s natural pause - Post-process with ffmpeg: split audio and insert silence
**WARNING:** `...` (ellipsis) is NOT a reliable pause — it can be vocalized as a word/sound. Do not use ellipsis as a pause mechanism.
### Pronunciation Control
**Phonetic spelling (any model):** Write words as you want them pronounced: - `Janus` → `Jan-us` - `nginx` → `engine-x` - Use dashes, capitals, apostrophes to guide pronunciation
**SSML phoneme tags (flash/turbo only):** ``` <phoneme alphabet="ipa" ph="ˈdʒeɪnəs">Janus</phoneme> ```
### Iterative Workflow
1. Generate → listen → identify pronunciation/pacing issues 2. Adjust: phonetic spellings, break tags, voice settings 3. Regenerate. If pauses aren't precise enough, add silence in post with ffmpeg rather than fighting the TTS engine.
## Voice Cloning
### Instant Voice Clone
```python with open("sample.mp3", "rb") as f: voice = client.voices.ivc.create( name="My Voice", files=[f], remove_background_noise=True ) print(f"Voice ID: {voice.voice_id}") ```
- Use `client.voices.ivc.create()` (not `client.voices.clone()`) - Pass file handles in binary mode (`"rb"`), not paths - Convert m4a first: `ffmpeg -i input.m4a -codec:a libmp3lame -qscale:a 2 output.mp3` - Multiple samples (2-3 clips) improve accuracy - Save voice ID for reuse
**Professional Voice Clone:** Requires Creator plan+, 30+ min audio. See [reference.md](reference.md).
## Sound Effects
Max 22 seconds per generation.
```python result = client.text_to_sound_effects.convert( text="Thunder rumbling followed by heavy rain", duration_seconds=10, prompt_influence=0.3 ) with open("thunder.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt tips:** Be specific — "Heavy footsteps on wooden floorboards, slow and deliberate, with creaking"
## Music Generation
10 seconds to 5 minutes. Use `client.music.compose()` (not `.generate()`).
```python result = client.music.compose( prompt="Upbeat indie rock, catchy guitar riff, energetic drums, travel vlog", music_length_ms=60000, force_instrumental=True ) with open("music.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt structure:** Genre, mood, instruments, tempo, use case. Add "no vocals" or use `force_instrumental=True` for background music.
## Remotion Integration
### Complete Workflow: Script to Synchronized Scene
``` VOICEOVER-SCRIPT.md → voiceover.py → public/audio/ → Remotion composition ↓ ↓ ↓ ↓ Scene narration Generate MP3 Audio files <Audio> component with durations per scene with timing synced to scenes ```
### Step 1: Generate Per-Scene Audio
Use the toolkit's voiceover tool to generate audio for each scene:
```bash # Generate voiceover files for each scene python tools/voiceover.py --scene-dir public/audio/scenes --json
# Output: # public/audio/scenes/ # ├── scene-01-title.mp3 # ├── scene-02-problem.mp3 # ├── scene-03-solution.mp3 # └── manifest.json (durations for each file) ```
The `manifest.json` contains timing info: ```json { "scenes": [ { "file": "scene-01-title.mp3", "duration": 4.2 }, { "file": "scene-02-problem.mp3", "duration": 12.8 }, { "file": "scene-03-solution.mp3", "duration": 15.3 } ], "totalDuration": 32.3 } ```
### Step 2: Use Audio in Remotion Composition
```tsx // src/Composition.tsx import { Audio, staticFile, Series, useVideoConfig } from 'remotion';
// Import scene components import { TitleSlide } from './scenes/TitleSlide'; import { ProblemSlide } from './scenes/ProblemSlide'; import { SolutionSlide } from './scenes/SolutionSlide';
// Scene durations (from manifest.json, converted to frames at 30fps) const SCENE_DURATIONS = { title: Math.ceil(4.2 * 30), // 126 frames problem: Math.ceil(12.8 * 30), // 384 frames solution: Math.ceil(15.3 * 30), // 459 frames };
export const MainComposition: React.FC = () => { return ( <> {/* Scene sequence */} <Series> <Series.Sequence durationInFrames={SCENE_DURATIONS.title}> <TitleSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.problem}> <ProblemSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.solution}> <SolutionSlide /> </Series.Sequence> </Series>
{/* Audio track - plays continuously across all scenes */} <Audio src={staticFile('audio/voiceover.mp3')} volume={1} />
{/* Optional: Background music at lower volume */} <Audio src={staticFile('audio/music.mp3')} volume={0.15} /> </> ); }; ```
### Step 3: Per-Scene Audio (Alternative)
For more control, add audio to each scene individually:
```tsx // src/scenes/ProblemSlide.tsx import { Audio, staticFile, useCurrentFrame } from 'remotion';
export const ProblemSlide: React.FC = () => { const frame = useCurrentFrame();
return ( <div style={{ /* slide styles */ }}> <h1>The Problem</h1> {/* Scene content */}
{/* Audio starts when this scene starts (frame 0 of this sequence) */} <Audio src={staticFile('audio/scenes/scene-02-problem.mp3')} /> </div> ); }; ```
### Syncing Visuals to Voiceover
Calculate scene duration from audio, not the other way around:
```tsx // src/config/timing.ts import manifest from '../../public/audio/scenes/manifest.json';
const FPS = 30;
// Convert audio durations to frame counts export const sceneDurations = manifest.scenes.reduce((acc, scene) => { const name = scene.file.replace(/^scene-\d+-/, '').replace('.mp3', ''); acc[name] = Math.ceil(scene.duration * FPS); return acc; }, {} as Record<string, number>);
// Usage in composition: // <Series.Sequence durationInFrames={sceneDurations.title}> ```
### Audio Timing Patterns
```tsx import { Audio, Sequence, interpolate, useCurrentFrame } from 'remotion';
// Fade in audio export const FadeInAudio: React.FC<{ src: string; fadeFrames?: number }> = ({ src, fadeFrames = 30 }) => { const frame = useCurrentFrame(); const volume = interpolate(frame, [0, fadeFrames], [0, 1], { extrapolateRight: 'clamp', }); return <Audio src={src} volume={volume} />; };
// Delayed audio start export const DelayedAudio: React.FC<{ src: string; delayFrames: number }> = ({ src, delayFrames }) => ( <Sequence from={delayFrames}> <Audio src={src} /> </Sequence> );
// Usage: // <FadeInAudio src={staticFile('audio/music.mp3')} fadeFrames={60} /> // <DelayedAudio src={staticFile('audio/sfx/whoosh.mp3')} delayFrames={45} /> ```
### Voiceover + Demo Video Sync
When a scene has both voiceover and demo video:
```tsx import { Audio, OffthreadVideo, staticFile, useVideoConfig } from 'remotion';
export const DemoScene: React.FC = () => { const { durationInFrames, fps } = useVideoConfig();
// Calculate playback rate to fit demo into voiceover duration const demoDuration = 45; // seconds (original demo length) const sceneDuration = durationInFrames / fps; // seconds (from voiceover) const playbackRate = demoDuration / sceneDuration;
return ( <> <OffthreadVideo src={staticFile('demos/feature-demo.mp4')} playbackRate={playbackRate} /> <Audio src={staticFile('audio/scenes/scene-04-demo.mp3')} /> </> ); }; ```
### Error Handling
```tsx import { Audio, staticFile, delayRender, continueRender } from 'remotion'; import { useEffect, useState } from 'react';
export const SafeAudio: React.FC<{ src: string }> = ({ src }) => { const [handle] = useState(() => delayRender()); const [audioReady, setAudioReady] = useState(false);
useEffect(() => { const audio = new window.Audio(src); audio.oncanplaythrough = () => { setAudioReady(true); continueRender(handle); }; audio.onerror = () => { console.error(`Failed to load audio: ${src}`); continueRender(handle); // Continue without audio rather than hang }; }, [src, handle]);
if (!audioReady) return null; return <Audio src={src} />; }; ```
### Toolkit Command: /generate-voiceover
The `/generate-voiceover` command handles the full workflow:
``` /generate-voiceover
1. Reads VOICEOVER-SCRIPT.md 2. Extracts narration for each scene 3. Generates audio via ElevenLabs API 4. Saves to public/audio/scenes/ 5. Creates manifest.json with durations 6. Updates project.json with timing info ```
## Popular Voices
- George: `JBFqnCBsd6RMkjVDRZzb` (warm narrator) - Rachel: `21m00Tcm4TlvDq8ikWAM` (clear female) - Adam: `pNInz6obpgDQGcFmaJgB` (professional male)
List all: `client.voices.get_all()`
For full API docs, see [reference.md](reference.md).
Source provenance
Decision snapshot
55,318 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for elevenlabs, ready for a manual X post.
elevenlabs: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating aud... 55.3K stars https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x
Listing + install path for elevenlabs: https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x Install: npx skills add calesthio/OpenMontage --skill elevenlabs
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to calesthio but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs/audit)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)calesthio
@calesthio
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
174.9K StarsTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
84.9K StarsVox Director
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K StarsCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
174.9K StarsSandbox only
Install targets
Codex install prompt
Install the "elevenlabs" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/elevenlabs. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"calesthio-elevenlabs","task":"Install elevenlabs","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Scenario
Design and creative
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add calesthio/OpenMontage --skill elevenlabs
Maintenance
fresh
16d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
55K
93/100 Quality · 80/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
55K GitHub stars
Repo activity
55K stars, 6.9K forks
Maintenance
16d since push
License
AGPL-3.0
Install
npx skills add calesthio/OpenMontage --skill elevenlabs
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add calesthio/OpenMontage --skill elevenlabsDo not use when
Alternative
174.9K Stars
npx skills add anthropics/skills --skill frontend-design
Alternative
84.9K Stars
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Alternative
1.8K Stars
npx skills add Alisa0808/vox-director --skill vox-director
Alternative
174.9K Stars
npx skills add anthropics/skills --skill canvas-design
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/calesthio-elevenlabs/install
Agent should check
Copy prompt
Task: Use elevenlabs in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20elevenlabs%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install
Install command: npx skills add calesthio/OpenMontage --skill elevenlabs
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/calesthio-elevenlabs/install
LLM text format
/api/skills/calesthio-elevenlabs/install?format=text
Find alternatives
/api/skills/search?q=elevenlabs&limit=3
Agent prompt
Use elevenlabs for this task. Review https://www.openagentskill.com/api/skills/calesthio-elevenlabs/install, then install with: npx skills add calesthio/OpenMontage --skill elevenlabsRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/calesthio-elevenlabs
LLM text
/api/registry/manifest/calesthio-elevenlabs?format=text
Install alias
/api/registry/install/calesthio-elevenlabs
Recommend
/api/registry/recommend?task=Use%20elevenlabs%20in%20an%20agent%20workflow&limit=3
Agent fit
Workflow automation
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Workflow automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS55K GitHub stars
Stars/forks activity
PASS55K stars, 6.9K forks; issue activity unavailable in current metadata
Recent maintenance
PASS16d since push
License clarity
PASSAGPL-3.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Parse messy files
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Design, build, test, and ship interfaces
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
--- name: elevenlabs description: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio synthesis task. ---
# ElevenLabs Audio Generation
## OpenMontage provider routing
Inspect the OpenMontage registry before choosing an authentication path.
- Prefer `fal_elevenlabs_tts` when it is available. It provides Eleven v3, Multilingual v2, and Turbo v2.5 through the centrally managed fal.ai connection; no separate ElevenLabs credential is needed. - Use `elevenlabs_tts` only when that direct provider is already reported as available by the registry. - In a shared installation, never tell the user to create a `.env`, export a key, or paste a credential. Report missing direct-provider access as an administrator setup request.
The direct API examples below require a centrally configured `ELEVENLABS_API_KEY`; they are not the default path when the fal.ai provider is available.
## Text-to-Speech
```python from elevenlabs.client import ElevenLabs from elevenlabs import save, VoiceSettings import os
client = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
audio = client.text_to_speech.convert( text="Welcome to my video!", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2", voice_settings=VoiceSettings( stability=0.5, similarity_boost=0.75, style=0.5, speed=1.0 ) ) save(audio, "voiceover.mp3") ```
### Models
| Model | Quality | SSML Support | Notes | |-------|---------|--------------|-------| | `eleven_multilingual_v2` | Highest consistency | None | Stable, production-ready, 29 languages | | `eleven_flash_v2_5` | Good | `<break>`, `<phoneme>` | Fast, supports pause/pronunciation tags | | `eleven_turbo_v2_5` | Good | `<break>`, `<phoneme>` | Fastest latency | | `eleven_v3` | Most expressive | None | Alpha — unreliable, needs prompt engineering |
**Choose:** multilingual_v2 for reliability, flash/turbo for SSML control, v3 for maximum expressiveness (expect retakes).
### Voice Settings by Style
| Style | stability | similarity | style | speed | |-------|-----------|------------|-------|-------| | Natural/professional | 0.75-0.85 | 0.9 | 0.0-0.1 | 1.0 | | Conversational | 0.5-0.6 | 0.85 | 0.3-0.4 | 0.9-1.0 | | Energetic/YouTuber | 0.3-0.5 | 0.75 | 0.5-0.7 | 1.0-1.1 |
### Pauses Between Sections
**With flash/turbo models:** Use SSML break tags inline: ``` ...end of section. <break time="1.5s" /> Start of next... ``` Max 3 seconds per break. Excessive breaks can cause speed artifacts.
**With multilingual_v2 / v3:** No SSML support. Options: - Paragraph breaks (blank lines) — creates ~0.3-0.5s natural pause - Post-process with ffmpeg: split audio and insert silence
**WARNING:** `...` (ellipsis) is NOT a reliable pause — it can be vocalized as a word/sound. Do not use ellipsis as a pause mechanism.
### Pronunciation Control
**Phonetic spelling (any model):** Write words as you want them pronounced: - `Janus` → `Jan-us` - `nginx` → `engine-x` - Use dashes, capitals, apostrophes to guide pronunciation
**SSML phoneme tags (flash/turbo only):** ``` <phoneme alphabet="ipa" ph="ˈdʒeɪnəs">Janus</phoneme> ```
### Iterative Workflow
1. Generate → listen → identify pronunciation/pacing issues 2. Adjust: phonetic spellings, break tags, voice settings 3. Regenerate. If pauses aren't precise enough, add silence in post with ffmpeg rather than fighting the TTS engine.
## Voice Cloning
### Instant Voice Clone
```python with open("sample.mp3", "rb") as f: voice = client.voices.ivc.create( name="My Voice", files=[f], remove_background_noise=True ) print(f"Voice ID: {voice.voice_id}") ```
- Use `client.voices.ivc.create()` (not `client.voices.clone()`) - Pass file handles in binary mode (`"rb"`), not paths - Convert m4a first: `ffmpeg -i input.m4a -codec:a libmp3lame -qscale:a 2 output.mp3` - Multiple samples (2-3 clips) improve accuracy - Save voice ID for reuse
**Professional Voice Clone:** Requires Creator plan+, 30+ min audio. See [reference.md](reference.md).
## Sound Effects
Max 22 seconds per generation.
```python result = client.text_to_sound_effects.convert( text="Thunder rumbling followed by heavy rain", duration_seconds=10, prompt_influence=0.3 ) with open("thunder.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt tips:** Be specific — "Heavy footsteps on wooden floorboards, slow and deliberate, with creaking"
## Music Generation
10 seconds to 5 minutes. Use `client.music.compose()` (not `.generate()`).
```python result = client.music.compose( prompt="Upbeat indie rock, catchy guitar riff, energetic drums, travel vlog", music_length_ms=60000, force_instrumental=True ) with open("music.mp3", "wb") as f: for chunk in result: f.write(chunk) ```
**Prompt structure:** Genre, mood, instruments, tempo, use case. Add "no vocals" or use `force_instrumental=True` for background music.
## Remotion Integration
### Complete Workflow: Script to Synchronized Scene
``` VOICEOVER-SCRIPT.md → voiceover.py → public/audio/ → Remotion composition ↓ ↓ ↓ ↓ Scene narration Generate MP3 Audio files <Audio> component with durations per scene with timing synced to scenes ```
### Step 1: Generate Per-Scene Audio
Use the toolkit's voiceover tool to generate audio for each scene:
```bash # Generate voiceover files for each scene python tools/voiceover.py --scene-dir public/audio/scenes --json
# Output: # public/audio/scenes/ # ├── scene-01-title.mp3 # ├── scene-02-problem.mp3 # ├── scene-03-solution.mp3 # └── manifest.json (durations for each file) ```
The `manifest.json` contains timing info: ```json { "scenes": [ { "file": "scene-01-title.mp3", "duration": 4.2 }, { "file": "scene-02-problem.mp3", "duration": 12.8 }, { "file": "scene-03-solution.mp3", "duration": 15.3 } ], "totalDuration": 32.3 } ```
### Step 2: Use Audio in Remotion Composition
```tsx // src/Composition.tsx import { Audio, staticFile, Series, useVideoConfig } from 'remotion';
// Import scene components import { TitleSlide } from './scenes/TitleSlide'; import { ProblemSlide } from './scenes/ProblemSlide'; import { SolutionSlide } from './scenes/SolutionSlide';
// Scene durations (from manifest.json, converted to frames at 30fps) const SCENE_DURATIONS = { title: Math.ceil(4.2 * 30), // 126 frames problem: Math.ceil(12.8 * 30), // 384 frames solution: Math.ceil(15.3 * 30), // 459 frames };
export const MainComposition: React.FC = () => { return ( <> {/* Scene sequence */} <Series> <Series.Sequence durationInFrames={SCENE_DURATIONS.title}> <TitleSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.problem}> <ProblemSlide /> </Series.Sequence> <Series.Sequence durationInFrames={SCENE_DURATIONS.solution}> <SolutionSlide /> </Series.Sequence> </Series>
{/* Audio track - plays continuously across all scenes */} <Audio src={staticFile('audio/voiceover.mp3')} volume={1} />
{/* Optional: Background music at lower volume */} <Audio src={staticFile('audio/music.mp3')} volume={0.15} /> </> ); }; ```
### Step 3: Per-Scene Audio (Alternative)
For more control, add audio to each scene individually:
```tsx // src/scenes/ProblemSlide.tsx import { Audio, staticFile, useCurrentFrame } from 'remotion';
export const ProblemSlide: React.FC = () => { const frame = useCurrentFrame();
return ( <div style={{ /* slide styles */ }}> <h1>The Problem</h1> {/* Scene content */}
{/* Audio starts when this scene starts (frame 0 of this sequence) */} <Audio src={staticFile('audio/scenes/scene-02-problem.mp3')} /> </div> ); }; ```
### Syncing Visuals to Voiceover
Calculate scene duration from audio, not the other way around:
```tsx // src/config/timing.ts import manifest from '../../public/audio/scenes/manifest.json';
const FPS = 30;
// Convert audio durations to frame counts export const sceneDurations = manifest.scenes.reduce((acc, scene) => { const name = scene.file.replace(/^scene-\d+-/, '').replace('.mp3', ''); acc[name] = Math.ceil(scene.duration * FPS); return acc; }, {} as Record<string, number>);
// Usage in composition: // <Series.Sequence durationInFrames={sceneDurations.title}> ```
### Audio Timing Patterns
```tsx import { Audio, Sequence, interpolate, useCurrentFrame } from 'remotion';
// Fade in audio export const FadeInAudio: React.FC<{ src: string; fadeFrames?: number }> = ({ src, fadeFrames = 30 }) => { const frame = useCurrentFrame(); const volume = interpolate(frame, [0, fadeFrames], [0, 1], { extrapolateRight: 'clamp', }); return <Audio src={src} volume={volume} />; };
// Delayed audio start export const DelayedAudio: React.FC<{ src: string; delayFrames: number }> = ({ src, delayFrames }) => ( <Sequence from={delayFrames}> <Audio src={src} /> </Sequence> );
// Usage: // <FadeInAudio src={staticFile('audio/music.mp3')} fadeFrames={60} /> // <DelayedAudio src={staticFile('audio/sfx/whoosh.mp3')} delayFrames={45} /> ```
### Voiceover + Demo Video Sync
When a scene has both voiceover and demo video:
```tsx import { Audio, OffthreadVideo, staticFile, useVideoConfig } from 'remotion';
export const DemoScene: React.FC = () => { const { durationInFrames, fps } = useVideoConfig();
// Calculate playback rate to fit demo into voiceover duration const demoDuration = 45; // seconds (original demo length) const sceneDuration = durationInFrames / fps; // seconds (from voiceover) const playbackRate = demoDuration / sceneDuration;
return ( <> <OffthreadVideo src={staticFile('demos/feature-demo.mp4')} playbackRate={playbackRate} /> <Audio src={staticFile('audio/scenes/scene-04-demo.mp3')} /> </> ); }; ```
### Error Handling
```tsx import { Audio, staticFile, delayRender, continueRender } from 'remotion'; import { useEffect, useState } from 'react';
export const SafeAudio: React.FC<{ src: string }> = ({ src }) => { const [handle] = useState(() => delayRender()); const [audioReady, setAudioReady] = useState(false);
useEffect(() => { const audio = new window.Audio(src); audio.oncanplaythrough = () => { setAudioReady(true); continueRender(handle); }; audio.onerror = () => { console.error(`Failed to load audio: ${src}`); continueRender(handle); // Continue without audio rather than hang }; }, [src, handle]);
if (!audioReady) return null; return <Audio src={src} />; }; ```
### Toolkit Command: /generate-voiceover
The `/generate-voiceover` command handles the full workflow:
``` /generate-voiceover
1. Reads VOICEOVER-SCRIPT.md 2. Extracts narration for each scene 3. Generates audio via ElevenLabs API 4. Saves to public/audio/scenes/ 5. Creates manifest.json with durations 6. Updates project.json with timing info ```
## Popular Voices
- George: `JBFqnCBsd6RMkjVDRZzb` (warm narrator) - Rachel: `21m00Tcm4TlvDq8ikWAM` (clear female) - Adam: `pNInz6obpgDQGcFmaJgB` (professional male)
List all: `client.voices.get_all()`
For full API docs, see [reference.md](reference.md).
Source provenance
Decision snapshot
55,318 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for elevenlabs, ready for a manual X post.
elevenlabs: Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating aud... 55.3K stars https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x
Listing + install path for elevenlabs: https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=x Install: npx skills add calesthio/OpenMontage --skill elevenlabs
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to calesthio but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs/audit)
[](https://www.openagentskill.com/skills/calesthio-elevenlabs?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)calesthio
@calesthio
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
174.9K StarsTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
84.9K StarsVox Director
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K StarsCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
174.9K StarsPermission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness