Creator · veniceai
Last updated · Sep 6, 2026
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/a
Creator · veniceai
Last updated · Sep 6, 2026
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/a
Creator · veniceai
Last updated · Sep 6, 2026
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/a
Creator · veniceai
Last updated · Sep 6, 2026
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/a
Sandbox only
Install targets
Codex install prompt
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"veniceai-venice-audio-speech","task":"Install venice-audio-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + OpenAI Agents + Browser agents
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add veniceai/skills --skill venice-audio-speech
Maintenance
fresh
6d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
139
68/100 Quality · 68/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
139 GitHub stars
Repo activity
139 stars, 20 forks
Maintenance
6d since push
License
MIT
Install
npx skills add veniceai/skills --skill venice-audio-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add veniceai/skills --skill venice-audio-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/veniceai-venice-audio-speech/install
Agent should check
Copy prompt
Task: Use venice-audio-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install
Install command: npx skills add veniceai/skills --skill venice-audio-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/veniceai-venice-audio-speech/install
LLM text format
/api/skills/veniceai-venice-audio-speech/install?format=text
Find alternatives
/api/skills/search?q=venice-audio-speech&limit=3
Agent prompt
Use venice-audio-speech for this task. Review https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install, then install with: npx skills add veniceai/skills --skill venice-audio-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/veniceai-venice-audio-speech
LLM text
/api/registry/manifest/veniceai-venice-audio-speech?format=text
Install alias
/api/registry/install/veniceai-venice-audio-speech
Recommend
/api/registry/recommend?task=Use%20venice-audio-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO139 GitHub stars
Stars/forks activity
CHECK139 stars, 20 forks; issue activity unavailable in current metadata
Recent maintenance
PASS6d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: venice-audio-speech description: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. ---
# Venice TTS (`/audio/speech`)
`POST /api/v1/audio/speech` converts text to an audio stream or file. OpenAI-compatible — the OpenAI SDK's `audio.speech.create()` works as a drop-in.
## Use when
- You want narration, voice replies, or UI audio from text. - You need a specific voice family (ElevenLabs, Kokoro, xAI, Qwen 3, Orpheus, Chatterbox, MiniMax, Inworld, Gemini Flash). - You want streaming audio returned sentence-by-sentence. - You need style/emotion control on supported models.
For music generation (lyrics + instrumental), see [`venice-audio-music`](../venice-audio-music/SKILL.md). For transcription (audio → text), see [`venice-audio-transcription`](../venice-audio-transcription/SKILL.md).
## Minimal request
```bash curl https://api.venice.ai/api/v1/audio/speech \ -H "Authorization: Bearer $VENICE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-xai-v1", "voice": "eve", "input": "Hello, welcome to Venice Voice.", "response_format": "mp3", "speed": 1.0, "streaming": false }' --output hello.mp3 ```
Response is the raw audio (`Content-Type` matches `response_format`).
## Request schema
| Field | Type | Default | Notes | |---|---|---|---| | `input` | string | — | **Required.** Up to **4096** characters. | | `model` | enum | `tts-kokoro` (OpenAPI schema default) | See model list below. `tts-xai-v1` is the recommended frontier default; pick the model that fits your voice + language needs. | | `voice` | string, ≤ 512 | model-specific (e.g. `eve` for `tts-xai-v1`) | **Voice is model-specific** — wrong combo = `400`. See voice families. Also accepts a cloned-voice handle (`vv_…`) from `POST /audio/voices`, paired with the same `model` that created it. | | `response_format` | `mp3` / `opus` / `aac` / `flac` / `wav` / `pcm` | `mp3` | `pcm` returns 24 kHz signed-16 LE for pipelines. | | `speed` | number | `1.0` | Range `0.25–4.0`. | | `streaming` | bool | `false` | `true` → streamed sentence-by-sentence as audio continues to generate. | | `language` | string | — | Optional hint. Accepted form depends on model (Qwen 3 = full names like `English`; xAI / ElevenLabs = ISO 639-1 like `en`; MiniMax = full names). Unsupported values silently ignored. | | `prompt` | string, ≤ 500 | — | Emotion / style cue. Only for models with `supportsPromptParam` (Qwen 3 currently). Examples: *"Very happy."*, *"Sad and slow."*. | | `temperature` | 0–2 | — | Sampling temperature. Only for models with `supportsTemperatureParam` (Qwen 3, Orpheus, Chatterbox HD). | | `top_p` | 0–1 | — | Only Qwen 3 currently. |
## Models
| Model ID | Family | Highlights | |---|---|---| | `tts-xai-v1` | xAI | **Recommended default.** Conversational style, ISO 639-1 language hints. | | `tts-kokoro` | Kokoro | OpenAPI schema default. Multilingual, many voices across languages. | | `tts-qwen3-0-6b` / `tts-qwen3-1-7b` | Qwen 3 | Emotion control via `prompt`, temperature, top_p. | | `tts-inworld-1-5-max` | Inworld | Character-driven voices (Craig, Ashley, …). | | `tts-chatterbox-hd` | Chatterbox | HD voices (Aurora, Blade, …), temperature. | | `tts-orpheus` | Orpheus | Conversational (tara, leah, jess, leo, …), temperature. | | `tts-elevenlabs-turbo-v2-5` | ElevenLabs Turbo | Rachel, Aria, Charlotte, Roger, … | | `tts-minimax-speech-02-hd` | MiniMax | WiseWoman, DeepVoiceMan, … Supports **persistent** voice cloning. | | `tts-gemini-3-1-flash` | Gemini Flash | Star-named voices (Achernar, Achird, Zephyr, …). | | `tts-gradium-v1` | Gradium | Multilingual across en/de/es/fr/pt, where the **voice picks the language** (there is no separate language parameter). Proprietary upstream, so `capabilities.private` is `false`. |
Always inspect the entry for your model in `GET /models?type=tts` — `model_spec.voices` is the authoritative voice list. Per-model toggles like `supportsPromptParam`, `supportsTemperatureParam`, `supportsTopPParam` live on the internal model definitions but are not currently exposed on `/models` — treat the request schema below (`instructions`, `temperature`, `top_p`) as the support matrix.
## Voice families (by prefix)
- **Kokoro** — lowercase + language/gender prefix: - `af_*`, `am_*` — American female / male - `bf_*`, `bm_*` — British female / male - `zf_*`, `zm_*` — Chinese - `ff_*`, `hf_*`, `hm_*`, `if_*`, `im_*`, `jf_*`, `jm_*`, `pf_*`, `pm_*`, `ef_*`, `em_*` — French, Hindi, Italian, Japanese, Portuguese, Spanish - Examples: `af_sky`, `af_bella`, `am_adam`, `bm_george`, `zf_xiaoxiao` - **Qwen 3** — `Vivian`, `Serena`, `Ono_Anna`, `Sohee`, `Uncle_Fu`, `Dylan`, `Eric`, `Ryan`, `Aiden` - **xAI** — 26 voices. Original five: `eve`, `ara`, `rex`, `sal`, `leo`. Flagship multilingual set: `altair`, `atlas`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`, `iris`, `kepler`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rigel`, `sirius`, `ursa`, `zagan`, `zenith` - **Orpheus** — `tara`, `leah`, `jess`, `mia`, `zoe`, `dan`, `zac` - **Inworld** — `Craig`, `Ashley`, `Olivia`, `Sarah`, `Elizabeth`, `Priya`, `Alex`, `Edward`, `Theodore`, `Ronald`, `Mark`, `Hades`, `Luna`, `Pixie` - **Chatterbox** — `Aurora`, `Britney`, `Siobhan`, `Vicky`, `Blade`, `Carl`, `Cliff`, `Richard`, `Rico` - **ElevenLabs Turbo** — `Rachel`, `Aria`, `Laura`, `Charlotte`, `Alice`, `Matilda`, `Jessica`, `Lily`, `Roger`, `Charlie`, `George`, `Callum`, `River`, `Liam`, `Will`, `Chris`, `Brian`, `Daniel`, `Bill` - **MiniMax** — `WiseWoman`, `FriendlyPerson`, `InspirationalGirl`, `CalmWoman`, `LivelyGirl`, `LovelyGirl`, `SweetGirl`, `ExuberantGirl`, `DeepVoiceMan`, `CasualGuy`, `PatientMan`, `YoungKnight`, `DeterminedMan`, `ImposingManner`, `ElegantMan` - **Gemini 3 Flash** — star names: `Achernar`, `Achird`, `Algenib`, `Algieba`, `Alnilam`, `Aoede`, `Autonoe`, `Callirrhoe`, `Charon`, `Despina`, `Enceladus`, `Erinome`, `Fenrir`, `Gacrux`, `Iapetus`, `Kore`, `Laomedeia`, `Leda`, `Orus`, `Pulcherrima`, `Puck`, `Rasalgethi`, `Sadachbia`, `Sadaltager`, `Schedar`, `Sulafat`, `Umbriel`, `Vindemiatrix`, `Zephyr`, `Zubenelgenubi` - **Gradium** — the voice selects the language. English: `Emma` (default), `Kent`, `Eva`, `Jack`. German: `Mia`, `Maximilian`. Spanish: `Valentina`, `Sergio`. French: `Elise`, `Leo`. Portuguese: `Alice`, `Davi`
Pass a voice that isn't in the chosen model's list and you get `400`.
## Voice cloning — `POST /audio/voices`
Clone a voice from an audio sample and get back a handle (`vv_…`) you can pass as `voice` on `/audio/speech`. `multipart/form-data` only.
```bash curl https://api.venice.ai/api/v1/audio/voices \ -H "Authorization: Bearer $VENICE_API_KEY" \ -F "model=tts-chatterbox-hd" \ -F "file=@sample.wav" ```
| Field | Notes | |---|---| | `file` | The voice sample, multipart field name `file`. Accepted containers depend on the model. Aim for a clean speech recording of at least 5–10 seconds. | | `model` | `tts-chatterbox-hd` (default) or `tts-minimax-speech-02-hd`. |
| Model | Containers | Persistence | Availability | |---|---|---|---| | `tts-chatterbox-hd` | MP3, WAV, FLAC, M4A | **Zero-shot.** No voice template is derived; the reference audio is stored with a TTL and re-read on every synthesis call. Handles expire after **7 days**, full stop. | Regular users. | | `tts-minimax-speech-02-hd` | MP3, WAV only | **Persistent.** The provider derives a voice template that survives across calls. Auto-deleted after **7 days without use**; each successful TTS request resets the window. | Limited access — contact support@venice.ai to have it enabled. |
A handle is bound to the model that created it. Pass a `vv_…` handle with a different `model` on `/audio/speech` and the call fails.
Samples in a container outside the per-model allowlist are rejected with `400` before anything is uploaded. Beyond the shared TTS error codes, this endpoint also returns `403` and `413` (sample too large).
## Streaming
```json { "model": "tts-xai-v1", "voice": "eve", "input": "Hello, this is a long document to narrate. ...", "streaming": true, "response_format": "mp3" } ```
With `streaming: true`, the HTTP body is a chunked audio stream. Decode as it arrives — useful for latency-sensitive UIs. `response_format: pcm` pairs well with browser Web Audio API for raw playback.
## OpenAI SDK
```ts import OpenAI from 'openai' import fs from 'node:fs/promises'
const client = new OpenAI({ apiKey: process.env.VENICE_API_KEY, baseURL: 'https://api.venice.ai/api/v1', })
const mp3 = await client.audio.speech.create({ model: 'tts-xai-v1', voice: 'eve', input: 'Hello from Venice.', response_format: 'mp3', })
await fs.writeFile('hello.mp3', Buffer.from(await mp3.arrayBuffer())) ```
## Emotion / style (Qwen 3 only)
```json { "model": "tts-qwen3-1-7b", "voice": "Vivian", "input": "We did it!", "prompt": "Excited and energetic.", "temperature": 0.9, "top_p": 0.95 } ```
For other families, emotion comes from the **voice choice itself** (e.g. Inworld `Hades` vs `Pixie`). `prompt` / `temperature` / `top_p` are silently ignored.
## Errors
| Code | Meaning | |---|---| | `400` | Bad voice/model combo, input too long (>4096), language hint rejected by a strict model, invalid voice for the chosen model. | | `401` | Auth / Pro-only model. | | `402` | Insufficient balance. | | `429` | Rate limited. | | `500` / `503` | Inference / capacity issue — retry with jitter. |
## Gotchas
- `input` hard cap is 4096 chars. For books / long content, split on sentence boundaries and concatenate audio client-side. - `streaming: true` + SDKs: some OpenAI SDK versions don't expose streaming for `audio.speech.create`; call the REST endpoint directly and consume the HTTP body. - `speed` compounds with model internal speech rate — extreme values (`0.25`, `4.0`) often sound unnatural; keep within `0.8–1.3` for narration. - Voice names are case-sensitive (`eve` ≠ `EVE`, `af_sky` ≠ `AF_SKY`). - Cloned voices expire. Chatterbox handles die 7 days after creation no matter what; MiniMax handles die after 7 days of no use. Re-clone rather than assuming a handle you stored last month still resolves. - Gradium has no `language` parameter. Pick the voice for the language you want.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for venice-audio-speech, ready for a manual X post.
venice-audio-speech: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. C... 139 stars https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x
Listing + install path for venice-audio-speech: https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x Install: npx skills add veniceai/skills --skill venice-audio-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to veniceai but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech/audit)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)veniceai
@veniceai
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"veniceai-venice-audio-speech","task":"Install venice-audio-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + OpenAI Agents + Browser agents
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add veniceai/skills --skill venice-audio-speech
Maintenance
fresh
6d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
139
68/100 Quality · 68/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
139 GitHub stars
Repo activity
139 stars, 20 forks
Maintenance
6d since push
License
MIT
Install
npx skills add veniceai/skills --skill venice-audio-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add veniceai/skills --skill venice-audio-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/veniceai-venice-audio-speech/install
Agent should check
Copy prompt
Task: Use venice-audio-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install
Install command: npx skills add veniceai/skills --skill venice-audio-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/veniceai-venice-audio-speech/install
LLM text format
/api/skills/veniceai-venice-audio-speech/install?format=text
Find alternatives
/api/skills/search?q=venice-audio-speech&limit=3
Agent prompt
Use venice-audio-speech for this task. Review https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install, then install with: npx skills add veniceai/skills --skill venice-audio-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/veniceai-venice-audio-speech
LLM text
/api/registry/manifest/veniceai-venice-audio-speech?format=text
Install alias
/api/registry/install/veniceai-venice-audio-speech
Recommend
/api/registry/recommend?task=Use%20venice-audio-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO139 GitHub stars
Stars/forks activity
CHECK139 stars, 20 forks; issue activity unavailable in current metadata
Recent maintenance
PASS6d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: venice-audio-speech description: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. ---
# Venice TTS (`/audio/speech`)
`POST /api/v1/audio/speech` converts text to an audio stream or file. OpenAI-compatible — the OpenAI SDK's `audio.speech.create()` works as a drop-in.
## Use when
- You want narration, voice replies, or UI audio from text. - You need a specific voice family (ElevenLabs, Kokoro, xAI, Qwen 3, Orpheus, Chatterbox, MiniMax, Inworld, Gemini Flash). - You want streaming audio returned sentence-by-sentence. - You need style/emotion control on supported models.
For music generation (lyrics + instrumental), see [`venice-audio-music`](../venice-audio-music/SKILL.md). For transcription (audio → text), see [`venice-audio-transcription`](../venice-audio-transcription/SKILL.md).
## Minimal request
```bash curl https://api.venice.ai/api/v1/audio/speech \ -H "Authorization: Bearer $VENICE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-xai-v1", "voice": "eve", "input": "Hello, welcome to Venice Voice.", "response_format": "mp3", "speed": 1.0, "streaming": false }' --output hello.mp3 ```
Response is the raw audio (`Content-Type` matches `response_format`).
## Request schema
| Field | Type | Default | Notes | |---|---|---|---| | `input` | string | — | **Required.** Up to **4096** characters. | | `model` | enum | `tts-kokoro` (OpenAPI schema default) | See model list below. `tts-xai-v1` is the recommended frontier default; pick the model that fits your voice + language needs. | | `voice` | string, ≤ 512 | model-specific (e.g. `eve` for `tts-xai-v1`) | **Voice is model-specific** — wrong combo = `400`. See voice families. Also accepts a cloned-voice handle (`vv_…`) from `POST /audio/voices`, paired with the same `model` that created it. | | `response_format` | `mp3` / `opus` / `aac` / `flac` / `wav` / `pcm` | `mp3` | `pcm` returns 24 kHz signed-16 LE for pipelines. | | `speed` | number | `1.0` | Range `0.25–4.0`. | | `streaming` | bool | `false` | `true` → streamed sentence-by-sentence as audio continues to generate. | | `language` | string | — | Optional hint. Accepted form depends on model (Qwen 3 = full names like `English`; xAI / ElevenLabs = ISO 639-1 like `en`; MiniMax = full names). Unsupported values silently ignored. | | `prompt` | string, ≤ 500 | — | Emotion / style cue. Only for models with `supportsPromptParam` (Qwen 3 currently). Examples: *"Very happy."*, *"Sad and slow."*. | | `temperature` | 0–2 | — | Sampling temperature. Only for models with `supportsTemperatureParam` (Qwen 3, Orpheus, Chatterbox HD). | | `top_p` | 0–1 | — | Only Qwen 3 currently. |
## Models
| Model ID | Family | Highlights | |---|---|---| | `tts-xai-v1` | xAI | **Recommended default.** Conversational style, ISO 639-1 language hints. | | `tts-kokoro` | Kokoro | OpenAPI schema default. Multilingual, many voices across languages. | | `tts-qwen3-0-6b` / `tts-qwen3-1-7b` | Qwen 3 | Emotion control via `prompt`, temperature, top_p. | | `tts-inworld-1-5-max` | Inworld | Character-driven voices (Craig, Ashley, …). | | `tts-chatterbox-hd` | Chatterbox | HD voices (Aurora, Blade, …), temperature. | | `tts-orpheus` | Orpheus | Conversational (tara, leah, jess, leo, …), temperature. | | `tts-elevenlabs-turbo-v2-5` | ElevenLabs Turbo | Rachel, Aria, Charlotte, Roger, … | | `tts-minimax-speech-02-hd` | MiniMax | WiseWoman, DeepVoiceMan, … Supports **persistent** voice cloning. | | `tts-gemini-3-1-flash` | Gemini Flash | Star-named voices (Achernar, Achird, Zephyr, …). | | `tts-gradium-v1` | Gradium | Multilingual across en/de/es/fr/pt, where the **voice picks the language** (there is no separate language parameter). Proprietary upstream, so `capabilities.private` is `false`. |
Always inspect the entry for your model in `GET /models?type=tts` — `model_spec.voices` is the authoritative voice list. Per-model toggles like `supportsPromptParam`, `supportsTemperatureParam`, `supportsTopPParam` live on the internal model definitions but are not currently exposed on `/models` — treat the request schema below (`instructions`, `temperature`, `top_p`) as the support matrix.
## Voice families (by prefix)
- **Kokoro** — lowercase + language/gender prefix: - `af_*`, `am_*` — American female / male - `bf_*`, `bm_*` — British female / male - `zf_*`, `zm_*` — Chinese - `ff_*`, `hf_*`, `hm_*`, `if_*`, `im_*`, `jf_*`, `jm_*`, `pf_*`, `pm_*`, `ef_*`, `em_*` — French, Hindi, Italian, Japanese, Portuguese, Spanish - Examples: `af_sky`, `af_bella`, `am_adam`, `bm_george`, `zf_xiaoxiao` - **Qwen 3** — `Vivian`, `Serena`, `Ono_Anna`, `Sohee`, `Uncle_Fu`, `Dylan`, `Eric`, `Ryan`, `Aiden` - **xAI** — 26 voices. Original five: `eve`, `ara`, `rex`, `sal`, `leo`. Flagship multilingual set: `altair`, `atlas`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`, `iris`, `kepler`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rigel`, `sirius`, `ursa`, `zagan`, `zenith` - **Orpheus** — `tara`, `leah`, `jess`, `mia`, `zoe`, `dan`, `zac` - **Inworld** — `Craig`, `Ashley`, `Olivia`, `Sarah`, `Elizabeth`, `Priya`, `Alex`, `Edward`, `Theodore`, `Ronald`, `Mark`, `Hades`, `Luna`, `Pixie` - **Chatterbox** — `Aurora`, `Britney`, `Siobhan`, `Vicky`, `Blade`, `Carl`, `Cliff`, `Richard`, `Rico` - **ElevenLabs Turbo** — `Rachel`, `Aria`, `Laura`, `Charlotte`, `Alice`, `Matilda`, `Jessica`, `Lily`, `Roger`, `Charlie`, `George`, `Callum`, `River`, `Liam`, `Will`, `Chris`, `Brian`, `Daniel`, `Bill` - **MiniMax** — `WiseWoman`, `FriendlyPerson`, `InspirationalGirl`, `CalmWoman`, `LivelyGirl`, `LovelyGirl`, `SweetGirl`, `ExuberantGirl`, `DeepVoiceMan`, `CasualGuy`, `PatientMan`, `YoungKnight`, `DeterminedMan`, `ImposingManner`, `ElegantMan` - **Gemini 3 Flash** — star names: `Achernar`, `Achird`, `Algenib`, `Algieba`, `Alnilam`, `Aoede`, `Autonoe`, `Callirrhoe`, `Charon`, `Despina`, `Enceladus`, `Erinome`, `Fenrir`, `Gacrux`, `Iapetus`, `Kore`, `Laomedeia`, `Leda`, `Orus`, `Pulcherrima`, `Puck`, `Rasalgethi`, `Sadachbia`, `Sadaltager`, `Schedar`, `Sulafat`, `Umbriel`, `Vindemiatrix`, `Zephyr`, `Zubenelgenubi` - **Gradium** — the voice selects the language. English: `Emma` (default), `Kent`, `Eva`, `Jack`. German: `Mia`, `Maximilian`. Spanish: `Valentina`, `Sergio`. French: `Elise`, `Leo`. Portuguese: `Alice`, `Davi`
Pass a voice that isn't in the chosen model's list and you get `400`.
## Voice cloning — `POST /audio/voices`
Clone a voice from an audio sample and get back a handle (`vv_…`) you can pass as `voice` on `/audio/speech`. `multipart/form-data` only.
```bash curl https://api.venice.ai/api/v1/audio/voices \ -H "Authorization: Bearer $VENICE_API_KEY" \ -F "model=tts-chatterbox-hd" \ -F "file=@sample.wav" ```
| Field | Notes | |---|---| | `file` | The voice sample, multipart field name `file`. Accepted containers depend on the model. Aim for a clean speech recording of at least 5–10 seconds. | | `model` | `tts-chatterbox-hd` (default) or `tts-minimax-speech-02-hd`. |
| Model | Containers | Persistence | Availability | |---|---|---|---| | `tts-chatterbox-hd` | MP3, WAV, FLAC, M4A | **Zero-shot.** No voice template is derived; the reference audio is stored with a TTL and re-read on every synthesis call. Handles expire after **7 days**, full stop. | Regular users. | | `tts-minimax-speech-02-hd` | MP3, WAV only | **Persistent.** The provider derives a voice template that survives across calls. Auto-deleted after **7 days without use**; each successful TTS request resets the window. | Limited access — contact support@venice.ai to have it enabled. |
A handle is bound to the model that created it. Pass a `vv_…` handle with a different `model` on `/audio/speech` and the call fails.
Samples in a container outside the per-model allowlist are rejected with `400` before anything is uploaded. Beyond the shared TTS error codes, this endpoint also returns `403` and `413` (sample too large).
## Streaming
```json { "model": "tts-xai-v1", "voice": "eve", "input": "Hello, this is a long document to narrate. ...", "streaming": true, "response_format": "mp3" } ```
With `streaming: true`, the HTTP body is a chunked audio stream. Decode as it arrives — useful for latency-sensitive UIs. `response_format: pcm` pairs well with browser Web Audio API for raw playback.
## OpenAI SDK
```ts import OpenAI from 'openai' import fs from 'node:fs/promises'
const client = new OpenAI({ apiKey: process.env.VENICE_API_KEY, baseURL: 'https://api.venice.ai/api/v1', })
const mp3 = await client.audio.speech.create({ model: 'tts-xai-v1', voice: 'eve', input: 'Hello from Venice.', response_format: 'mp3', })
await fs.writeFile('hello.mp3', Buffer.from(await mp3.arrayBuffer())) ```
## Emotion / style (Qwen 3 only)
```json { "model": "tts-qwen3-1-7b", "voice": "Vivian", "input": "We did it!", "prompt": "Excited and energetic.", "temperature": 0.9, "top_p": 0.95 } ```
For other families, emotion comes from the **voice choice itself** (e.g. Inworld `Hades` vs `Pixie`). `prompt` / `temperature` / `top_p` are silently ignored.
## Errors
| Code | Meaning | |---|---| | `400` | Bad voice/model combo, input too long (>4096), language hint rejected by a strict model, invalid voice for the chosen model. | | `401` | Auth / Pro-only model. | | `402` | Insufficient balance. | | `429` | Rate limited. | | `500` / `503` | Inference / capacity issue — retry with jitter. |
## Gotchas
- `input` hard cap is 4096 chars. For books / long content, split on sentence boundaries and concatenate audio client-side. - `streaming: true` + SDKs: some OpenAI SDK versions don't expose streaming for `audio.speech.create`; call the REST endpoint directly and consume the HTTP body. - `speed` compounds with model internal speech rate — extreme values (`0.25`, `4.0`) often sound unnatural; keep within `0.8–1.3` for narration. - Voice names are case-sensitive (`eve` ≠ `EVE`, `af_sky` ≠ `AF_SKY`). - Cloned voices expire. Chatterbox handles die 7 days after creation no matter what; MiniMax handles die after 7 days of no use. Re-clone rather than assuming a handle you stored last month still resolves. - Gradium has no `language` parameter. Pick the voice for the language you want.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for venice-audio-speech, ready for a manual X post.
venice-audio-speech: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. C... 139 stars https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x
Listing + install path for venice-audio-speech: https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x Install: npx skills add veniceai/skills --skill venice-audio-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to veniceai but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech/audit)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)veniceai
@veniceai
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"veniceai-venice-audio-speech","task":"Install venice-audio-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + OpenAI Agents + Browser agents
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add veniceai/skills --skill venice-audio-speech
Maintenance
fresh
6d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
139
68/100 Quality · 68/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
139 GitHub stars
Repo activity
139 stars, 20 forks
Maintenance
6d since push
License
MIT
Install
npx skills add veniceai/skills --skill venice-audio-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add veniceai/skills --skill venice-audio-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/veniceai-venice-audio-speech/install
Agent should check
Copy prompt
Task: Use venice-audio-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install
Install command: npx skills add veniceai/skills --skill venice-audio-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/veniceai-venice-audio-speech/install
LLM text format
/api/skills/veniceai-venice-audio-speech/install?format=text
Find alternatives
/api/skills/search?q=venice-audio-speech&limit=3
Agent prompt
Use venice-audio-speech for this task. Review https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install, then install with: npx skills add veniceai/skills --skill venice-audio-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/veniceai-venice-audio-speech
LLM text
/api/registry/manifest/veniceai-venice-audio-speech?format=text
Install alias
/api/registry/install/veniceai-venice-audio-speech
Recommend
/api/registry/recommend?task=Use%20venice-audio-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO139 GitHub stars
Stars/forks activity
CHECK139 stars, 20 forks; issue activity unavailable in current metadata
Recent maintenance
PASS6d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: venice-audio-speech description: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. ---
# Venice TTS (`/audio/speech`)
`POST /api/v1/audio/speech` converts text to an audio stream or file. OpenAI-compatible — the OpenAI SDK's `audio.speech.create()` works as a drop-in.
## Use when
- You want narration, voice replies, or UI audio from text. - You need a specific voice family (ElevenLabs, Kokoro, xAI, Qwen 3, Orpheus, Chatterbox, MiniMax, Inworld, Gemini Flash). - You want streaming audio returned sentence-by-sentence. - You need style/emotion control on supported models.
For music generation (lyrics + instrumental), see [`venice-audio-music`](../venice-audio-music/SKILL.md). For transcription (audio → text), see [`venice-audio-transcription`](../venice-audio-transcription/SKILL.md).
## Minimal request
```bash curl https://api.venice.ai/api/v1/audio/speech \ -H "Authorization: Bearer $VENICE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-xai-v1", "voice": "eve", "input": "Hello, welcome to Venice Voice.", "response_format": "mp3", "speed": 1.0, "streaming": false }' --output hello.mp3 ```
Response is the raw audio (`Content-Type` matches `response_format`).
## Request schema
| Field | Type | Default | Notes | |---|---|---|---| | `input` | string | — | **Required.** Up to **4096** characters. | | `model` | enum | `tts-kokoro` (OpenAPI schema default) | See model list below. `tts-xai-v1` is the recommended frontier default; pick the model that fits your voice + language needs. | | `voice` | string, ≤ 512 | model-specific (e.g. `eve` for `tts-xai-v1`) | **Voice is model-specific** — wrong combo = `400`. See voice families. Also accepts a cloned-voice handle (`vv_…`) from `POST /audio/voices`, paired with the same `model` that created it. | | `response_format` | `mp3` / `opus` / `aac` / `flac` / `wav` / `pcm` | `mp3` | `pcm` returns 24 kHz signed-16 LE for pipelines. | | `speed` | number | `1.0` | Range `0.25–4.0`. | | `streaming` | bool | `false` | `true` → streamed sentence-by-sentence as audio continues to generate. | | `language` | string | — | Optional hint. Accepted form depends on model (Qwen 3 = full names like `English`; xAI / ElevenLabs = ISO 639-1 like `en`; MiniMax = full names). Unsupported values silently ignored. | | `prompt` | string, ≤ 500 | — | Emotion / style cue. Only for models with `supportsPromptParam` (Qwen 3 currently). Examples: *"Very happy."*, *"Sad and slow."*. | | `temperature` | 0–2 | — | Sampling temperature. Only for models with `supportsTemperatureParam` (Qwen 3, Orpheus, Chatterbox HD). | | `top_p` | 0–1 | — | Only Qwen 3 currently. |
## Models
| Model ID | Family | Highlights | |---|---|---| | `tts-xai-v1` | xAI | **Recommended default.** Conversational style, ISO 639-1 language hints. | | `tts-kokoro` | Kokoro | OpenAPI schema default. Multilingual, many voices across languages. | | `tts-qwen3-0-6b` / `tts-qwen3-1-7b` | Qwen 3 | Emotion control via `prompt`, temperature, top_p. | | `tts-inworld-1-5-max` | Inworld | Character-driven voices (Craig, Ashley, …). | | `tts-chatterbox-hd` | Chatterbox | HD voices (Aurora, Blade, …), temperature. | | `tts-orpheus` | Orpheus | Conversational (tara, leah, jess, leo, …), temperature. | | `tts-elevenlabs-turbo-v2-5` | ElevenLabs Turbo | Rachel, Aria, Charlotte, Roger, … | | `tts-minimax-speech-02-hd` | MiniMax | WiseWoman, DeepVoiceMan, … Supports **persistent** voice cloning. | | `tts-gemini-3-1-flash` | Gemini Flash | Star-named voices (Achernar, Achird, Zephyr, …). | | `tts-gradium-v1` | Gradium | Multilingual across en/de/es/fr/pt, where the **voice picks the language** (there is no separate language parameter). Proprietary upstream, so `capabilities.private` is `false`. |
Always inspect the entry for your model in `GET /models?type=tts` — `model_spec.voices` is the authoritative voice list. Per-model toggles like `supportsPromptParam`, `supportsTemperatureParam`, `supportsTopPParam` live on the internal model definitions but are not currently exposed on `/models` — treat the request schema below (`instructions`, `temperature`, `top_p`) as the support matrix.
## Voice families (by prefix)
- **Kokoro** — lowercase + language/gender prefix: - `af_*`, `am_*` — American female / male - `bf_*`, `bm_*` — British female / male - `zf_*`, `zm_*` — Chinese - `ff_*`, `hf_*`, `hm_*`, `if_*`, `im_*`, `jf_*`, `jm_*`, `pf_*`, `pm_*`, `ef_*`, `em_*` — French, Hindi, Italian, Japanese, Portuguese, Spanish - Examples: `af_sky`, `af_bella`, `am_adam`, `bm_george`, `zf_xiaoxiao` - **Qwen 3** — `Vivian`, `Serena`, `Ono_Anna`, `Sohee`, `Uncle_Fu`, `Dylan`, `Eric`, `Ryan`, `Aiden` - **xAI** — 26 voices. Original five: `eve`, `ara`, `rex`, `sal`, `leo`. Flagship multilingual set: `altair`, `atlas`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`, `iris`, `kepler`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rigel`, `sirius`, `ursa`, `zagan`, `zenith` - **Orpheus** — `tara`, `leah`, `jess`, `mia`, `zoe`, `dan`, `zac` - **Inworld** — `Craig`, `Ashley`, `Olivia`, `Sarah`, `Elizabeth`, `Priya`, `Alex`, `Edward`, `Theodore`, `Ronald`, `Mark`, `Hades`, `Luna`, `Pixie` - **Chatterbox** — `Aurora`, `Britney`, `Siobhan`, `Vicky`, `Blade`, `Carl`, `Cliff`, `Richard`, `Rico` - **ElevenLabs Turbo** — `Rachel`, `Aria`, `Laura`, `Charlotte`, `Alice`, `Matilda`, `Jessica`, `Lily`, `Roger`, `Charlie`, `George`, `Callum`, `River`, `Liam`, `Will`, `Chris`, `Brian`, `Daniel`, `Bill` - **MiniMax** — `WiseWoman`, `FriendlyPerson`, `InspirationalGirl`, `CalmWoman`, `LivelyGirl`, `LovelyGirl`, `SweetGirl`, `ExuberantGirl`, `DeepVoiceMan`, `CasualGuy`, `PatientMan`, `YoungKnight`, `DeterminedMan`, `ImposingManner`, `ElegantMan` - **Gemini 3 Flash** — star names: `Achernar`, `Achird`, `Algenib`, `Algieba`, `Alnilam`, `Aoede`, `Autonoe`, `Callirrhoe`, `Charon`, `Despina`, `Enceladus`, `Erinome`, `Fenrir`, `Gacrux`, `Iapetus`, `Kore`, `Laomedeia`, `Leda`, `Orus`, `Pulcherrima`, `Puck`, `Rasalgethi`, `Sadachbia`, `Sadaltager`, `Schedar`, `Sulafat`, `Umbriel`, `Vindemiatrix`, `Zephyr`, `Zubenelgenubi` - **Gradium** — the voice selects the language. English: `Emma` (default), `Kent`, `Eva`, `Jack`. German: `Mia`, `Maximilian`. Spanish: `Valentina`, `Sergio`. French: `Elise`, `Leo`. Portuguese: `Alice`, `Davi`
Pass a voice that isn't in the chosen model's list and you get `400`.
## Voice cloning — `POST /audio/voices`
Clone a voice from an audio sample and get back a handle (`vv_…`) you can pass as `voice` on `/audio/speech`. `multipart/form-data` only.
```bash curl https://api.venice.ai/api/v1/audio/voices \ -H "Authorization: Bearer $VENICE_API_KEY" \ -F "model=tts-chatterbox-hd" \ -F "file=@sample.wav" ```
| Field | Notes | |---|---| | `file` | The voice sample, multipart field name `file`. Accepted containers depend on the model. Aim for a clean speech recording of at least 5–10 seconds. | | `model` | `tts-chatterbox-hd` (default) or `tts-minimax-speech-02-hd`. |
| Model | Containers | Persistence | Availability | |---|---|---|---| | `tts-chatterbox-hd` | MP3, WAV, FLAC, M4A | **Zero-shot.** No voice template is derived; the reference audio is stored with a TTL and re-read on every synthesis call. Handles expire after **7 days**, full stop. | Regular users. | | `tts-minimax-speech-02-hd` | MP3, WAV only | **Persistent.** The provider derives a voice template that survives across calls. Auto-deleted after **7 days without use**; each successful TTS request resets the window. | Limited access — contact support@venice.ai to have it enabled. |
A handle is bound to the model that created it. Pass a `vv_…` handle with a different `model` on `/audio/speech` and the call fails.
Samples in a container outside the per-model allowlist are rejected with `400` before anything is uploaded. Beyond the shared TTS error codes, this endpoint also returns `403` and `413` (sample too large).
## Streaming
```json { "model": "tts-xai-v1", "voice": "eve", "input": "Hello, this is a long document to narrate. ...", "streaming": true, "response_format": "mp3" } ```
With `streaming: true`, the HTTP body is a chunked audio stream. Decode as it arrives — useful for latency-sensitive UIs. `response_format: pcm` pairs well with browser Web Audio API for raw playback.
## OpenAI SDK
```ts import OpenAI from 'openai' import fs from 'node:fs/promises'
const client = new OpenAI({ apiKey: process.env.VENICE_API_KEY, baseURL: 'https://api.venice.ai/api/v1', })
const mp3 = await client.audio.speech.create({ model: 'tts-xai-v1', voice: 'eve', input: 'Hello from Venice.', response_format: 'mp3', })
await fs.writeFile('hello.mp3', Buffer.from(await mp3.arrayBuffer())) ```
## Emotion / style (Qwen 3 only)
```json { "model": "tts-qwen3-1-7b", "voice": "Vivian", "input": "We did it!", "prompt": "Excited and energetic.", "temperature": 0.9, "top_p": 0.95 } ```
For other families, emotion comes from the **voice choice itself** (e.g. Inworld `Hades` vs `Pixie`). `prompt` / `temperature` / `top_p` are silently ignored.
## Errors
| Code | Meaning | |---|---| | `400` | Bad voice/model combo, input too long (>4096), language hint rejected by a strict model, invalid voice for the chosen model. | | `401` | Auth / Pro-only model. | | `402` | Insufficient balance. | | `429` | Rate limited. | | `500` / `503` | Inference / capacity issue — retry with jitter. |
## Gotchas
- `input` hard cap is 4096 chars. For books / long content, split on sentence boundaries and concatenate audio client-side. - `streaming: true` + SDKs: some OpenAI SDK versions don't expose streaming for `audio.speech.create`; call the REST endpoint directly and consume the HTTP body. - `speed` compounds with model internal speech rate — extreme values (`0.25`, `4.0`) often sound unnatural; keep within `0.8–1.3` for narration. - Voice names are case-sensitive (`eve` ≠ `EVE`, `af_sky` ≠ `AF_SKY`). - Cloned voices expire. Chatterbox handles die 7 days after creation no matter what; MiniMax handles die after 7 days of no use. Re-clone rather than assuming a handle you stored last month still resolves. - Gradium has no `language` parameter. Pick the voice for the language you want.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for venice-audio-speech, ready for a manual X post.
venice-audio-speech: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. C... 139 stars https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x
Listing + install path for venice-audio-speech: https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x Install: npx skills add veniceai/skills --skill venice-audio-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to veniceai but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech/audit)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)veniceai
@veniceai
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsSandbox only
Install targets
Codex install prompt
Install the "venice-audio-speech" agent skill from https://github.com/veniceai/skills/tree/main/skills/venice-audio-speech. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"veniceai-venice-audio-speech","task":"Install venice-audio-speech","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + OpenAI Agents + Browser agents
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add veniceai/skills --skill venice-audio-speech
Maintenance
fresh
6d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
139
68/100 Quality · 68/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
139 GitHub stars
Repo activity
139 stars, 20 forks
Maintenance
6d since push
License
MIT
Install
npx skills add veniceai/skills --skill venice-audio-speech
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add veniceai/skills --skill venice-audio-speechDo not use when
Agent safety v2
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/veniceai-venice-audio-speech/install
Agent should check
Copy prompt
Task: Use venice-audio-speech in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20venice-audio-speech%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install
Install command: npx skills add veniceai/skills --skill venice-audio-speech
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/veniceai-venice-audio-speech/install
LLM text format
/api/skills/veniceai-venice-audio-speech/install?format=text
Find alternatives
/api/skills/search?q=venice-audio-speech&limit=3
Agent prompt
Use venice-audio-speech for this task. Review https://www.openagentskill.com/api/skills/veniceai-venice-audio-speech/install, then install with: npx skills add veniceai/skills --skill venice-audio-speechRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/veniceai-venice-audio-speech
LLM text
/api/registry/manifest/veniceai-venice-audio-speech?format=text
Install alias
/api/registry/install/veniceai-venice-audio-speech
Recommend
/api/registry/recommend?task=Use%20venice-audio-speech%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents, Browser agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO139 GitHub stars
Stars/forks activity
CHECK139 stars, 20 forks; issue activity unavailable in current metadata
Recent maintenance
PASS6d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Scrape, clean, and reuse web data
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: venice-audio-speech description: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per family, cloned-voice handles, output formats (mp3/opus/aac/flac/wav/pcm), streaming, prompt/emotion styling, temperature/top_p, and language hints. ---
# Venice TTS (`/audio/speech`)
`POST /api/v1/audio/speech` converts text to an audio stream or file. OpenAI-compatible — the OpenAI SDK's `audio.speech.create()` works as a drop-in.
## Use when
- You want narration, voice replies, or UI audio from text. - You need a specific voice family (ElevenLabs, Kokoro, xAI, Qwen 3, Orpheus, Chatterbox, MiniMax, Inworld, Gemini Flash). - You want streaming audio returned sentence-by-sentence. - You need style/emotion control on supported models.
For music generation (lyrics + instrumental), see [`venice-audio-music`](../venice-audio-music/SKILL.md). For transcription (audio → text), see [`venice-audio-transcription`](../venice-audio-transcription/SKILL.md).
## Minimal request
```bash curl https://api.venice.ai/api/v1/audio/speech \ -H "Authorization: Bearer $VENICE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tts-xai-v1", "voice": "eve", "input": "Hello, welcome to Venice Voice.", "response_format": "mp3", "speed": 1.0, "streaming": false }' --output hello.mp3 ```
Response is the raw audio (`Content-Type` matches `response_format`).
## Request schema
| Field | Type | Default | Notes | |---|---|---|---| | `input` | string | — | **Required.** Up to **4096** characters. | | `model` | enum | `tts-kokoro` (OpenAPI schema default) | See model list below. `tts-xai-v1` is the recommended frontier default; pick the model that fits your voice + language needs. | | `voice` | string, ≤ 512 | model-specific (e.g. `eve` for `tts-xai-v1`) | **Voice is model-specific** — wrong combo = `400`. See voice families. Also accepts a cloned-voice handle (`vv_…`) from `POST /audio/voices`, paired with the same `model` that created it. | | `response_format` | `mp3` / `opus` / `aac` / `flac` / `wav` / `pcm` | `mp3` | `pcm` returns 24 kHz signed-16 LE for pipelines. | | `speed` | number | `1.0` | Range `0.25–4.0`. | | `streaming` | bool | `false` | `true` → streamed sentence-by-sentence as audio continues to generate. | | `language` | string | — | Optional hint. Accepted form depends on model (Qwen 3 = full names like `English`; xAI / ElevenLabs = ISO 639-1 like `en`; MiniMax = full names). Unsupported values silently ignored. | | `prompt` | string, ≤ 500 | — | Emotion / style cue. Only for models with `supportsPromptParam` (Qwen 3 currently). Examples: *"Very happy."*, *"Sad and slow."*. | | `temperature` | 0–2 | — | Sampling temperature. Only for models with `supportsTemperatureParam` (Qwen 3, Orpheus, Chatterbox HD). | | `top_p` | 0–1 | — | Only Qwen 3 currently. |
## Models
| Model ID | Family | Highlights | |---|---|---| | `tts-xai-v1` | xAI | **Recommended default.** Conversational style, ISO 639-1 language hints. | | `tts-kokoro` | Kokoro | OpenAPI schema default. Multilingual, many voices across languages. | | `tts-qwen3-0-6b` / `tts-qwen3-1-7b` | Qwen 3 | Emotion control via `prompt`, temperature, top_p. | | `tts-inworld-1-5-max` | Inworld | Character-driven voices (Craig, Ashley, …). | | `tts-chatterbox-hd` | Chatterbox | HD voices (Aurora, Blade, …), temperature. | | `tts-orpheus` | Orpheus | Conversational (tara, leah, jess, leo, …), temperature. | | `tts-elevenlabs-turbo-v2-5` | ElevenLabs Turbo | Rachel, Aria, Charlotte, Roger, … | | `tts-minimax-speech-02-hd` | MiniMax | WiseWoman, DeepVoiceMan, … Supports **persistent** voice cloning. | | `tts-gemini-3-1-flash` | Gemini Flash | Star-named voices (Achernar, Achird, Zephyr, …). | | `tts-gradium-v1` | Gradium | Multilingual across en/de/es/fr/pt, where the **voice picks the language** (there is no separate language parameter). Proprietary upstream, so `capabilities.private` is `false`. |
Always inspect the entry for your model in `GET /models?type=tts` — `model_spec.voices` is the authoritative voice list. Per-model toggles like `supportsPromptParam`, `supportsTemperatureParam`, `supportsTopPParam` live on the internal model definitions but are not currently exposed on `/models` — treat the request schema below (`instructions`, `temperature`, `top_p`) as the support matrix.
## Voice families (by prefix)
- **Kokoro** — lowercase + language/gender prefix: - `af_*`, `am_*` — American female / male - `bf_*`, `bm_*` — British female / male - `zf_*`, `zm_*` — Chinese - `ff_*`, `hf_*`, `hm_*`, `if_*`, `im_*`, `jf_*`, `jm_*`, `pf_*`, `pm_*`, `ef_*`, `em_*` — French, Hindi, Italian, Japanese, Portuguese, Spanish - Examples: `af_sky`, `af_bella`, `am_adam`, `bm_george`, `zf_xiaoxiao` - **Qwen 3** — `Vivian`, `Serena`, `Ono_Anna`, `Sohee`, `Uncle_Fu`, `Dylan`, `Eric`, `Ryan`, `Aiden` - **xAI** — 26 voices. Original five: `eve`, `ara`, `rex`, `sal`, `leo`. Flagship multilingual set: `altair`, `atlas`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`, `iris`, `kepler`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rigel`, `sirius`, `ursa`, `zagan`, `zenith` - **Orpheus** — `tara`, `leah`, `jess`, `mia`, `zoe`, `dan`, `zac` - **Inworld** — `Craig`, `Ashley`, `Olivia`, `Sarah`, `Elizabeth`, `Priya`, `Alex`, `Edward`, `Theodore`, `Ronald`, `Mark`, `Hades`, `Luna`, `Pixie` - **Chatterbox** — `Aurora`, `Britney`, `Siobhan`, `Vicky`, `Blade`, `Carl`, `Cliff`, `Richard`, `Rico` - **ElevenLabs Turbo** — `Rachel`, `Aria`, `Laura`, `Charlotte`, `Alice`, `Matilda`, `Jessica`, `Lily`, `Roger`, `Charlie`, `George`, `Callum`, `River`, `Liam`, `Will`, `Chris`, `Brian`, `Daniel`, `Bill` - **MiniMax** — `WiseWoman`, `FriendlyPerson`, `InspirationalGirl`, `CalmWoman`, `LivelyGirl`, `LovelyGirl`, `SweetGirl`, `ExuberantGirl`, `DeepVoiceMan`, `CasualGuy`, `PatientMan`, `YoungKnight`, `DeterminedMan`, `ImposingManner`, `ElegantMan` - **Gemini 3 Flash** — star names: `Achernar`, `Achird`, `Algenib`, `Algieba`, `Alnilam`, `Aoede`, `Autonoe`, `Callirrhoe`, `Charon`, `Despina`, `Enceladus`, `Erinome`, `Fenrir`, `Gacrux`, `Iapetus`, `Kore`, `Laomedeia`, `Leda`, `Orus`, `Pulcherrima`, `Puck`, `Rasalgethi`, `Sadachbia`, `Sadaltager`, `Schedar`, `Sulafat`, `Umbriel`, `Vindemiatrix`, `Zephyr`, `Zubenelgenubi` - **Gradium** — the voice selects the language. English: `Emma` (default), `Kent`, `Eva`, `Jack`. German: `Mia`, `Maximilian`. Spanish: `Valentina`, `Sergio`. French: `Elise`, `Leo`. Portuguese: `Alice`, `Davi`
Pass a voice that isn't in the chosen model's list and you get `400`.
## Voice cloning — `POST /audio/voices`
Clone a voice from an audio sample and get back a handle (`vv_…`) you can pass as `voice` on `/audio/speech`. `multipart/form-data` only.
```bash curl https://api.venice.ai/api/v1/audio/voices \ -H "Authorization: Bearer $VENICE_API_KEY" \ -F "model=tts-chatterbox-hd" \ -F "file=@sample.wav" ```
| Field | Notes | |---|---| | `file` | The voice sample, multipart field name `file`. Accepted containers depend on the model. Aim for a clean speech recording of at least 5–10 seconds. | | `model` | `tts-chatterbox-hd` (default) or `tts-minimax-speech-02-hd`. |
| Model | Containers | Persistence | Availability | |---|---|---|---| | `tts-chatterbox-hd` | MP3, WAV, FLAC, M4A | **Zero-shot.** No voice template is derived; the reference audio is stored with a TTL and re-read on every synthesis call. Handles expire after **7 days**, full stop. | Regular users. | | `tts-minimax-speech-02-hd` | MP3, WAV only | **Persistent.** The provider derives a voice template that survives across calls. Auto-deleted after **7 days without use**; each successful TTS request resets the window. | Limited access — contact support@venice.ai to have it enabled. |
A handle is bound to the model that created it. Pass a `vv_…` handle with a different `model` on `/audio/speech` and the call fails.
Samples in a container outside the per-model allowlist are rejected with `400` before anything is uploaded. Beyond the shared TTS error codes, this endpoint also returns `403` and `413` (sample too large).
## Streaming
```json { "model": "tts-xai-v1", "voice": "eve", "input": "Hello, this is a long document to narrate. ...", "streaming": true, "response_format": "mp3" } ```
With `streaming: true`, the HTTP body is a chunked audio stream. Decode as it arrives — useful for latency-sensitive UIs. `response_format: pcm` pairs well with browser Web Audio API for raw playback.
## OpenAI SDK
```ts import OpenAI from 'openai' import fs from 'node:fs/promises'
const client = new OpenAI({ apiKey: process.env.VENICE_API_KEY, baseURL: 'https://api.venice.ai/api/v1', })
const mp3 = await client.audio.speech.create({ model: 'tts-xai-v1', voice: 'eve', input: 'Hello from Venice.', response_format: 'mp3', })
await fs.writeFile('hello.mp3', Buffer.from(await mp3.arrayBuffer())) ```
## Emotion / style (Qwen 3 only)
```json { "model": "tts-qwen3-1-7b", "voice": "Vivian", "input": "We did it!", "prompt": "Excited and energetic.", "temperature": 0.9, "top_p": 0.95 } ```
For other families, emotion comes from the **voice choice itself** (e.g. Inworld `Hades` vs `Pixie`). `prompt` / `temperature` / `top_p` are silently ignored.
## Errors
| Code | Meaning | |---|---| | `400` | Bad voice/model combo, input too long (>4096), language hint rejected by a strict model, invalid voice for the chosen model. | | `401` | Auth / Pro-only model. | | `402` | Insufficient balance. | | `429` | Rate limited. | | `500` / `503` | Inference / capacity issue — retry with jitter. |
## Gotchas
- `input` hard cap is 4096 chars. For books / long content, split on sentence boundaries and concatenate audio client-side. - `streaming: true` + SDKs: some OpenAI SDK versions don't expose streaming for `audio.speech.create`; call the REST endpoint directly and consume the HTTP body. - `speed` compounds with model internal speech rate — extreme values (`0.25`, `4.0`) often sound unnatural; keep within `0.8–1.3` for narration. - Voice names are case-sensitive (`eve` ≠ `EVE`, `af_sky` ≠ `AF_SKY`). - Cloned voices expire. Chatterbox handles die 7 days after creation no matter what; MiniMax handles die after 7 days of no use. Re-clone rather than assuming a handle you stored last month still resolves. - Gradium has no `language` parameter. Pick the voice for the language you want.
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for venice-audio-speech, ready for a manual X post.
venice-audio-speech: Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. C... 139 stars https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x
Listing + install path for venice-audio-speech: https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=x Install: npx skills add veniceai/skills --skill venice-audio-speech
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to veniceai but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech/audit)
[](https://www.openagentskill.com/skills/veniceai-venice-audio-speech?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)veniceai
@veniceai
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsPermission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness