AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Live registry search

Results for “speaker-id

Exact name and slug matches are checked against the live registry before ranked alternatives.

Clear search

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-16 of 17 ranked candidates matching "speaker-id"

Best blend of relevance, quality, freshness, and verified outcomes

1

Speaker

STRONG · 72TRUST · 75SAFE · REVIEWEDCODING

Speaker is a Codex skill project for academic presentations: read real.pptx, combine text extraction, PPTX structure parsing, page-by-page rendering, OCR, and visual rev…

$ npx skills add AI272/speaker
413 stars49 quality75 trustReviewed with permission notesClaude Code + OpenAI Agents2mo since pushNeeds review

Scenario Coding agents

Claude Code + OpenAI Agents · 4 targets

pythoncodex
by AI272DetailsQuick view
2

ChatGPT OpenAI Smart Speaker

NEEDS REVIEW · 45TRUST · 72SAFE · EXPERIMENTALDESIGN

This AI Smart Speaker uses speech recognition, TTS (text-to-speech), and STT (speech-to-text) to enable voice and vision-driven conversations, with additional web search…

$ npx skills add Olney1/ChatGPT-OpenAI-Smart-Speaker
314 stars36 quality72 trustExperimentalOpenAI Agents + LangChain2y since pushNeeds review

Scenario Multimodal media

OpenAI Agents + LangChain · 4 targets

pythonvoice
by Olney1DetailsQuick view
3

Whisper Diarization

VERIFIEDEXCELLENT · 93TRUST · 85SAFE · REVIEWEDDESIGN

Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

$ npx skills add MahmoudAshraf97/whisper-diarization
5.6K stars61 quality85 trustReviewedOpenAI Agents6mo since pushSafe to try

Scenario Multimodal media

OpenAI Agents + CLI · 4 targets

jupyter-notebookspeech
by MahmoudAshraf97DetailsQuick view
4

Izwi

STRONG · 71TRUST · 77SAFE · REVIEWEDDESIGN

Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.

$ npx skills add izwi-ai/izwi
340 stars48 quality77 trustReviewed with permission notesOpenAI Agents2mo since pushNeeds review

Scenario Multimodal media

OpenAI Agents + CLI · 4 targets

rustvoice
by izwi-aiDetailsQuick view
5

Ppt Master

VERIFIEDEXCELLENT · 100TRUST · 89SAFE · REVIEWEDPRESENTATION

AI generates a real, editable PowerPoint from any document — native shapes & animations, speaker notes voiced as audio narration, and the option to follow your own .pptx…

$ npx skills add hugohe3/ppt-master
37.2K stars72 quality89 trustReviewed with permission notes2mo since pushNeeds review

Scenario Presentation generation

CLI + Codex · 4 targets

pythonai-agents
by hugohe3DetailsQuick view
6

Sherpa Onnx

VERIFIEDEXCELLENT · 100TRUST · 90SAFE · VERIFIEDDESIGN

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…

$ npx skills add k2-fsa/sherpa-onnx
13.1K stars69 quality90 trustVerified2mo since pushSafe to try

Scenario Multimodal media

CLI + Codex · 4 targets

c++voice
by k2-fsaDetailsQuick view
7

PaddleSpeech

VERIFIEDEXCELLENT · 100TRUST · 90SAFE · REVIEWEDDESIGN

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…

$ npx skills add PaddlePaddle/PaddleSpeech
12.6K stars69 quality90 trustReviewed2mo since pushSafe to try

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by PaddlePaddleDetailsQuick view
8

MOSS TTS

VERIFIEDEXCELLENT · 100TRUST · 88SAFE · VERIFIEDDESIGN

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…

$ npx skills add OpenMOSS/MOSS-TTS
3.5K stars63 quality88 trustVerified2mo since pushSafe to try

Scenario Multimodal media

CLI + Codex · 4 targets

pythonvoice
by OpenMOSSDetailsQuick view
9

Fun ASR

VERIFIEDEXCELLENT · 91TRUST · 87SAFE · REVIEWEDDESIGN

End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.

$ npx skills add FunAudioLLM/Fun-ASR
1.3K stars61 quality87 trustReviewed2mo since pushSafe to try

Scenario Multimodal media

CLI + Codex · 4 targets

cspeech
by FunAudioLLMDetailsQuick view
10

VibeVoice ComfyUI

VERIFIEDSTRONG · 79TRUST · 84SAFE · REVIEWEDDESIGN

A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your C…

$ npx skills add Enemyx-net/VibeVoice-ComfyUI
1.5K stars53 quality84 trustReviewed with permission notes6mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonvoice
by Enemyx-netDetailsQuick view
11

TranscriptionSuite

STRONG · 72TRUST · 79SAFE · REVIEWEDDATA

A fully local and private Speech-To-Text app with cross-platform support, speaker diarization, Audio Notebook mode, LM Studio integration, and both longform and live tra…

$ npx skills add homelab-00/TranscriptionSuite
519 stars50 quality79 trustReviewed with permission notes2mo since pushNeeds review

Scenario Data analysis

CLI + Codex · 4 targets

typescriptnotebook
by homelab-00DetailsQuick view
12

Slides_maker

STRONG · 77TRUST · 80SAFE · REVIEWEDPRESENTATION

Turn papers, code, and docs into presentation-ready, natively editable PPTX in Codex / Claude Code. Native charts and equations, speaker notes, click-build animations, a…

$ npx skills add addsumtech/slides_maker
379 stars52 quality80 trustReviewedClaude Code + OpenAI Agents15d since pushSafe to try

Scenario Presentation generation

Claude Code + OpenAI Agents · 4 targets

python
by addsumtechDetailsQuick view
13

ComfyUI OmniVoice TTS

STRONG · 72TRUST · 78SAFE · REVIEWEDDESIGN

OmniVoice TTS nodes for ComfyUI - Zero-shot multilingual text-to-speech with voice cloning, voice design, and multi-speaker dialogue

$ npx skills add Saganaki22/ComfyUI-OmniVoice-TTS
432 stars49 quality78 trustReviewed with permission notes2mo since pushNeeds review

Scenario Design and creative

CLI + Codex · 4 targets

pythonvoice
by Saganaki22DetailsQuick view
14

ComfyUI VibeVoice

PROMISING · 61TRUST · 78SAFE · REVIEWEDDESIGN

ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio

$ npx skills add wildminder/ComfyUI-VibeVoice
588 stars42 quality78 trustReviewed with permission notes11mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonvoice
by wildminderDetailsQuick view
15

Real Time Voice Translator

PROMISING · 67TRUST · 77SAFE · REVIEWEDDESIGN

A desktop application that uses AI to translate voice between languages in real time, while preserving the speaker's tone and emotion.

$ npx skills add SamirPaulb/real-time-voice-translator
410 stars45 quality77 trustReviewed with permission notes4mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

tclvoice
by SamirPaulbDetailsQuick view
16

Bluestrike

NEEDS REVIEW · 46TRUST · 67SAFE · EXPERIMENTALCODING

Bluestrike: CLI tool to hack Bluetooth devices through speaker jamming, traffic spoofing & device hijacking (In the making)

$ npx skills add StealthIQ/Bluestrike
417 stars37 quality67 trustExperimental3y since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

pythoniot
by StealthIQDetailsQuick view

Page 1

Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.

Try the agent resolve API