AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Browse by scenario

Real skills, grouped by the work your agent needs to finish.

Each directory entry links to real skill pages with GitHub adoption, trust score, install handoff, risk notes, and agent-readable metadata. Use these as starting points when you want a shortlist before asking an agent to install anything.

Design to deployment

Frontend and UI skills

Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.

  • :speech_balloon: Easy way to create conversation chats

    Trust 80Quality 67support-automation
  • A fully local and private Speech-To-Text app with cross-platform support, speaker diarization, Audio Notebook…

    Trust 79Quality 75data-analysis
  • Inference9.3K stars

    Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and mult…

    Trust 87Quality 100ml-automation

Codex, Claude Code, Cursor

Coding agent skills

Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.

  • Skills122 stars

    Agent Skills for the Venice.ai API. One folder per surface area, each with a SKILL.md for agent runtimes (Cur…

    Trust 84Quality 87utility
  • OpenClaw/Codex Skill: Auto video editing for talk/vlog videos — speech recognition, sentence splitting, subti…

    Trust 81Quality 83media
  • Openai Edge Tts1.9K stars

    Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs

    Trust 80Quality 71media-automation

Documents and knowledge

Research and RAG skills

Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.

  • Dot1.9K stars

    Text-To-Speech, RAG, and LLMs. All local!

    Trust 79Quality 67data
  • Llamatik151 stars

    True on-device AI for Kotlin Multiplatform (Android, iOS, Desktop, JVM, WASM). LLM, Speech-to-Text and Image…

    Trust 78Quality 70rag-knowledge
  • Kokoro Tts1.6K stars

    A CLI text-to-speech tool using the Kokoro model, supporting multiple languages, voices (with blending), and…

    Trust 82Quality 91document-processing

Crawlers and extraction

Web scraping skills

Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.

  • Lue787 stars

    Terminal eBook Reader with Audiobook-Quality Text-to-Speech — Supports EPUB, PDF, DOCX, HTML, RTF, TXT, and M…

    Trust 78Quality 80document-processing
  • Cboard738 stars

    Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser

    Trust 79Quality 77media-automation
  • Dictionariez647 stars

    📚 A customizable dictionary extension that supports double-click lookups in 20+ languages, 1000+ dictionarie…

    Trust 79Quality 76media-automation

Prompts, B-roll, explainers

Video creation skills

Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.

  • FunClip5.8K stars

    Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integra…

    Trust 89Quality 100media-automation
  • VoxCPM31K stars

    VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloni…

    Trust 90Quality 100media-automation
  • Video Analyzer1.5K stars

    Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition

    Trust 84Quality 91media-automation

Images, video, UI

Design and creative skills

Image, video, creative production, UI design, multimodal generation, and visual workflow skills.

  • Kokoro FastAPI5.1K stars

    Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch s…

    Trust 87Quality 99media-automation
  • Mlx Audio7.4K stars

    A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framewor…

    Trust 89Quality 100media-automation
  • A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality sin…

    Trust 86Quality 87media-automation

Analysis and pipelines

Data and analytics skills

Data analysis, analytics, ETL, notebooks, databases, tables, charts, and reporting workflows.

  • TTS10K stars

    :robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/…

    Trust 82Quality 76media-automation
  • Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

    Trust 85Quality 93media-automation
  • Vosk API15K stars

    Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

    Trust 88Quality 100media-automation

Contracts and policy

Legal and compliance skills

Contract review, policy analysis, compliance research, audit support, and legal document workflows.

  • My Translator1.2K stars

    Real-time speech translation — macOS & Windows, free TTS, no server, your API keys only

    Trust 86Quality 93legal-compliance
  • Typewhisper Mac1.4K stars

    Local speech-to-text for macOS on-device AI, fully private, optional cloud

    Trust 85Quality 94legal-compliance
  • Pindrop570 stars

    A native macOS menu bar dictation app using local speech-to-text with WhisperKit

    Trust 82Quality 79legal-compliance

Supply tracks

Build the registry by domain, not just by count.

Coding

Coding and developer agents

174

Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.

174 quality168 maintained

Research

Research and knowledge work

71

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

71 quality68 maintained

Presentation

Presentation and deck workflows

1

PPTX generation, HTML slides, pitch decks, speaker notes, and presentation workflow skills.

1 quality1 maintained

Finance

Finance and quant workflows

13

Market data, SEC filings, portfolio analysis, quant research, backtesting, and risk workflows.

13 quality13 maintained

Marketing

Marketing and growth automation

10

SEO, content operations, lead generation, CRM, email automation, analytics, and growth workflows.

10 quality10 maintained

Design

Design and creative production

144

Design assets, images, video, audio, multimodal media, presentation, and creative production skills.

144 quality118 maintained

Data

Data, BI, and analytics

32

CSV, SQL, notebooks, dashboards, data pipelines, BI, ETL, and spreadsheet analysis.

32 quality32 maintained

Legal

Legal, policy, and compliance

14

Contract analysis, privacy, policy review, compliance checks, governance, and document risk review.

14 quality14 maintained

Education

Education and tutoring

5

Tutoring, course generation, quizzes, learning analytics, classrooms, and teaching workflows.

5 quality5 maintained

World Cup

Football and World Cup analytics

2

Football data, World Cup dashboards, xG, match prediction, scouting, and sports analytics.

2 quality2 maintained

High-intent entry points

Start from the task, not a keyword list.

These shortcuts use the same trust, supply, and relevance signals as the registry API, so humans and agents land on a useful shortlist faster.

Agent-readable index

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-16 of 455 ranked candidates matching "speech-to-speech"

Best blend of quality, stars, freshness, and agent usage

1

Speech To Speech

VERIFIEDEXCELLENT · 99TRUST · 86SAFE · REVIEWEDDESIGN

Build local voice agents with open-source models

$ npx skills add huggingface/speech-to-speech
4.9K stars68 quality86 trustReviewed1mo since pushSafe to try
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustStrong OpenAgentSkill Trust Score across adoption, recent maintenance, license clarity, documentation, dependency/…Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonmachine-learning
by huggingfaceDetailsQuick view
2

Speech Swift

STRONG · 78TRUST · 81SAFE · REVIEWEDDESIGN

AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML

$ npx skills add soniqo/speech-swift
894 stars54 quality81 trustReviewed2mo since pushSafe to try
QualitySolid option that is likely worth shortlisting for production workflows.
TrustGood trust signals with a few areas worth checking before rollout.Review: Quality score needs review
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

swiftspeech
by soniqoDetailsQuick view
3

Mlx Audio

VERIFIEDEXCELLENT · 100TRUST · 89SAFE · REVIEWEDDESIGN

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

$ npx skills add Blaizzy/mlx-audio
7.4K stars69 quality89 trustReviewed1mo since pushSafe to try
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustStrong OpenAgentSkill Trust Score across adoption, recent maintenance, license clarity, documentation, dependency/…Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by BlaizzyDetailsQuick view
4

PaddleSpeech

VERIFIEDEXCELLENT · 100TRUST · 90SAFE · REVIEWEDDESIGN

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…

$ npx skills add PaddlePaddle/PaddleSpeech
12.6K stars72 quality90 trustReviewed1mo since pushSafe to try
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustStrong OpenAgentSkill Trust Score across adoption, recent maintenance, license clarity, documentation, dependency/…Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by PaddlePaddleDetailsQuick view
5

Speechbrain

VERIFIEDEXCELLENT · 100TRUST · 88SAFE · REVIEWEDDESIGN

A PyTorch-based Speech Toolkit

$ npx skills add speechbrain/speechbrain
11.6K stars72 quality88 trustReviewed2mo since pushSafe to try
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustStrong OpenAgentSkill Trust Score across adoption, recent maintenance, license clarity, documentation, dependency/…Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by speechbrainDetailsQuick view
6

Speech Recognition

VERIFIEDEXCELLENT · 100TRUST · 89SAFE · REVIEWEDDESIGN

Speech recognition module for Python, supporting several engines and APIs, online and offline.

$ npx skills add Uberi/speech_recognition
9.0K stars69 quality89 trustReviewed2mo since pushSafe to try
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustStrong OpenAgentSkill Trust Score across adoption, recent maintenance, license clarity, documentation, dependency/…Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by UberiDetailsQuick view
7

ASRT SpeechRecognition

VERIFIEDEXCELLENT · 99TRUST · 86SAFE · REVIEWEDDESIGN

A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统

$ npx skills add nl8590687/ASRT_SpeechRecognition
8.4K stars66 quality86 trustReviewed4mo since pushSafe to try
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustStrong OpenAgentSkill Trust Score across adoption, recent maintenance, license clarity, documentation, dependency/…Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by nl8590687DetailsQuick view
8

CleanS2S

STRONG · 72TRUST · 78SAFE · REVIEWEDDESIGN

High-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!

$ npx skills add opendilab/CleanS2S
528 stars50 quality78 trustReviewed with permission notesOpenAI Agents4mo since pushNeeds review
QualitySolid option that is likely worth shortlisting for production workflows.
TrustGood trust signals with a few areas worth checking before rollout.Review: Quality score needs review
Safety gateUsable candidate, but the agent should surface permission and audit notes before installation.

Scenario Multimodal media

OpenAI Agents + CLI · 4 targets

pythonspeech
by opendilabDetailsQuick view
9

Automatic Speech Recognition

VERIFIEDPROMISING · 69TRUST · 78SAFE · REVIEWEDDESIGN

End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow

$ npx skills add zzw922cn/Automatic_Speech_Recognition
2.8K stars51 quality78 trustReviewed with permission notes3y since pushNeeds review
QualityUseful candidate, but compare it with alternatives before adopting.Check: Repository looks stale
TrustGood trust signals with a few areas worth checking before rollout.Review: Repository looks stale
Safety gateUsable candidate, but the agent should surface permission and audit notes before installation.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by zzw922cnDetailsQuick view
10

Tensorflow Speech Recognition

VERIFIEDPROMISING · 63TRUST · 76SAFE · EXPERIMENTALDESIGN

🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks

$ npx skills add pannous/tensorflow-speech-recognition
2.2K stars50 quality76 trustExperimental3y since pushNeeds review
QualityUseful candidate, but compare it with alternatives before adopting.Check: Repository looks stale
TrustGood trust signals with a few areas worth checking before rollout.Review: License is unclear
Safety gateSparse or mixed signals. Useful for discovery, but not for autonomous installation.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by pannousDetailsQuick view
11

StreamSpeech

VERIFIEDPROMISING · 69TRUST · 81SAFE · REVIEWEDDESIGN

StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

$ npx skills add ictnlp/StreamSpeech
1.3K stars52 quality81 trustReviewed with permission notes1y since pushNeeds review
QualityUseful candidate, but compare it with alternatives before adopting.Check: Repository looks stale
TrustGood trust signals with a few areas worth checking before rollout.Review: Repository looks stale
Safety gateUsable candidate, but the agent should surface permission and audit notes before installation.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by ictnlpDetailsQuick view
12

Expo Speech Recognition

STRONG · 76TRUST · 78SAFE · REVIEWEDDESIGN

Speech Recognition for React Native Expo projects

$ npx skills add jamsch/expo-speech-recognition
637 stars53 quality78 trustReviewed2mo since pushSafe to try
QualitySolid option that is likely worth shortlisting for production workflows.
TrustGood trust signals with a few areas worth checking before rollout.Review: Quality score needs review
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

typescriptspeech
by jamschDetailsQuick view
13

Speech To Text Benchmark

STRONG · 74TRUST · 77SAFE · REVIEWEDDESIGN

speech to text benchmark framework

$ npx skills add Picovoice/speech-to-text-benchmark
693 stars51 quality77 trustReviewed with permission notes5mo since pushNeeds review
QualitySolid option that is likely worth shortlisting for production workflows.
TrustGood trust signals with a few areas worth checking before rollout.Review: Quality score needs review
Safety gateUsable candidate, but the agent should surface permission and audit notes before installation.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by PicovoiceDetailsQuick view
14

Speech AI Forge

VERIFIEDEXCELLENT · 91TRUST · 85SAFE · REVIEWEDDESIGN

🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.

$ npx skills add lenML/Speech-AI-Forge
1.4K stars61 quality85 trustReviewed2mo since pushSafe to try
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustGood trust signals with a few areas worth checking before rollout.Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.

Scenario Multimodal media

CLI + Codex · 4 targets

pythonvoice
by lenMLDetailsQuick view
15

Speechgpt

VERIFIEDPROMISING · 69TRUST · 78SAFE · REVIEWEDCODING

💬 SpeechGPT is a web application that enables you to converse with ChatGPT.

$ npx skills add hahahumble/speechgpt
2.8K stars51 quality78 trustReviewed with permission notesOpenAI Agents3y since pushNeeds review
QualityUseful candidate, but compare it with alternatives before adopting.Check: Repository looks stale
TrustGood trust signals with a few areas worth checking before rollout.Review: Repository looks stale
Safety gateUsable candidate, but the agent should surface permission and audit notes before installation.

Scenario Coding agents

OpenAI Agents + CLI · 4 targets

typescriptchatbot
by hahahumbleDetailsQuick view
16

Cognitive Speech TTS

VERIFIEDSTRONG · 80TRUST · 81SAFE · REVIEWEDDESIGN

Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.

$ npx skills add Azure-Samples/Cognitive-Speech-TTS
1.0K stars56 quality81 trustReviewed with permission notes5mo since pushNeeds review
QualitySolid option that is likely worth shortlisting for production workflows.
TrustGood trust signals with a few areas worth checking before rollout.Review: License is unclear
Safety gateUsable candidate, but the agent should surface permission and audit notes before installation.

Scenario Multimodal media

CLI + Codex · 4 targets

c#voice
by Azure-SamplesDetailsQuick view

Page 1

Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.

Try the agent resolve API