TTS
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
33–48 / 84
Results: 84
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Renumics
Interactively explore unstructured datasets from your dataframe.
libAudioFlux
A library for audio and music analysis, feature extraction.
milvus-io
Dealing with all unstructured data, such as reverse image search, audio search, molecular search, video analysis, question and answer systems, NLP, etc.
FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
CPJKU
Python audio and music signal processing library
mozilla
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
pexoai
A collection of open-source Agent Skills for content creation — images, audio, and video.
tyiannak
Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications
lauragift21
🔥 Awesome list of resources on Web Development.
watzon
A native macOS menu bar dictation app using local speech-to-text with WhisperKit
backblaze-labs
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.
apocas
RESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics et…
GauravSingh9356
Personal Assistant built using python libraries. It does almost anything which includes sending emails, Optical Text Recognition, Dynamic News Reporting at any time with…
wladradchenko
Wunjo CE: Face Swap, Lip Sync, Control Remove Objects & Text & Background, Restyling, Audio Separator, Clone Voice, Video Generation. Open Source, Local & Free.
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.