Open Speech Corpora
coqui-ai
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Browse all public registry entries, one page at a time. Counts include resources; MCP-only resources are omitted from displayed skills. Listing is not a safety or compatibility guarantee.
Page 20 · 16 shown · 828 public entries
Results: 828
coqui-ai
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
lucidrains
Implementation of Video Diffusion Models, Jonathan Ho's new paper extending DDPMs to Video Generation - in Pytorch
ali-vilab
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
wernerturing
Detect and remove logos from videos, even if they change position several times
mayuelala
[AAAI 2024] Follow-Your-Pose: This repo is the official implementation of "Follow-Your-Pose : Pose-Guided Text-to-Video Generation using Pose-Free Videos"
receyuki
A simple standalone viewer for reading prompts from Stable Diffusion generated image outside the webui.
CSTR-Edinburgh
This is now the official location of the Merlin project.
mit-han-lab
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
jishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
zqbxdev
OpenAI-compatible Web Chat API proxy with GPT, Grok, Gemini account management and Docker self-hosting
Uminosachi
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
cubist38
A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPI framework, it provides an effic…
OEvortex
LLM4Free — All-in-one Python toolkit for web search, AI interaction (40+ free providers), digital utilities, and more. Formerly WebScout.
myccarl
End-to-end AI short-video production pipeline. FastAPI orchestration + Spring Boot gateway with multi-model failover, circuit breaker, metering, and full-stack observabi…
ictnlp
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
RedSkill · 流白Livo · 1.0.0
Describe a rain curtain, growing flowering branches or a flock of swallows. This RedSkill package adapts three p5.js templates into sketch.js code with custom colors, density and speed.
Noncommercial use only · Local runtime recording · Use on the source platform