Kokoro FastAPI
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
81–96 / 398
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 398
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
Breakthrough
:movie_camera: Python and OpenCV-based scene cut/transition detection program & library.
Picovoice
On-device wake word detection powered by deep learning
espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
phillipi
Image-to-image translation with conditional adversarial nets
mozilla
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
hao-ai-lab
A unified inference and post-training framework for accelerated video generation.
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.