Stable Diffusion.Cpp
leejet
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
65–80 / 98
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 98
leejet
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
cubist38
A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPI framework, it provides an effic…
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
Tencent-Hunyuan
HunyuanVideo-1.5: A leading lightweight video generation model
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
sanchit-gandhi
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
SkyworkAI
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
metavoiceio
Foundational model for human-like, expressive TTS
lenML
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
devnen
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice clonin…
diodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
ai-forever
Kandinsky 2 — multilingual text2image latent diffusion model