CleanS2S
opendilab
High-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
225–240 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
opendilab
High-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!
Echo-Team-Joy-Future-Academy-JD
A Simple Baseline for Video World Models with Memory
rapidaai
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channe…
Agions
scene-fab - AI 影视解说创作工具 | 智能拆条 · AI 解说生成 · 一键配音合成
vilassn
Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android
lidge-jun
Minimal CLI + web UI for OpenAI GPT Image 2 generation. Dual auth: API Key (paid) or OAuth via ChatGPT (free). Text-to-image, image-to-image, parallel gen, custom sizes.
helloianneo
Xiaohei 2.0 Codex Skill for Chinese real-object article illustrations and long-scroll story images
izwi-ai
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.
githubharald
Connectionist Temporal Classification (CTC) decoder with dictionary and language model.
Aratako
Multilingual TTS model with voice cloning and duration control, based on T5Gemma encoder-decoder LLM
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
duixcom
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
debpalash
The open-source ElevenLabs alternative for local voice cloning, design, create, dubbing and dictation Desktop App
modelscope
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
rsxdalv
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Au…
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.