Izwi
izwi-ai
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
433–448 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
izwi-ai
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.
rockbenben
🎨 AI Image Prompt Generator with 5000+ prompts in 18 languages. Create perfect prompts fin your native language. 极简的图像提示词编辑器
Xeron2000
故事想法 → 多智能体协作 → 漫剧成片 | 基于 LangGraph 的 AI 漫剧生成平台
echogarden-project
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of engines for speech synthesis, speech recognition, forced alignme…
gudaochangsheng
[ECCV 2026] Official PyTorch implementation of RefAlign: Representation Alignment for Reference-to-Video Generation
travisvn
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned speech anywhere the OpenAI API is used (e.g. Open WebUI, Anyt…
woheller69
Android Input Method Editor (IME) based on Whisper
PowerBeef
Vocello — a local, private voice studio for Apple Silicon. Write a script, pick or describe a voice, and generate speech on-device. macOS today, iPhone coming soon. (For…
martin-rizzo
A set of ComfyUI nodes designed specifically for the Z-Image / Z-Image Turbo model.
Djdefrag
RealScaler - image/video AI upscaler app (Real-ESRGAN)
InternRobotics
[ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling
wildminder
ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio
zzh-tech
[TPAMI][ECCV2024 Oral] Clearer anytime frame interpolation & Manipulated interpolation of anything
ChaitanyaEswarRajeshJakki
A fully autonomous AI Agent/Python pipeline that utilizes Large Language Models (LLMs) like Gemini to generate content, produce videos, and automatically upload educatio…
SamirPaulb
A desktop application that uses AI to translate voice between languages in real time, while preserving the speaker's tone and emotion.
Correr-Zhou
[ICML 2026] ByteDance's All-in-One Video Generation Model for Human-Object Interaction Video Generation
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.