OmniGen
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 25
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 25
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
wernerturing
Detect and remove logos from videos, even if they change position several times
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
netease-youdao
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
myshell-ai
High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
huanngzh
[ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
bytedance
SALMONN family: A suite of advanced multi-modal LLMs
fluxions-ai
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime o…
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
diodiogod
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio…
FireRedTeam
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity gene…
Enemyx-net
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your C…