UNO
bytedance
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
33–48 / 351
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 351
bytedance
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
sandrohanea
Whisper.net. Speech to text made simple using Whisper Models
AutoArk
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
thu-ml
Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (ICML 2025) , UltraViCo (ICLR 2026) and UltraImage
kaltura
The Kaltura Platform Backend. To install Kaltura, visit the install packages repository.
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
NVIDIA
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
junyanz
Software that can generate photos from paintings, turn horses into zebras, perform style transfer, and more.
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
IAHispano
A simple, high-quality voice conversion tool focused on ease of use and performance.
abhiTronix
A High-performance cross-platform Video Processing Python framework powerpacked with unique trailblazing features :fire:
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.