OmniGen
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 335
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 335
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
R3gm
Synchronized Translation for Videos. Video dubbing
denizsafak
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
ali-vilab
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
bytedance
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
algolia
🗣 An overlay that gets your user’s voice permission and input as text in a customizable UI
OEvortex
LLM4Free — All-in-one Python toolkit for web search, AI interaction (40+ free providers), digital utilities, and more. Formerly WebScout.
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning