WhisperX
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 319
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 319
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
HKUDS
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
sanchit-gandhi
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
junyanz
Image-to-Image Translation in PyTorch
phillipi
Image-to-image translation with conditional adversarial nets
OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
AaronFeng753
Video, Image and GIF upscale/enlarge(Super-Resolution) and Video frame interpolation. Achieved with Waifu2x, Real-ESRGAN, Real-CUGAN, RTX Video Super Resolution VSR, SRM…
index-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Blaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
nateshmbhat
Offline Text To Speech synthesis for python
pydn
A powerful tool that translates ComfyUI workflows into executable Python code.
SkyworkAI
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.