Argmax Oss Swift
argmaxinc
On-device Speech AI for Apple Silicon
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
17–32 / 335
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 335
argmaxinc
On-device Speech AI for Apple Silicon
m-bain
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
open-mmlab
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get…
babysor
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
unslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
OpenBMB
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
invoke-ai
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest…
AIDC-AI
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
Tencent-Hunyuan
HunyuanVideo: A Systematic Framework For Large Video Generation Model
AaronFeng753
Video, Image and GIF upscale/enlarge(Super-Resolution) and Video frame interpolation. Achieved with Waifu2x, Real-ESRGAN, Real-CUGAN, RTX Video Super Resolution VSR, SRM…
index-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System