FireRedASR
FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
289–304 / 478
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 478
FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
bytedance
🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
shivammehta25
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
devnen
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice clonin…
letoram
Arcan - [Display Server, Multimedia Framework, Game Engine] -> "Desktop Engine"
muxinc
The easiest way to add video in your Nextjs app.
Vchitect
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
s60sc
ESP32 Camera motion capture application to record JPEGs to SD card as AVI files and stream to browser as MJPEG. If a microphone is installed then a WAV file is also crea…
lhotse-speech
Tools for handling multimodal data in machine learning projects.
FoundationVision
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
OpenGVLab
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, S…
tnfe
A fast video processing library based on node.js (一个基于node.js的高速视频制作库)
DigitalPhonetics
Controllable and fast Text-to-Speech for over 7000 languages!