Fun ASR
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
145–160 / 462
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 462
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
devnen
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice clonin…
nvidia-cosmos
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the…
Tencent
High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
numz
Official SeedVR2 Video Upscaler for ComfyUI
letoram
Arcan - [Display Server, Multimedia Framework, Game Engine] -> "Desktop Engine"
muxinc
The easiest way to add video in your Nextjs app.
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
ThioJoe
Automatically translates the text of a video based on a subtitle file, and then uses AI voice services to create a new dubbed & translated audio track where the speech i…
Vchitect
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
s60sc
ESP32 Camera motion capture application to record JPEGs to SD card as AVI files and stream to browser as MJPEG. If a microphone is installed then a WAV file is also crea…
janvarev
Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.
lhotse-speech
Tools for handling multimodal data in machine learning projects.