Cosmos Predict2.5
nvidia-cosmos
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
65–80 / 382
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 382
nvidia-cosmos
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the…
OpenGVLab
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, S…
k2-fsa
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Su…
supertone-inc
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
Wan-Video
Wan: Open and Advanced Large-Scale Video Generative Models
FireRedTeam
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity gene…
leandromoreira
FFmpeg libav tutorial - learn how media works from basic to transmuxing, transcoding and more. Translations: 🇺🇸 🇨🇳 🇰🇷 🇪🇸 🇻🇳 🇧🇷 🇷🇺
openvinotoolkit
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
HKUDS
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
duixcom
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
nari-labs
A TTS model capable of generating ultra-realistic dialogue in one pass.
myshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.