Audio
pytorch
Data manipulation and transformation for audio signal processing, powered by PyTorch
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 2 · 16 shown · 223 public entries
Results: 223
pytorch
Data manipulation and transformation for audio signal processing, powered by PyTorch
bghira
A general fine-tuning kit geared toward image/video/audio diffusion models.
yuanzhongqiao
短剧平台 AI Short Film Motion Comic Generation Platform Industrial AI Motion Comic & Video Workbench
showlab
[ICML 2026] Video generation via code
NVIDIA-AI-Blueprints
The NVIDIA VSS Blueprint is a suite of reference architectures for building GPU-accelerated vision agents and AI-powered video analytics applications.
Skardyy
Terminal image, video, and Markdown viewer
hakanyalcinkaya
Kodluyoruz için Hazırladığım Video Eğitim Seti Repo'sudur. Tüm Eğitimlerime: https://linktr.ee/hakanyalcinkaya adresinden ulaşabilirsiniz.
nvidia-cosmos
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the…
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
Tencent-Hunyuan
HunyuanVideo-1.5: A leading lightweight video generation model
myshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.
GVCLab
[CVPR 2026] PersonaLive! : Expressive Portrait Image Animation for Live Streaming
HKUDS
[KDD'2026] "VideoRAG: Chat with Your Videos"
hkchengrex
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
ThioJoe
Automatically translates the text of a video based on a subtitle file, and then uses AI voice services to create a new dubbed & translated audio track where the speech i…
byjlw
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition