No visual example yet
Explore the skillFastVideo
hao-ai-lab
A unified inference and post-training framework for accelerated video generation.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
11 Skills
Results: 11
No visual example yet
Explore the skillhao-ai-lab
A unified inference and post-training framework for accelerated video generation.
No visual example yet
Explore the skillyl4579
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
No visual example yet
Explore the skillunslothai
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
No visual example yet
Explore the skillFunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
No visual example yet
Explore the skillNVIDIA
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference appl…
No visual example yet
Explore the skillhkchengrex
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
No visual example yet
Explore the skillcoqui-ai
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
No visual example yet
Explore the skillyeyupiaoling
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inf…
No visual example yet
Explore the skillthu-ml
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
No visual example yet
Explore the skillEvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
No visual example yet
Explore the skillFinrandojin
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptered M4B,…