Modelscope
modelscope
ModelScope: bring the notion of Model-as-a-Service to life.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 22
Results: 22
modelscope
ModelScope: bring the notion of Model-as-a-Service to life.
Tencent-Hunyuan
HunyuanVideo: A Systematic Framework For Large Video Generation Model
vllm-project
A framework for efficient model inference with omni-modality models
PKU-YuanGroup
Helios: Real Real-Time Long Video Generation Model
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
Rimagination
A hardware-aware Codex skill for local MiniMax H3 video generation through ComfyUI
RVC-Boss
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
ggml-org
Port of OpenAI's Whisper model in C/C++
FunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,…
nari-labs
A TTS model capable of generating ultra-realistic dialogue in one pass.
myshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.
GetStream
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
leejet
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching