SenseVoice
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Search retrieves candidates across the registry and ranks a bounded shortlist by task fit. This count is matching candidates, not the registry total. No suitable match? Try a specific tool or task.
1–16 / 95
Results: 95
FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
Blaizzy
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
PINTO0309
A repository for storing models that have been inter-converted between various frameworks. Supported frameworks are TensorFlow, PyTorch, ONNX, OpenVINO, TFJS, TFTRT, Ten…
FunAudioLLM
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
mlabonne
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
modelscope
ModelScope: bring the notion of Model-as-a-Service to life.
vllm-project
A framework for efficient model inference with omni-modality models
Pluviobyte
Generate real-timestamp subtitle artifacts from final narration audio or merged video with a caption quality gate.
huggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
huggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
py-why
DoWhy is a Python library for causal inference that supports explicit modeling and testing of causal assumptions. DoWhy is based on a unified language for causal inferen…
anthropics
Complete and link income statement, balance sheet, and cash flow statement model templates with formulas and checks.
Tencent-Hunyuan
HunyuanVideo: A Systematic Framework For Large Video Generation Model
abus-aikorea
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTu…
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.