🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
每个推荐都保留与其仓库、审计和安装路径的明确关联。
搜索结果: audio-editing
英文目录AI generates a real, editable PowerPoint from any document — native shapes & animations, speaker notes voiced as audio narration, and the option to follow your own .pptx template, not slide images · by Hugo He
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
A tool that lets AI agents like Claude Code edit videos by cutting filler words, color grading, adding subtitles, and more, all via natural language commands.
A fully open-source headless CMS that supports Markdown and Visual Editing
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Video editing with Python
🎛 🔊 A Python library for audio.
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale