🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Skill-Verzeichnis
Wiederverwendbare Skills für AI Agents entdecken.
Jede Empfehlung bleibt mit ihrem Repository, Audit und Installationspfad nachvollziehbar.
Suchergebnisse: audio-inference
Englisches VerzeichnisMNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
AI generates a real, editable PowerPoint from any document — native shapes & animations, speaker notes voiced as audio narration, and the option to follow your own .pptx template, not slide images · by Hugo He
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
A tool that lets AI agents like Claude Code edit videos by cutting filler words, color grading, adding subtitles, and more, all via natural language commands.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Run agents like Hermes and OpenClaw more securely inside NVIDIA OpenShell with managed inference
Open source alternative to AWS. Elastic compute, block storage (non replicated), firewall and load balancer, managed Postgres, K8s, AI inference, and IAM services.
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
The Triton Inference Server provides an optimized cloud and edge inferencing solution.