Directorio de skills

Descubre skills reutilizables para AI agents.

Busca skills reales de GitHub por tarea y revisa stars, confianza, auditoría, categoría y ruta de instalación antes de utilizarlos.

Cada recomendación conserva un vínculo claro con su repositorio, auditoría y ruta de instalación.

Resultados de búsqueda: audio-recording

Directorio en inglés

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

162K
Stars
87/100
Confianza
Categoría: ml-automationAuditoría

AI generates a real, editable PowerPoint from any document — native shapes & animations, speaker notes voiced as audio narration, and the option to follow your own .pptx template, not slide images · by Hugo He

37K
Stars
81/100
Confianza
Categoría: agent-frameworksAuditoría

🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.

34K
Stars
87/100
Confianza
Categoría: media-automationAuditoría

Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.

8.6K
Stars
77/100
Confianza
Categoría: media-automationAuditoría

A tool that lets AI agents like Claude Code edit videos by cutting filler words, color grading, adding subtitles, and more, all via natural language commands.

17K
Stars
85/100
Confianza
Categoría: mediaAuditoría

Android in docker solution with noVNC supported and video recording

15K
Stars
76/100
Confianza
Categoría: devopsAuditoría

Desktop app that records your on-screen work session and uses the GitHub Copilot CLI to reconstruct it as an intent + ordered steps, then builds a reusable Skill or Automation for Microsoft Scout, Microsoft Copilot Cowork, or Copilot Studio.

3.0K
Stars
83/100
Confianza
Categoría: utilityAuditoría

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

12K
Stars
87/100
Confianza
Categoría: media-automationAuditoría

A curated collection of reusable agent skills for designers and builders to generate UI prompts and workflows using AI coding agents.

4.8K
Stars
86/100
Confianza
Categoría: design-creativeAuditoría

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

7.4K
Stars
86/100
Confianza
Categoría: media-automationAuditoría

🎛 🔊 A Python library for audio.

6.2K
Stars
77/100
Confianza
Categoría: ml-automationAuditoría

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

5.6K
Stars
86/100
Confianza
Categoría: ml-automationAuditoría