No visual example yet
Explore the skillAutodistill
autodistill
Images to inference with no labeling (use foundation models to train supervised models).
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
70 Skills
Results: 70
No visual example yet
Explore the skillautodistill
Images to inference with no labeling (use foundation models to train supervised models).
No visual example yet
Explore the skillDamRsn
Audio Plugin for Audio to MIDI transcription using deep learning.
No visual example yet
Explore the skillHitachi-Automotive-And-Industry-Lab
Web labeling tool for bitmap images and point clouds
No visual example yet
Explore the skillMiteshPuthran
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
No visual example yet
Explore the skilljishengpeng
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
No visual example yet
Explore the skillmravanelli
SincNet is a neural architecture for efficiently processing raw audio samples.
No visual example yet
Explore the skillPluviobyte
Generate a controlled local narration workflow with auditions, version tracking, and subtitle-ready final audio.

hugohe3
AI generates a real, editable PowerPoint from any document — native shapes & animations, speaker notes voiced as audio narration, and the option to follow your own .pptx…
View previews · 2No visual example yet
Explore the skillbackblaze-labs
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.
No visual example yet
Explore the skillrapidaai
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channe…
No visual example yet
Explore the skillinfiniV
Local voice dictation and meeting recorder for Windows + Linux. Hold a hotkey to dictate, or record long-form meetings with system audio. Whisper transcription, bring-yo…
No visual example yet
Explore the skillhuggingface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
No visual example yet
Explore the skillEmily2040
Direct the model. Do not micro-manage the frame.
No visual example yet
Explore the skillmyshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.
No visual example yet
Explore the skillEventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
No visual example yet
Explore the skillwhitphx
Real-time video and audio processing on Streamlit