Wav2letter
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
33–48 / 67
Results: 67
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
CPJKU
Python audio and music signal processing library
IliasHad
Local-first Video Knowledge Base. Index your video library with multi-modal analysis (YOLO, DeepFace, Whisper), search semantically via natural language, Docker-ready.
mozilla
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
tyiannak
Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications
watzon
A native macOS menu bar dictation app using local speech-to-text with WhisperKit
backblaze-labs
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.
video-db
AI video agents framework for next-gen video interactions and workflows.
aqm857886159
Open-source, local-first desktop app for AI video creation: write a script → generate images & video → edit on a timeline → export. Bring your own model & API key — ever…
GauravSingh9356
Personal Assistant built using python libraries. It does almost anything which includes sending emails, Optical Text Recognition, Dynamic News Reporting at any time with…
unum-cloud
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
wladradchenko
Wunjo CE: Face Swap, Lip Sync, Control Remove Objects & Text & Background, Restyling, Audio Separator, Clone Voice, Video Generation. Open Source, Local & Free.
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
Saik0s
The open-source iOS app that's making quality voice transcription more accessible on mobile devices.
Agions
scene-fab - AI 影视解说创作工具 | 智能拆条 · AI 解说生成 · 一键配音合成
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.