Pandrator
lukaszliniewicz
Turn PDFs and EPUBs into audiobooks; subtitles or videos into dubbed videos (including translation), and more. For free. Pandrator uses local models, including voice-clo…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
273–288 / 466
Results: 466
lukaszliniewicz
Turn PDFs and EPUBs into audiobooks; subtitles or videos into dubbed videos (including translation), and more. For free. Pandrator uses local models, including voice-clo…
aetaric
Checkrr Scans your library files for corrupt media and optionally replaces the files via sonarr and radarr
UniversalViewer
A community-developed open source project on a mission to help you share your 📚📜📰📽️📻🗿 with the 🌎
algolia
🗣 An overlay that gets your user’s voice permission and input as text in a customizable UI
FireRedTeam
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switchin…
VolcanicArts
A modular node-programming language, program creator, animation system, toolkit, router, and debugger made for VRChat
PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models
nobody132
中文语音识别; Mandarin Automatic Speech Recognition;
altunenes
Engine for creating compute shader apps. easy multi pass pipelines, hot reload, audio/video input, and frame export
travisvn
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
homelab-00
A fully local and private Speech-To-Text app with cross-platform support, speaker diarization, Audio Notebook mode, LM Studio integration, and both longform and live tra…
julius-speech
Open-Source Large Vocabulary Continuous Speech Recognition Engine
EDCD
Companion application for Elite Dangerous
astorfi
:unlock: Lip Reading - Cross Audio-Visual Recognition using 3D Architectures
archinetai
A timeline of the latest AI models for audio generation, starting in 2023!
syhw
Attempt at tracking states of the arts and recent results (bibliography) on speech recognition.
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Video-generation prompts, B-roll, Vox-style explainers, camera direction, captions, and creative-production workflows.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.