No visual example yet
Explore the skillParlor
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
129–144 / 466
Results: 466
No visual example yet
Explore the skillfikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
No visual example yet
Explore the skillNatively-AI-assistant
Natively — Free open-source AI meeting assistant, interview copilot, and note taker. The best alternative to Cluely, Otter, Granola, Final Round AI, Fireflies, and Inter…
No visual example yet
Explore the skillkaldi-asr
kaldi-asr/kaldi is the official location of the Kaldi project.
No visual example yet
Explore the skillRHVoice
a free and open source speech synthesizer for Russian and other languages
No visual example yet
Explore the skillmltframework
MLT Multimedia Framework
No visual example yet
Explore the skillespeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
No visual example yet
Explore the skillmicrosoft
A Unified Semi-Supervised Learning Codebase (NeurIPS'22)
No visual example yet
Explore the skillOpenShot
OpenShot Video Library (libopenshot) is a free, open-source project dedicated to delivering high quality video editing, animation, and playback solutions to the world. A…
No visual example yet
Explore the skillmkiol
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
No visual example yet
Explore the skillMahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
No visual example yet
Explore the skillbytedance
SALMONN family: A suite of advanced multi-modal LLMs
No visual example yet
Explore the skillcoqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
No visual example yet
Explore the skilllenML
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
No visual example yet
Explore the skillARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
No visual example yet
Explore the skillTypeWhisper
Local speech-to-text for macOS on-device AI, fully private, optional cloud
No visual example yet
Explore the skillmet4citizen
Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.
Design direction, Figma implementation, UI review, React performance, browser QA, and safe preview deployments.
Code review, repo inspection, testing, planning, shipping, and engineering workflows for coding agents.
Research, recent-events briefings, PDF parsing, markdown conversion, RAG ingestion, and knowledge workflows.
Stock analysis, market research, quant backtesting, financial data, and investment research skills.
Crawling, scraping, extraction, browser automation, structured data capture, and website-to-markdown workflows.
Presentation generation, editable PPTX, slide decks, speaker notes, and visual storytelling workflows.
Choose React motion graphics, generated footage or editing; compare dependencies and inspect actual outputs.
Image, video, creative production, UI design, multimodal generation, and visual workflow skills.