Dsnote
mkiol
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 6 · 16 shown · 743 public entries
Results: 743
mkiol
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
bytedance
SALMONN family: A suite of advanced multi-modal LLMs
coqui-ai
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
lenML
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
TypeWhisper
Local speech-to-text for macOS on-device AI, fully private, optional cloud
met4citizen
Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.
shivammehta25
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching
FunAudioLLM
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
devnen
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice clonin…
Renumics
Interactively explore unstructured datasets from your dataframe.
Yuan-ManX
Your AI Game Dev Hub. The ultimate resource hub for AI-powered game development tools. Discover cutting-edge LLMs, World Model, Agent, Code, Image, Texture, Shader, 3D M…
kubeai-project
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
modal-labs
myshell-ai
Instant voice cloning by MIT and MyShell. Audio foundation model.