Dsnote
mkiol
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
145β160 / 462
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 462
mkiol
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
Purfview
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
fudan-generative-vision
[ECCV 2024] Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
Picsart-AI-Research
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
ali-vilab
Official implementations for paper: Anydoor: zero-shot object-level image customization
bytedance
SALMONN family: A suite of advanced multi-modal LLMs
metavoiceio
Foundational model for human-like, expressive TTS
huggingface
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
HITsz-TMG
π AI ε ¨θͺε¨εθ§ι’ηζεε·₯ | Your First AIGC Coworker. Chat an Idea. Get a Film. π¦
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
TensorSpeech
:stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, Germanβ¦
lenML
π¦ Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning β natively on MLX. Unsloth-compatible API.
Janspiry
Unofficial implementation of Image Super-Resolution via Iterative Refinement by Pytorch
met4citizen
Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.
bytedance
π₯ [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity