CogVideo
zai-org
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 19
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 19
zai-org
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
gradio-app
The python library for real-time communication
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
thu-ml
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, i…
ARahim3
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
FireRedTeam
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering ou…
thu-ml
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
FoundationVision
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
BinWang28
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
PurpleDoubleD
Local AI desktop app — chat, agent mode, image gen, video gen. Supports Ollama, Gemma 4, Llama, Qwen, OpenAI, Anthropic. Single .exe, no Docker.
taco-group
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
lone-cloud
A desktop app for running Large Language Models locally.
modelscope
Open-source, accurate and easy-to-use video speech recognition & clipping tool. LLM-based AI clipping integrated.
fluxions-ai
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime o…