VAR
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
65–80 / 382
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 382
FoundationVision
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation…
phillipi
Image-to-image translation with conditional adversarial nets
mozilla
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
MahmoudAshraf97
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
hao-ai-lab
A unified inference and post-training framework for accelerated video generation.
OpenMOSS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and…
IAHispano
A simple, high-quality voice conversion tool focused on ease of use and performance.
flashlight
Facebook AI Research's Automatic Speech Recognition Toolkit
Tencent-Hunyuan
HunyuanVideo-1.5: A leading lightweight video generation model
elevenlabs
The official Python SDK for the ElevenLabs API.
netease-youdao
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine