SALMONN
bytedance
SALMONN family: A suite of advanced multi-modal LLMs
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
145–160 / 460
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 460
bytedance
SALMONN family: A suite of advanced multi-modal LLMs
met4citizen
Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.
devnen
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice clonin…
s60sc
ESP32 Camera motion capture application to record JPEGs to SD card as AVI files and stream to browser as MJPEG. If a microphone is installed then a WAV file is also crea…
pluja
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
junyanz
Image-to-Image Translation in PyTorch
camenduru
stable diffusion webui colab
VectorSpaceLab
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
leejet
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
junyanz
Software that can generate photos from paintings, turn horses into zebras, perform style transfer, and more.
vllm-project
A framework for efficient model inference with omni-modality models
KnpLabs
PHP library allowing thumbnail, snapshot or PDF generation from a url or a html page. Wrapper for wkhtmltopdf/wkhtmltoimage
phillipi
Image-to-image translation with conditional adversarial nets
carson-katri
Stable Diffusion built-in to Blender
open-mmlab
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion mod…