No visual example yet
Explore the skillMultimodal RAG Survey
llm-lab-org
A Survey on Multimodal Retrieval-Augmented Generation
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 84
Results: 84
No visual example yet
Explore the skillllm-lab-org
A Survey on Multimodal Retrieval-Augmented Generation
No visual example yet
Explore the skillaimagelab
This is the official repository for the paper "Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing". ICCV 2023
No visual example yet
Explore the skillbytedance
Run multimodal agents that operate desktop interfaces
No visual example yet
Explore the skilldexhunter
Write precise Seedance 2.0 prompts for multimodal video, camera movement, editing, music, and product storytelling.
No visual example yet
Explore the skillfikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
No visual example yet
Explore the skillTencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
No visual example yet
Explore the skillAIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
No visual example yet
Explore the skillmlfoundations
An open-source framework for training large multimodal models.
No visual example yet
Explore the skillbytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
No visual example yet
Explore the skillinvictus717
Meta-Transformer for Unified Multimodal Learning
No visual example yet
Explore the skillAutoArk
EVA OS — A real-time multimodal AIOS for next-generation hardware, enabling your devices being “alive” and as intelligent as a real brain.
No visual example yet
Explore the skillEvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
No visual example yet
Explore the skillTIGER-AI-Lab
This repo contains the code for "VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks" [ICLR 2025]
No visual example yet
Explore the skillpexoai
Expert prompt engineering for Seedance 2.0. Use when the user wants to generate a video with multimodal assets (images, videos, audio) and needs the best possible prompt.
No visual example yet
Explore the skillshenhao-stu
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
No visual example yet
Explore the skillautogluon
Multi-Agent System Powered by LLMs for End-to-end Multimodal ML Automation