Multimodal RAG Survey
llm-lab-org
A Survey on Multimodal Retrieval-Augmented Generation
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 84
Results: 84
llm-lab-org
A Survey on Multimodal Retrieval-Augmented Generation
aimagelab
This is the official repository for the paper "Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing". ICCV 2023
bytedance
Run multimodal agents that operate desktop interfaces
dexhunter
Write precise Seedance 2.0 prompts for multimodal video, camera movement, editing, music, and product storytelling.
fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
mlfoundations
An open-source framework for training large multimodal models.
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
invictus717
Meta-Transformer for Unified Multimodal Learning
AutoArk
EVA OS — A real-time multimodal AIOS for next-generation hardware, enabling your devices being “alive” and as intelligent as a real brain.
EvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
TIGER-AI-Lab
This repo contains the code for "VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks" [ICLR 2025]
pexoai
Expert prompt engineering for Seedance 2.0. Use when the user wants to generate a video with multimodal assets (images, videos, audio) and needs the best possible prompt.
shenhao-stu
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
autogluon
Multi-Agent System Powered by LLMs for End-to-end Multimodal ML Automation