No visual example yet
Explore the skillMultimodal RAG Survey
llm-lab-org
A Survey on Multimodal Retrieval-Augmented Generation
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 21
Results: 21
No visual example yet
Explore the skillllm-lab-org
A Survey on Multimodal Retrieval-Augmented Generation
No visual example yet
Explore the skillTencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
No visual example yet
Explore the skillbytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
No visual example yet
Explore the skillEvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
No visual example yet
Explore the skillgorse-io
AI powered open source recommender system engine supports classical/LLM rankers and multimodal content via embedding
No visual example yet
Explore the skillOthersideAI
A framework to enable a multimodal model to operate a computer.
No visual example yet
Explore the skillTHUDM
Towards Large Multimodal Models as Visual Foundation Agents
No visual example yet
Explore the skillTencentQQGYLab
AppAgent: Multimodal Agents as Smartphone Users, an LLM-based multimodal agent framework designed to operate smartphone apps.
No visual example yet
Explore the skillhkchengrex
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
No visual example yet
Explore the skillJIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
No visual example yet
Explore the skillRQLuo
MixTeX multimodal LaTeX, ZhEn, and, Table OCR. It performs efficient CPU-based inference in a local offline on Windows.
No visual example yet
Explore the skillTIGER-AI-Lab
Official Repo for "TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding" [ACL 2025 oral]
No visual example yet
Explore the skillmbzuai-oryx
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with…
No visual example yet
Explore the skillkohjingyu
🧀 Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs".
No visual example yet
Explore the skillkohjingyu
🐟 Code and models for the NeurIPS 2023 paper "Generating Images with Multimodal Language Models".
No visual example yet
Explore the skillopengeos
A multimodal AI agent for geospatial data analysis and interactive visualization