No visual example yet
Explore the skillOmAgent
om-ai-lab
[EMNLP-2024] Build multimodal language agents for fast prototype and production
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
49–64 / 84
Results: 84
No visual example yet
Explore the skillom-ai-lab
[EMNLP-2024] Build multimodal language agents for fast prototype and production
No visual example yet
Explore the skilldeclare-lab
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
No visual example yet
Explore the skillunum-cloud
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
No visual example yet
Explore the skillTencent-Hunyuan
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
No visual example yet
Explore the skillRQLuo
MixTeX multimodal LaTeX, ZhEn, and, Table OCR. It performs efficient CPU-based inference in a local offline on Windows.
No visual example yet
Explore the skillTIGER-AI-Lab
Official Repo for "TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding" [ACL 2025 oral]
No visual example yet
Explore the skillmims-harvard
Therapeutics Commons (TDC): Multimodal Foundation for Therapeutic Science
No visual example yet
Explore the skillgo-kratos
Blades is a Go-based multimodal AI Agent framework.
No visual example yet
Explore the skillHaervwe
Open‑WebUI Tools is a modular toolkit designed to extend and enrich your Open WebUI instance, turning it into a powerful AI workstation. With a suite of over 15 speciali…
No visual example yet
Explore the skilllancedb
Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs
No visual example yet
Explore the skillframersai
Build autonomous AI agents with adaptive intelligence and emergent behaviors, included with multimodal RAG and optional HEXACO personalities.
No visual example yet
Explore the skillHorizonWind2004
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models…
No visual example yet
Explore the skillmbzuai-oryx
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with…
No visual example yet
Explore the skillfoxglove
Multimodal visualization and data platform
No visual example yet
Explore the skillberabuddies
Multimodal document deep analysis tool based on Zhipu GLM-OCR, GLM-4.7, and GLM-4.6V.
No visual example yet
Explore the skillduoan
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models