Multi Modality Arena
OpenGVLab
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–15 / 15
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 15
OpenGVLab
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4…
AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
shikiw
[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
ictnlp
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
cambrian-mllm
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
declare-lab
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
Zaki-1052
A feature-rich portal to chat with GPT-4, Claude, Gemini, Mistral, & OpenAI Assistant APIs via a lightweight Node.js web app; supports customizable multimodality for voi…
InternRobotics
[ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
alan-ai
The Self-Coding System for Your App — Alan AI SDK for Ionic
mbzuai-oryx
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrai…
alan-ai
The Self-Coding System for Your App — Alan AI SDK for Cordova
alan-ai
The Self-Coding System for Your App — Alan AI SDK for React Native
RLHF-V
[CVPR'25 highlight] RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness