Ovis
AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–8 / 8
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 8
AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
cambrian-mllm
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
shikiw
[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
theopenconversationkit
Tock, the open source conversational AI toolkit.
Blaizzy
MLX-Embeddings is the best package for running Vision and Language Embedding models locally on your Mac using MLX.
OpenGVLab
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4…
ictnlp
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
yuanze-lin
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"