No visual example yet
Explore the skillMMMU
MMMU-Benchmark
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
17–32 / 84
Results: 84
No visual example yet
Explore the skillMMMU-Benchmark
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
No visual example yet
Explore the skillcaiyuanhao1998
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions (NeurIPS 2025)
No visual example yet
Explore the skillpliang279
[NeurIPS 2021] Multiscale Benchmarks for Multimodal Representation Learning
No visual example yet
Explore the skillhuggingface
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and traini…
No visual example yet
Explore the skilldeepset-ai
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit con…
No visual example yet
Explore the skillRoffyS
Convert files (PDF, image, Word, PPT, Excel, notebooks, code snippets) to markdown using powerful multimodal LLM
No visual example yet
Explore the skillxiincs
为 Claude Code 赋能多模态视觉能力,支持豆包、通义千问、GPT-4o 等模型,用于截图 / UI / 图表分析;适配 DeepSeek 等无视觉底座,搭配 browser-harness 可做前端布局自动化检查。
No visual example yet
Explore the skillzeyofu
This repo contains evaluation code for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive". https://arxiv.org/abs/2404.12390 [ECCV 2024]
No visual example yet
Explore the skillapache
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
No visual example yet
Explore the skilllancedb
Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.
No visual example yet
Explore the skillgorse-io
AI powered open source recommender system engine supports classical/LLM rankers and multimodal content via embedding
No visual example yet
Explore the skillxorbitsai
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all throug…
No visual example yet
Explore the skillactiveloopai
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
No visual example yet
Explore the skilljina-ai
☁️ Build multimodal AI applications with cloud-native stack
No visual example yet
Explore the skillOthersideAI
A framework to enable a multimodal model to operate a computer.
No visual example yet
Explore the skillEventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale