No visual example yet
Explore the skillOvis
AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–8 / 8
Results: 8
No visual example yet
Explore the skillAIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
No visual example yet
Explore the skillshenhao-stu
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
No visual example yet
Explore the skillOthersideAI
A framework to enable a multimodal model to operate a computer.
No visual example yet
Explore the skillEventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
No visual example yet
Explore the skillOpenAdaptAI
Open Source Generative Process Automation (i.e. Generative RPA). AI-First Process Automation with Large ([Language (LLMs) / Action (LAMs) / Multimodal (LMMs)] / Visual L…
No visual example yet
Explore the skillQIN2DIM
🥂 Gracefully face hCaptcha challenge with multimodal large language model.
No visual example yet
Explore the skilldeclare-lab
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
No visual example yet
Explore the skillWisconsinAIVision
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts