Ovis
AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–8 / 8
Results: 8
AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
shenhao-stu
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
OthersideAI
A framework to enable a multimodal model to operate a computer.
Eventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
OpenAdaptAI
Open Source Generative Process Automation (i.e. Generative RPA). AI-First Process Automation with Large ([Language (LLMs) / Action (LAMs) / Multimodal (LMMs)] / Visual L…
QIN2DIM
🥂 Gracefully face hCaptcha challenge with multimodal large language model.
declare-lab
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
WisconsinAIVision
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts