暂未收录效果图
查看技能说明Multimodal RAG Survey
llm-lab-org
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
84 Skills
搜索结果: 84
暂未收录效果图
查看技能说明llm-lab-org
暂未收录效果图
查看技能说明aimagelab
This is the official repository for the paper "Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing". ICCV 2023
暂未收录效果图
查看技能说明bytedance
暂未收录效果图
查看技能说明dexhunter
Write precise Seedance 2.0 prompts for multimodal video, camera movement, editing, music, and product storytelling.
暂未收录效果图
查看技能说明fikrikarim
On-device, real-time multimodal AI. Have natural voice and vision conversations with an AI that runs entirely on your machine. Powered by Gemma 4 E2B and Kokoro.
暂未收录效果图
查看技能说明Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
暂未收录效果图
查看技能说明AIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
暂未收录效果图
查看技能说明mlfoundations
An open-source framework for training large multimodal models.
暂未收录效果图
查看技能说明bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
暂未收录效果图
查看技能说明invictus717
Meta-Transformer for Unified Multimodal Learning
暂未收录效果图
查看技能说明AutoArk
EVA OS — A real-time multimodal AIOS for next-generation hardware, enabling your devices being “alive” and as intelligent as a real brain.
暂未收录效果图
查看技能说明EvolvingLMMs-Lab
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
暂未收录效果图
查看技能说明TIGER-AI-Lab
This repo contains the code for "VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks" [ICLR 2025]
暂未收录效果图
查看技能说明pexoai
Expert prompt engineering for Seedance 2.0. Use when the user wants to generate a video with multimodal assets (images, videos, audio) and needs the best possible prompt.
暂未收录效果图
查看技能说明shenhao-stu
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
暂未收录效果图
查看技能说明autogluon
Multi-Agent System Powered by LLMs for End-to-end Multimodal ML Automation