暂未收录效果图
查看技能说明OmAgent
om-ai-lab
[EMNLP-2024] Build multimodal language agents for fast prototype and production
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
84 Skills
搜索结果: 84
暂未收录效果图
查看技能说明om-ai-lab
[EMNLP-2024] Build multimodal language agents for fast prototype and production
暂未收录效果图
查看技能说明declare-lab
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
暂未收录效果图
查看技能说明unum-cloud
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
暂未收录效果图
查看技能说明Tencent-Hunyuan
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
暂未收录效果图
查看技能说明RQLuo
MixTeX multimodal LaTeX, ZhEn, and, Table OCR. It performs efficient CPU-based inference in a local offline on Windows.
暂未收录效果图
查看技能说明TIGER-AI-Lab
Official Repo for "TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding" [ACL 2025 oral]
暂未收录效果图
查看技能说明mims-harvard
Therapeutics Commons (TDC): Multimodal Foundation for Therapeutic Science
暂未收录效果图
查看技能说明go-kratos
Blades is a Go-based multimodal AI Agent framework.
暂未收录效果图
查看技能说明Haervwe
Open‑WebUI Tools is a modular toolkit designed to extend and enrich your Open WebUI instance, turning it into a powerful AI workstation. With a suite of over 15 speciali…
暂未收录效果图
查看技能说明lancedb
Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs
暂未收录效果图
查看技能说明framersai
Build autonomous AI agents with adaptive intelligence and emergent behaviors, included with multimodal RAG and optional HEXACO personalities.
暂未收录效果图
查看技能说明HorizonWind2004
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models…
暂未收录效果图
查看技能说明mbzuai-oryx
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with…
暂未收录效果图
查看技能说明foxglove
Multimodal visualization and data platform
暂未收录效果图
查看技能说明duoan
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
暂未收录效果图
查看技能说明berabuddies
Multimodal document deep analysis tool based on Zhipu GLM-OCR, GLM-4.7, and GLM-4.6V.