HunyuanImage 3.0
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–9 / 9
Results: 9
Tencent-Hunyuan
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
xiincs
为 Claude Code 赋能多模态视觉能力,支持豆包、通义千问、GPT-4o 等模型,用于截图 / UI / 图表分析;适配 DeepSeek 等无视觉底座,搭配 browser-harness 可做前端布局自动化检查。
jina-ai
☁️ Build multimodal AI applications with cloud-native stack
JIA-Lab-research
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
kohjingyu
🧀 Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs".
kohjingyu
🐟 Code and models for the NeurIPS 2023 paper "Generating Images with Multimodal Language Models".
opengeos
A multimodal AI agent for geospatial data analysis and interactive visualization
DjangoPeng
This repository is a hub for AI Agent projects, including GitHub Sentinel, LanguageMentor, and ChatPPT, designed to enhance enterprise workflows, language learning, and…
adityasoni9998
Code for the paper "Coding Agents with Multimodal Browsing are Generalist Problem Solvers"