Run multimodal agents that operate desktop interfaces
Skill 디렉토리
AI Agent를 위한 재사용 가능한 Skill을 찾으세요.
모든 추천은 리포지토리, 감사, 설치 경로와 명확하게 연결됩니다.
검색 결과: bytedance
영문 디렉토리An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
FlowGram is an extensible workflow development framework with built-in canvas, form, variable, and materials that helps developers build AI workflow platforms faster and simpler.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
Official Repo For Pixel-LLM Codebase: Sa2VA (Arxiv-25), SAMTok (CVPR-26), VRT, SaSaSa2VA (1-st solution for LSVOS)
SALMONN family: A suite of advanced multi-modal LLMs
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
Two Claude Skills that turn agents into AI film directors, providing cinematic dramaturgy and exact prompt syntax for major video/image models.
Appshark is a static taint analysis platform to scan vulnerabilities in an Android app.
A Claude Code custom skill for generating structured Chinese prompts for ByteDance's Seedance 2.0 AI video generation platform.
🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning