暂未收录效果图
查看技能说明Lance
bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
5 Skills
搜索结果: 5
暂未收录效果图
查看技能说明bytedance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
暂未收录效果图
查看技能说明nvidia-cosmos
Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.
暂未收录效果图
查看技能说明FunAudioLLM
Multilingual speech understanding: ASR + emotion recognition + audio event detection. 50+ languages, 15x faster than Whisper, non-autoregressive.
暂未收录效果图
查看技能说明unum-cloud
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
暂未收录效果图
查看技能说明menyifang
Official implementation of "MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling"