No visual example yet
Explore the skillCambrian
cambrian-mllm
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
OPENAGENTSKILL / DIRECTORY
Finde den passenden Skill für deine nächste Aufgabe mit Codex, Claude Code, Cursor und mehr.
5 Skills
Ergebnisse: 5
No visual example yet
Explore the skillcambrian-mllm
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
No visual example yet
Explore the skillcambrian-mllm
Cambrian-S: Towards Spatial Supersensing in Video
No visual example yet
Explore the skillAIDC-AI
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
No visual example yet
Explore the skillbytedance
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
No visual example yet
Explore the skillOS-Agent-Survey
This is the repo for the paper "OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use" (ACL 2025 Oral).