A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
Skill 디렉토리
AI Agent를 위한 재사용 가능한 Skill을 찾으세요.
모든 추천은 리포지토리, 감사, 설치 경로와 명확하게 연결됩니다.
검색 결과: mllm
영문 디렉토리Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Cambrian-S: Towards Spatial Supersensing in Video
This is the repo for the paper "OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use" (ACL 2025 Oral).