A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
Annuaire de skills
Découvrez des skills réutilisables pour les AI agents.
Chaque recommandation reste clairement reliée à son dépôt, son audit et son chemin d’installation.
Résultats de recherche: mllm
Annuaire en anglaisCambrian-1 is a family of multimodal LLMs with a vision-centric design.
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Cambrian-S: Towards Spatial Supersensing in Video
This is the repo for the paper "OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use" (ACL 2025 Oral).