A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
Directorio de skills
Descubre skills reutilizables para AI agents.
Cada recomendación conserva un vínculo claro con su repositorio, auditoría y ruta de instalación.
Resultados de búsqueda: mllm
Directorio en inglésCambrian-1 is a family of multimodal LLMs with a vision-centric design.
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Cambrian-S: Towards Spatial Supersensing in Video
This is the repo for the paper "OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use" (ACL 2025 Oral).