Unisondb
ankur-anand
A streaming multimodal database for Edge AI, and Edge Computing.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
65–80 / 84
Results: 84
ankur-anand
A streaming multimodal database for Edge AI, and Edge Computing.
jaketae
Multimodal AI Story Teller, built with Stable Diffusion, GPT, and neural text-to-speech
bowang-lab
BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model | NeurIPS '25
kohjingyu
🧀 Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs".
kohjingyu
🐟 Code and models for the NeurIPS 2023 paper "Generating Images with Multimodal Language Models".
opengeos
A multimodal AI agent for geospatial data analysis and interactive visualization
JIA-Lab-research
This project is the official implementation of 'LLMGA: Multimodal Large Language Model based Generation Assistant', ECCV2024 Oral
WisconsinAIVision
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
idwts
🦦 Crayotter: A Multimodal AI-Agent for Video-Editing, Video-Composing, and Video Production. Powered by Multimodal LLMs for autonomous Text-to-Video agentic framework.…
wanshuiyin
Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal generati…
veniceai
Call POST /chat/completions on Venice. Covers the OpenAI-compatible request shape, Venice-only venice_parameters (web search, E2EE, characters, thinking control, X searc…
firebase
Official skill for integrating Firebase AI Logic (Gemini API) into web applications. Covers setup, multimodal inference, structured output, and security.
Jason904
Design banners for social media, ads, website heroes, creative assets, and print. Multiple art direction options with AI-generated visuals. Actions: design, create, gene…
mbzuai-oryx
[CVPR 2025 🔥]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
DjangoPeng
This repository is a hub for AI Agent projects, including GitHub Sentinel, LanguageMentor, and ChatPPT, designed to enhance enterprise workflows, language learning, and…
sunshine-lang
从多模态素材到选题、核验、脚本、配图和小红书发布文案的 Codex Skill