CLIP
openai
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Search retrieves candidates across the registry and ranks a bounded shortlist by task fit. This count is matching candidates, not the registry total. No suitable match? Try a specific tool or task.
33β48 / 95
Results: 95
openai
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
explosion
π« Industrial-strength Natural Language Processing (NLP) in Python
jianchang512
Translate the video from one language to another and embed dubbing & subtitles.
meta-llama
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to eβ¦
PaddlePaddle
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System,β¦
dataelement
BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI workflow, RAG, Agent, Unified mβ¦
windingwind
Translate PDF, EPub, webpage, metadata, annotations, notes to the target language. Support 20+ translate services.
datalab-to
OCR model that handles complex tables, forms, handwriting with full layout.
GetStream
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
Pluviobyte
Render, preview, validate, and burn consistent production subtitles from quality-checked timestamp artifacts.
roboflow
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
enricoros
AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas, AGI functions, world-class Beam multi-model chats, text-to-imaβ¦
leejet
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
Eventual-Inc
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
remsky
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.