技能目录

为 AI Agent 发现可复用技能。

按任务搜索真实的 GitHub 技能,并在使用前查看 Stars、信任、审计、分类和安装路径。

每个推荐都保留与其仓库、审计和安装路径的明确关联。

搜索结果: paddleocr-vl

英文目录

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

83K
Stars
86/100
信任
分类: document-processing审计

A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge models like RF-DETR, YOLO11, SAM 3, and Qwen3-VL.

9.5K
Stars
76/100
信任
分类: ml-automation审计

Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.

4.3K
Stars
80/100
信任
分类: rag-knowledge审计

基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.

1.8K
Stars
83/100
信任
分类: document-processing审计

PaddleOCR inference in PyTorch. Converted from [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)

1.2K
Stars
84/100
信任
分类: document-processing审计

Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python.

2.2K
Stars
73/100
信任
分类: document-processing审计

OCR离线图片文字识别命令行windows程序,以JSON字符串形式输出结果,方便别的程序调用。提供各种语言API。由 PaddleOCR C++ 编译。

1.5K
Stars
70/100
信任
分类: document-processing审计

A reusable skill for Claude Code that adds multimodal vision capabilities for analyzing screenshots, UI, and charts using various models.

92
Stars
67/100
信任
分类: coding-agents审计

ConvMAE: Masked Convolution Meets Masked Autoencoders

531
Stars
63/100
信任
分类: robotics-iot审计

“Dive Into OCR” is a textbook developed by the PaddleOCR community that integrates OCR theory and practice.

258
Stars
63/100
信任
分类: document-processing审计