DeepKE
zjunlp
[EMNLP 2022] An Open Toolkit for Knowledge Graph Extraction and Construction
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–16 / 46
Results: 46
zjunlp
[EMNLP 2022] An Open Toolkit for Knowledge Graph Extraction and Construction
pymupdf
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
firecrawl
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
CatchTheTornado
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any docume…
NanoNets
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extract…
ispras
Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta info…
adbar
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
landing-ai
This tool has been deprecated. Use Agentic Document Extraction instead.
opendataloader-project
Procedure for extracting structured data from PDFs with opendataloader-pdf (ODL) correctly: read the installed tool's own help to discover its current options, build the…
landing-ai
Python library for Agentic Document Extraction (ADE).
codelucas
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
getmaxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
Achno
A tool to convert a Wallpaper's color scheme / palette, OCR with VLM's Traditional & Hybrid, Image Compression ,color palette extraction, image upsacling with Adversaria…
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
unjs
📄 PDF extraction and rendering across all JavaScript runtimes
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.