Pdf Inspector
firecrawl
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
33โ48 / 65
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 65
firecrawl
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
CIRCL
AIL framework - Analysis Information Leak framework. Project moved to https://github.com/ail-project
unjs
๐ PDF extraction and rendering across all JavaScript runtimes
AndyTheFactory
๐ฐ Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.
soxoj
โ๏ธ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
openai
Automate a real browser from the terminal for navigation, form filling, snapshots, screenshots, extraction, and UI-flow debugging.
ankidroid
AnkiDroid: Anki flashcards on Android. Your secret trick to achieve superhuman information retention.
LAION-AI
OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
LLMQuant
QuantMind is an intelligent knowledge extraction and retrieval framework for quantitative finance.
zaidmukaddam
Scira (Formerly MiniPerplx) is a minimalistic AI-powered search engine that helps you find information on the internet and cites it too. Powered by Vercel AI SDK!
adbar
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
external-secrets
External Secrets Operator reads information from a third-party service like AWS Secrets Manager and automatically injects the values as Kubernetes Secrets.
grobidOrg
A machine learning software for extracting information from scholarly documents