Extractor
lightfeed
Use LLMs to robustly extract web data
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 18
Results: 18
lightfeed
Use LLMs to robustly extract web data
extractus
To extract main article from given URL with Node.js
YaoFANGUK
视频硬字幕提取,生成srt文件。无需申请第三方API,本地实现文本识别。基于深度学习的视频字幕提取框架,包含字幕区域检测、字幕内容提取。A GUI tool for extracting hard-coded subtitle (hardsub) from videos and generating srt files.
skylander86
AWS Lambda functions to extract text from various binary formats.
gamemaker1
Yet another library to extract text from MS Office and PDF files
soxoj
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
scambier
A (companion) plugin to facilitate the extraction of text from images (OCR) and PDFs.
ahmetozlu
A super lightweight image processing algorithm for detection and extraction of overlapped handwritten signatures on scanned documents using OpenCV and scikit-image.
nainiayoub
PDF text data extraction web app with OCR for scanned documents
fhamborg
news-please - an integrated web crawler and information extractor for news that just works
floriandiud
Facebook Group Members Extractor. Download Facebook group members in CSV.
opendatalab
MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.
StanGirard
SEO & Security Audit for Websites. Lighthouse & Security Headers crawler, Sitemap/Keywords/Images Extractor, Summarizer, etc ...
ropensci
Bindings for Tabula PDF Table Extractor Library
CeON
opensearch-project
Ingest documents at scale into Amazon OpenSearch using OpenSearch Ingestion Service (OSIS) pipelines. Upload pre-generated JSONL chunks to S3 and OSIS indexes them — opt…