News Please
fhamborg
news-please - an integrated web crawler and information extractor for news that just works
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
Page 19 · 16 shown · 1,601 public entries
Results: 1601
fhamborg
news-please - an integrated web crawler and information extractor for news that just works
eikek
Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
DerwenAI
Python implementation of TextRank algorithms ("textgraphs") for phrase extraction
modesty
converts binary PDF to JSON and text, for server-side PDF processing and command-line use. Zero dependency.
sirfz
A Python wrapper for the tesseract-ocr API
TimmyOVO
Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python.
zorlan
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
adithya-s-k
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks
NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
lukas-blecher
pix2tex: Using a ViT to convert images of equations into LaTeX code.
salomonelli
:necktie: :briefcase: Build fast :rocket: and easy multiple beautiful resumes and create your best CV ever! Made with Vue and LESS.
extractus
To extract main article from given URL with Node.js
alvarobartt
Financial Data Extraction from Investing.com with Python
Lulzx
Minimal PDF creation library. <400 LOC, zero dependencies, makes real PDFs.
AlibabaResearch
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab,…
camelot-dev
A web interface to extract tabular data from PDFs