pdf-extract
Fast, zero-AI text extraction from PDFs that have a text layer (digitally created PDFs from Word, Typst, WeasyPrint, wkhtmltopdf, LaTeX, etc). Uses pymupdf (fitz) - instant and deterministic. Use when you need to quickly pull raw text from a known text-layer PDF, e.g. "extract te
供給アセットの概要
デザインとクリエイティブ制作
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
シナリオ
Multimodal media
I need my agent to process images, video, or audio and extract useful information.
Agent 適合
Claude Code + CLI + Codex
Codex、Claude Code、Cursor、CLI、またはカスタム Agent に対応します。
インストール
準備完了
npx skills add henkisdabro/wookstar-claude-plugins --skill pdf-extract
メンテナンス
新しい
最終プッシュから 2 日
リスク
要レビュー
Dependency or permission surface needs review
GitHub 品質
83
66/100 品質 · 71/100 信頼
対象タグ
レビュー注記
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent 導入スコアカード
信頼、監査、インストール準備状況を一目で確認
公開リポジトリのメタデータ、OpenAgentSkill のレビューシグナル、保守の鮮度、インストール準備状況を組み合わせたスコアです。候補選定の目安であり、人によるレビューの代替ではありません。
品質
有望有用な候補ですが、採用前に代替と比較してください。
信頼
サンドボックス限定信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。
監査
要レビューインストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。
OpenAgentSkill Trust Score v5
インストール前に人のレビュー
実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。
スター
GitHub スター 83
リポジトリ活動
スター 83、フォーク 11
メンテナンス
最終プッシュから 2 日
ライセンス
MIT
インストール
npx skills add henkisdabro/wookstar-claude-plugins --skill pdf-extract
インストール安全性
標準パッケージまたはランタイムのインストールパス
権限範囲
secrets or environment access, shell or command execution
Agent の成果
Agent の成果データはまだありません
ドキュメント
README/SKILL.md の文脈が十分です
リスク概要
本番前にレビュー
- Quality score needs review
- Permission surface needs review: secrets or environment access, shell or command execution
- GitHub adoption: 83 GitHub stars
- Stars/forks activity: 83 stars, 11 forks; issue activity unavailable in current metadata
インストール準備状況
インストールパスを利用可能
- インストールパスを利用できます
- リポジトリの根拠を利用できます
- ライセンスが明示されています
- Agent-Proven の成果エビデンスはまだありません
Agent 可読メタデータ
このスキルの機械可読な判断データ。
このブロックまたは埋め込み JSON を使い、Agent がこのスキルをインストールすべきか、代替を選ぶべきか、先に人のレビューを求めるべきかを判断できます。
適したタスク
- Document processing ワークフロー
- Claude Code チーム
- builders willing to evaluate younger projects
- Read uploaded files
適した Agent
インストール判断
- コマンド
- npx skills add henkisdabro/wookstar-claude-plugins --skill pdf-extract
- ポリシー
- ブロック
- 人によるレビュー
- はい
信頼とリスク
- 信頼
- 63/100
- 監査
- 77/100
- リスクレベル
- 要レビュー
成果ループ
- エンドポイント
- /api/agent/outcome
- イベント ID
- resolve
- 成果
- 5
インストールコマンド
npx skills add henkisdabro/wookstar-claude-plugins --skill pdf-extract使わない場合
- ベンダー提供の SLA が必要なチーム
- 内部セキュリティレビューのない高コンプライアンス環境
- 現在のメタデータに重大なリスクシグナルはありません
- 高リスク権限のヒント: Shell or command execution, Secrets or environment access
- Dependency or permission surface needs review
代替スキル
Frontend Design
171.1K スター
npx skills add anthropics/skills --skill frontend-design
代替スキル
Taste Skill: Anti-Slop Frontend
79.4K スター
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
代替スキル
Canvas Design
171.1K スター
npx skills add anthropics/skills --skill canvas-design
代替スキル
Anthropic Brand Guidelines
171.1K スター
npx skills add anthropics/skills --skill brand-guidelines
Agent セーフティ v2
33/100 · 自動インストールを避ける
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
高
Shell またはコマンド実行
Skill メタデータに端末、CLI、Shell、サブプロセス、またはコマンド実行のワークフローが含まれます。
中
Browser automation
Skill may drive a browser or interact with web pages.
中
ネットワークアクセス
Skill はリモートページ、API、リポジトリ、外部サービスにアクセスする可能性があります。
中
ファイルシステムアクセス
Skill はプロジェクトファイル、ドキュメント、生成物、ローカルワークスペース状態を読み書きする可能性があります。
- 高リスク権限のヒント: Shell or command execution, Secrets or environment access
- Dependency or permission surface needs review
インストール先
Agent ワークフローにこのスキルをインストール
公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install henkisdabro-pdf-extractAgent 解決プラン
インストール前に Agent に適合性を検証させます。
Resolve API は第一候補、代替、安全ポリシー、監査メモ、インストール先、Agent がそのまま使えるプロンプトを返します。
JSON を開く
/api/agent/resolve?task=Use%20pdf-extract%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve テキスト
/api/agent/resolve?task=Use%20pdf-extract%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
インストール引き継ぎ
/api/skills/henkisdabro-pdf-extract/install
Agent が確認すべきこと
- Resolve API でタスク適合と代替を確認。
- 監査・信頼スコアと安全ポリシーの警告を確認。
- Codex、Claude Code、Cursor、CLI のインストール先互換性を確認。
プロンプトをコピー
Task: Use pdf-extract in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20pdf-extract%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/henkisdabro-pdf-extract/install
Install command: npx skills add henkisdabro/wookstar-claude-plugins --skill pdf-extract
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent 引き継ぎ
別のディレクトリではなく、インストール経路を Agent に渡します。
公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。
インストール引き継ぎ
/api/skills/henkisdabro-pdf-extract/install
LLM テキスト形式
/api/skills/henkisdabro-pdf-extract/install?format=text
代替を探す
/api/skills/search?q=pdf-extract&limit=3
Agent プロンプト
Use pdf-extract for this task. Review https://www.openagentskill.com/api/skills/henkisdabro-pdf-extract/install, then install with: npx skills add henkisdabro/wookstar-claude-plugins --skill pdf-extractRegistry メタデータ
自動スキル選択用の Agent 可読プロファイル。
Registry API 経由で判断、信頼、監査、ユースケース、インストールのシグナルを提供し、UI をスクレイピングせずに Agent が順位付けできます。
Manifest
/api/registry/manifest/henkisdabro-pdf-extract
LLM テキスト
/api/registry/manifest/henkisdabro-pdf-extract?format=text
インストール別名
/api/registry/install/henkisdabro-pdf-extract
推奨
/api/registry/recommend?task=Use%20pdf-extract%20in%20an%20agent%20workflow&limit=3
Agent 適合
Document processing
プラットフォーム
Claude Code
Agent 判断パネル
Fallback candidate for Document processing
まずこのスキルでプロトタイプを作り、代替候補を用意してください。
スタック内の役割
代替候補
主な適合
Document processing
信頼ラベル
まずプロトタイプ
インストールパス
コマンド準備済み
使う場面
- Document processing ワークフロー
- Claude Code チーム
- builders willing to evaluate younger projects
根拠
- 最近のリポジトリ活動
- インストールコマンドまたは GitHub リポジトリが利用可能
- 品質プロファイル 66/100
- OpenAgentSkill エンゲージメント 5 件
先にレビュー
- 現在のメタデータに重大なリスクシグナルはありません
実装パス
- 1サンドボックスの Agent にインストールし、Document processing タスクを一度最初から最後まで実行します。
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
信頼プロファイル
サンドボックス限定
信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。
GitHub 採用度
確認GitHub スター 83
スター/フォーク活動
確認スター 83、フォーク 11; 現在のメタデータでは Issue 活動を利用できません
最近のメンテナンス
合格最終プッシュから 2 日
ライセンスの明確さ
合格MIT
良いシグナル
- AI レビュー承認済み
- インストールパスを利用できます
- リポジトリの根拠を利用できます
- 最近保守されたリポジトリ
- インストールコマンドに明確な高リスクパターンはありません
- 成果ループは準備済みですが、最初の実行が必要です
インストール前にレビュー
- Quality score needs review
- Permission surface needs review: secrets or environment access, shell or command execution
- GitHub adoption: 83 GitHub stars
- Stars/forks activity: 83 stars, 11 forks; issue activity unavailable in current metadata
- Dependency/runtime risk: command execution surface, credential or environment access
- Permission surface: secrets or environment access, shell or command execution
- 実際の Agent 成果レポートはまだありません
- 無人インストールの前に人によるレビューが必要です
推奨アクション
実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。
品質プロファイル
有望 Agent ワークフロー向けの候補
有用な候補ですが、採用前に代替と比較してください。
ワークフロー適合
このスキルを使うシナリオ
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Process rich media
Multimodal media
I need my agent to process images, video, or audio and extract useful information.
Collect structured data
Web scraping
I need my agent to scrape websites and extract structured data from pages.
ワークフロー適合
完全なワークフローに追加
Scrape, clean, and reuse web data
Web data pipeline
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Ingest, retrieve, and cite
RAG knowledge base
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
代替候補
インストール前に比較
このタスクに適する可能性のある類似スキル。
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Taste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Canvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
Anthropic Brand Guidelines
Apply Anthropic official brand colors, typography, and visual standards to appropriate Anthropic-related artifacts.
概要
--- name: pdf-extract description: Fast, zero-AI text extraction from PDFs that have a text layer (digitally created PDFs from Word, Typst, WeasyPrint, wkhtmltopdf, LaTeX, etc). Uses pymupdf (fitz) - instant and deterministic. Use when you need to quickly pull raw text from a known text-layer PDF, e.g. "extract text from this PDF", "read this PDF", "get the content of", "what does this PDF say", "quickly read this PDF". Do NOT use for scanned/image PDFs or when you need structured output (tables, headings, OCR, AI analysis) - use the pdf-processing-pro skill in this plugin for those cases. allowed-tools: Bash, Read, Write ---
# PDF Text Extraction
Extract text from PDF files using pymupdf via `uv run --with pymupdf`.
## Prerequisites
- `uv` installed - see <https://docs.astral.sh/uv/getting-started/installation/> - No venv or pre-installation needed - `uv run --with` handles caching automatically
## Extract text from a single PDF
```bash uv run --with pymupdf python3 -c " import fitz doc = fitz.open('/path/to/file.pdf') for page in doc: text = page.get_text().strip() if text: print(text) print() " ```
## Extract and save to file
```bash uv run --with pymupdf python3 -c " import fitz
doc = fitz.open('/path/to/file.pdf') pages = [] for page in doc: text = page.get_text().strip() if text: pages.append(text)
with open('/path/to/output.txt', 'w') as f: f.write('\n\n'.join(pages))
print(f'Extracted {len(pages)} pages') " ```
## Extract specific pages
```bash uv run --with pymupdf python3 -c " import fitz
doc = fitz.open('/path/to/file.pdf') # Pages are 0-indexed for i in range(2, 5): # Pages 3-5 text = doc[i].get_text().strip() if text: print(text) " ```
## Batch extract from multiple PDFs
```bash uv run --with pymupdf python3 -c " import fitz import glob import os
for pdf_path in glob.glob('/path/to/folder/*.pdf'): doc = fitz.open(pdf_path) text = '\n\n'.join(p.get_text().strip() for p in doc if p.get_text().strip()) out_path = pdf_path.rsplit('.', 1)[0] + '.txt' with open(out_path, 'w') as f: f.write(text) print(f'{os.path.basename(pdf_path)}: {len(doc)} pages extracted') " ```
## Get PDF metadata
```bash uv run --with pymupdf python3 -c " import fitz doc = fitz.open('/path/to/file.pdf') meta = doc.metadata print(f'Title: {meta.get(\"title\", \"N/A\")}') print(f'Author: {meta.get(\"author\", \"N/A\")}') print(f'Pages: {len(doc)}') print(f'Creator: {meta.get(\"creator\", \"N/A\")}') " ```
## Key notes
- pymupdf is imported as `fitz` (legacy naming from the MuPDF library) - Pages are 0-indexed: `doc[0]` is the first page - `get_text()` returns plain text; use `get_text("blocks")` for positioned blocks - `get_text("html")` returns HTML with formatting preserved - The package caches after the first `uv run --with pymupdf` invocation - subsequent runs are instant
## When to use this vs pdf-processing-pro
| Use pdf-extract | Use pdf-processing-pro | |---|---| | PDF created digitally (Word, Typst, LaTeX, wkhtmltopdf, WeasyPrint) | Scanned or image-based PDF (photo, fax, scan) | | Need raw text quickly - less than a second | Need structured output: tables, headings, forms | | Bulk/batch extraction without AI cost | OCR required (scanned documents) | | Offline, no API key, no extra dependencies | Form filling, validation, batch workflows | | Simple text content, no tables needed | Tables or structured layout are important |
技術詳細
- バージョン
- 1.0.0
- ライセンス
- MIT
- 最終更新
- 2026年8月21日
- 公開日
- 2026年8月21日
判断の要約
代替候補
最近のリポジトリ活動
Agent 実証エビデンス
Agent 実証エビデンス
Resolve、レビュー、インストール、限定実行後の成果レポート。
- 成功率
- —
- 直近の失敗
- —
- 成果
- 0
- 出力品質
- —
- 失敗
- 0
- 非該当
- 0
- インストール数
- 0
- リスクによりブロック
- 0
- 設定が必要
- 0
- 本番
- 0
Agent の実行結果はまだありません。最初の実行では /api/agent/outcome を通じて成功、設定要件、リスクによるブロック、失敗、非該当を報告できます。
成長ループ
共有キット
pdf-extract 用のシナリオベース草案です。X へ手動投稿できます。
pdf-extract: Fast, zero-AI text extraction from PDFs that have a text layer (digitally created PDFs from W... 83 stars https://www.openagentskill.com/skills/henkisdabro-pdf-extract?ref=x
任意:インストールコマンド付きの返信
Listing + install path for pdf-extract: https://www.openagentskill.com/skills/henkisdabro-pdf-extract?ref=x Install: npx skills add henkisdabro/wookstar-claude-plugins --skill pdf-extract
掲載元
Registry により登録
この掲載は公開ソースから登録されており、メンテナー申請が承認されるまで公式として表示されません。
- 作成者
- henkisdabro
- インデックス作成者
- OpenAgentSkill コミュニティインデックス
帰属は公開リポジトリまたは作成者プロフィールにリンクされています。作成者は掲載を申請して所有権シグナルを更新できます。
このスキルを申請所有者の申請
このスキル掲載を申請
この Registry により登録 掲載は henkisdabro に帰属していますが、まだ公式として表示されていません。申請すると、確認済み所有者シグナルが追加され、今後の公開、インストール、監査更新の信頼性が高まります。
クリエイター被リンクキット
README にエビデンスバッジを追加
開発者がリポジトリを評価する場所で、正規掲載、現在の信頼・監査シグナル、実際の Agent-Proven エビデンスを表示します。
[](https://www.openagentskill.com/skills/henkisdabro-pdf-extract)
[](https://www.openagentskill.com/skills/henkisdabro-pdf-extract)
[](https://www.openagentskill.com/skills/henkisdabro-pdf-extract/audit)
[](https://www.openagentskill.com/skills/henkisdabro-pdf-extract)作者
henkisdabro
@henkisdabro
プラットフォーム適合
健全性シグナル
- GitHub スター
- 83
- 品質スコア
- 37/100
- 最終 GitHub プッシュ
- 2026年8月21日
- フレームワークのヒント
- 不明
- OpenAgentSkill 閲覧数
- 5
- インストールコピー数
- 0
- 外部クリック
- 0
コミュニティシグナル
このスキルが Agent ワークフローに役立つかを共有してください。集約されたフィードバックがランキングを改善します。
信頼と安全性
サンドボックス限定
- GitHub 採用度GitHub スター 83確認
- スター/フォーク活動スター 83、フォーク 11; 現在のメタデータでは Issue 活動を利用できません確認
- 最近のメンテナンス最終プッシュから 2 日合格
- ライセンスの明確さMIT合格
- README/SKILL.md の完全性メタデータには十分な利用・ワークフロー文脈があります合格
- 依存関係/ランタイムのリスクcommand execution surface, credential or environment access確認
関連スキル
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
171.1K スターTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
79.4K スターCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
171.1K スターAnthropic Brand Guidelines
Apply Anthropic official brand colors, typography, and visual standards to appropriate Anthropic-related artifacts.
171.1K スター