给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding,…
QualityHigh-confidence pick with strong adoption and healthy maintenance signals.
TrustGood trust signals with a few areas worth checking before rollout.Review: Documentation summary is thin
Safety gateGood audit and safety signals with no high-risk permission hints in public metadata.
Scenario Document processing · I need my agent to read PDFs, extract tables, and turn documents into structured data.
Claude Code + OpenAI Agents · 4 targets