Stepfun Vision Skill
Codex Skill:让纯文本模型(DeepSeek)借助 StepFun step-3.7-flash 获得看图能力 | Give text-only Codex models (DeepSeek) image understanding via StepFun step-3.7-flash
Supply asset profile
Coding and developer agents
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
Coding agents
I need a coding agent that can understand a repository, edit code, and review pull requests.
Agent fit
Claude Code + OpenAI Agents + Cursor
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add jwangkun/stepfun-vision-skill
Maintenance
fresh
19d since push
Risk
Needs review
Low GitHub adoption signal
GitHub quality
14
66/100 Quality · 76/100 Trust
Coverage tags
Review notes
Low GitHub adoption signal · Quality score needs review
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
14 GitHub stars
Repo activity
14 stars, 1 forks
Maintenance
19d since push
License
MIT
Install
npx skills add jwangkun/stepfun-vision-skill
Install safety
standard package or runtime install path
Permission surface
shell or command execution, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Review before production
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 14 GitHub stars
- Stars/forks activity: 14 stars, 1 forks; issue activity unavailable in current metadata
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- Multimodal media workflows
- Claude Code teams
- builders willing to evaluate younger projects
- Read media metadata
Suited agents
Install decision
- Command
- npx skills add jwangkun/stepfun-vision-skill
- Policy
- review
- Human review
- yes
Trust and risk
- Trust
- 68/100
- Audit
- 80/100
- Risk level
- Needs review
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add jwangkun/stepfun-vision-skillDo not use when
- teams that need a vendor-supported SLA
- production agents without a repository review
- Low GitHub adoption signal
- No OpenAgentSkill engagement data yet
- High-risk permission hints: Shell or command execution
Agent safety v2
52/100 · Avoid automatic install
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Shell or command execution
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Filesystem access
Skill may read or write project files, documents, generated artifacts, or local workspace state.
- High-risk permission hints: Shell or command execution
- Low GitHub adoption signal
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install jwangkun-stepfun-vision-skillAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20Stepfun%20Vision%20Skill%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20Stepfun%20Vision%20Skill%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/jwangkun-stepfun-vision-skill/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use Stepfun Vision Skill in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20Stepfun%20Vision%20Skill%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/jwangkun-stepfun-vision-skill/install
Install command: npx skills add jwangkun/stepfun-vision-skill
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/jwangkun-stepfun-vision-skill/install
LLM text format
/api/skills/jwangkun-stepfun-vision-skill/install?format=text
Find alternatives
/api/skills/search?q=Stepfun%20Vision%20Skill&limit=3
Agent prompt
Use Stepfun Vision Skill for this task. Review https://www.openagentskill.com/api/skills/jwangkun-stepfun-vision-skill/install, then install with: npx skills add jwangkun/stepfun-vision-skillRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/jwangkun-stepfun-vision-skill
LLM text
/api/registry/manifest/jwangkun-stepfun-vision-skill?format=text
Install alias
/api/registry/install/jwangkun-stepfun-vision-skill
Recommend
/api/registry/recommend?task=Use%20Stepfun%20Vision%20Skill%20in%20an%20agent%20workflow&limit=3
Agent fit
Multimodal media
Use-case tags
Platforms
JavaScript, Claude Code, OpenAI Agents, Cursor
Audit report
Needs review · 80/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Fallback candidate for Multimodal media
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Multimodal media
Trust label
Prototype first
Install path
Command ready
Use when
- Multimodal media workflows
- Claude Code teams
- builders willing to evaluate younger projects
Evidence
- recent repository activity
- install command or GitHub repo available
- 66/100 quality profile
review first
- Low GitHub adoption signal
- No OpenAgentSkill engagement data yet
Implementation path
- 1Install it in a sandbox agent and run one Multimodal media task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Sandbox only
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
FIX14 GitHub stars
Stars/forks activity
FIX14 stars, 1 forks; issue activity unavailable in current metadata
Recent maintenance
PASS19d since push
License clarity
PASSMIT
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 14 GitHub stars
- Stars/forks activity: 14 stars, 1 forks; issue activity unavailable in current metadata
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Promising candidate for agent workflows
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Use this skill in these scenarios
Process rich media
Multimodal media
I need my agent to process images, video, or audio and extract useful information.
Build and ship code
Coding agents
I need a coding agent that can understand a repository, edit code, and review pull requests.
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Workflow fit
Add it to a complete workflow
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
Khazix Skills
A collection of practical, installable AI agent skills for disk cleanup, AI news retrieval, and project management, following the Agent Skills standard.
Awesome Claude Skills
A curated list of resources and tools for enhancing Claude AI workflows.
Claude Scientific Skills
A comprehensive collection of ready-to-use scientific and research skills for AI agents.
Overview
# stepfun-vision-skill
   
给 **纯文本主模型(如 Codex 接入的 DeepSeek `deepseek-v4-flash`)** 外挂“看图”能力的 Codex Skill: 当主模型不支持图片输入时,把图片交给 **StepFun `step-3.7-flash`(原生多模态推理模型)** 转成文字描述,主模型再基于描述继续推理、写代码。
``` 用户粘贴图片 / 本地图片 / 图片 URL │ ▼ scripts/describe-image.js(零依赖 Node.js) 1. 守卫:读 ~/.codex/config.toml,仅主模型为 deepseek-v4-* 时启用 2. --latest:从 Codex 会话文件恢复用户粘贴的图片(base64 重建) 3. 调 StepFun /v1/chat/completions(step-3.7-flash,OpenAI 兼容) │ ▼ 文字描述 → 返回给主模型继续干活 ```
## 项目简介 / About
**为什么需要它?** 很多人把 Codex 接到 **DeepSeek(`deepseek-v4-flash` 等纯文本模型)** 上使用——便宜、快、中文好,但这类模型**不支持图片输入**:粘贴截图会显示 `image content omitted because you do not support image input`,看不了 UI 稿、报错截图、图表和白板照片。
**怎么解决?** 本 Skill 把图片交给 **StepFun `step-3.7-flash`**(原生多模态推理模型,OpenAI 兼容接口)转成精准的文字描述,再让 DeepSeek 基于描述继续推理、写代码、排障。视觉模型只负责“看”和“转录”,**结论仍由主模型得出**——成本低、接入快、不挑主模型。
**核心亮点:**
- 粘贴即用:自动从 Codex 会话文件恢复你粘贴的图片(`--latest`),无需手动保存 - 支持本地文件、图片 URL、多张图、带问题识别(OCR / 细节追问) - 零依赖 Node.js;一键安装器自动把 key 写入 `config.json`,可选写入环境变量 - Provider 守卫:只在 DeepSeek 主模型下启用,避免误用浪费 - 推理模型适配:`reasoning_effort` 控制思考开销 + `content`/`reasoning` 回退,保证必出结果
**适用场景:** UI/前端还原、报错截图排查、设计稿评审、图表/白板转数据、票据/文档 OCR、网页截图分析等。
> **English**: *stepfun-vision-skill adds image understanding to text-only coding models (e.g. DeepSeek `deepseek-v4-flash` inside Codex) by relaying images to StepFun's natively multimodal `step-3.7-flash` model. The vision model transcribes/describes the image into text; the main model then reasons on that text. Zero-dependency Node.js, one-command installer, DeepSeek-only provider guard, and `--latest` recovery of pasted images from Codex session files.*
> 本仓库是 [deepseek-vision-skill](https://github.com/iuiaeng2005/deepseek-vision-skill)(MIT)的 StepFun 适配版:换了默认接口/
Platform compatibility
Technical details
- Version
- 1.0.0
- License
- MIT
- Last updated
- Aug 18, 2026
- Published
- Aug 6, 2026
Frameworks & tools
Decision snapshot
Fallback candidate
recent repository activity
Audit
Install review
Install and adoption review
- Security
- 86/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for Stepfun Vision Skill, ready for a manual X post.
A practical pick for design or creative work: Stepfun Vision Skill: Codex Skill:让纯文本模型(DeepSeek)借助 StepFun step-3.7-flash 获得看图能力 | Give text-only Codex models (DeepSeek) image understanding v... 14 stars https://www.openagentskill.com/skills/jwangkun-stepfun-vision-skill?ref=x
Optional reply with install command
Listing + install path for Stepfun Vision Skill: https://www.openagentskill.com/skills/jwangkun-stepfun-vision-skill?ref=x Install: npx skills add jwangkun/stepfun-vision-skill
Listing source
Community indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- jwangkun
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Community indexed listing is attributed to jwangkun but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/jwangkun-stepfun-vision-skill)
[](https://www.openagentskill.com/skills/jwangkun-stepfun-vision-skill)
[](https://www.openagentskill.com/skills/jwangkun-stepfun-vision-skill/audit)
[](https://www.openagentskill.com/skills/jwangkun-stepfun-vision-skill)Author
jwangkun
@jwangkun
Platform fit
Health signals
- GitHub stars
- 14
- Quality score
- 39/100
- Last GitHub push
- Aug 3, 2026
- Framework hints
- 1
- OpenAgentSkill views
- 0
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Sandbox only
- GitHub adoption14 GitHub starsFIX
- Stars/forks activity14 stars, 1 forks; issue activity unavailable in current metadataFIX
- Recent maintenance19d since pushPASS
- License clarityMITPASS
- README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
- Dependency/runtime riskno major dependency risk hints in public metadataPASS
Related skills
Khazix Skills
A collection of practical, installable AI agent skills for disk cleanup, AI news retrieval, and project management, following the Agent Skills standard.
19.6K StarsAwesome Claude Skills
A curated list of resources and tools for enhancing Claude AI workflows.
65.9K StarsClaude Scientific Skills
A comprehensive collection of ready-to-use scientific and research skills for AI agents.
31.2K Stars