Skill audit report
Use when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. Also use when the user mentions prompt A/B testing, prompt comparison, prompt optimization validation, "did my prompt change help," or prompt regression testing. Outputs per-dimension win rates with statistical significance using OpenJudge PairwiseAnalyzer.
OpenAgentSkill Trust Score
The Trust Score helps an agent decide whether a skill is safe enough to shortlist before installation.
GitHub adoption
INFO76
816 GitHub stars
Stars/forks activity
INFO71
816 stars, 65 forks; issue activity unavailable in current metadata
Recent maintenance
PASS88
1mo since push
License clarity
PASS86
Apache-2.0
README/SKILL.md completeness
PASS86
Metadata includes enough usage and workflow context
Dependency/runtime risk
INFO72
command execution surface
Install availability
PASS92
npx skills add agentscope-ai/OpenJudge --skill prompt-regression
Install command safety
PASS92
standard package or runtime install path
Permission surface
INFO64
shell or command execution, database access
Repository evidence
PASS86
https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression
Review status
INFO66
AI review data available
Agent Proven outcomes
INFO54
No agent outcome data yet
Checks
Install path
92
npx skills add agentscope-ai/OpenJudge --skill prompt-regression
Repository
88
https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression
License
86
Apache-2.0
Maintenance
88
1mo since push
AI review
55
The full implementation of scripts/pairwise.py is not visible in the excerpt, but the provided portion indicates a well-structured, stdlib-only script with clear usage and exit codes.
README/SKILL.md completeness
86
Usable description available
Dependency risk
72
command execution surface
Install command safety
92
standard package or runtime install path
Permission surface
64
shell or command execution, database access
Stars/forks activity
71
816 stars, 65 forks; issue activity unavailable in current metadata
Adoption
88
816 GitHub stars
Warnings
Method
This report combines public metadata, AI review output, repository freshness, install readiness, OpenAgentSkill events, quality scoring, trust checks, and the agent safety gate. It is not a full source-code security review.
Compare nearby options
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
169K Stars · Audit report
Platform to build admin panels, internal tools, and dashboards. Integrates with 25+ databases and any API.
41K Stars · Audit report
Implement work from an approved spec or ticket set, run focused and full tests, invoke code review, and commit the result to the current branch.
176K Stars · Audit report