Skill audit report
Use when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining multiple metrics into a composite score, or building an automated evaluation pipeline. Also use when the user mentions grader selection, metric design, judge prompt engineering, rubric design, evaluation pipeline code, or "how to evaluate [X] automatically." Outputs executable OpenJudge pipeline code.
OpenAgentSkill Trust Score
The Trust Score helps an agent decide whether a skill is safe enough to shortlist before installation.
GitHub adoption
INFO76
816 GitHub stars
Stars/forks activity
INFO71
816 stars, 65 forks; issue activity unavailable in current metadata
Recent maintenance
PASS88
1mo since push
License clarity
PASS86
Apache-2.0
README/SKILL.md completeness
PASS86
Metadata includes enough usage and workflow context
Dependency/runtime risk
INFO72
external package install surface, database surface
Install availability
PASS92
npx skills add agentscope-ai/OpenJudge --skill metric-design
Install command safety
PASS92
standard package or runtime install path
Permission surface
PASS88
database access
Repository evidence
PASS86
https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/02-metric-design
Review status
INFO66
AI review data available
Agent Proven outcomes
INFO54
No agent outcome data yet
Checks
Install path
92
npx skills add agentscope-ai/OpenJudge --skill metric-design
Repository
88
https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/02-metric-design
License
86
Apache-2.0
Maintenance
88
1mo since push
AI review
55
The SKILL.md excerpt is truncated; the full document may contain additional details, but the provided content is coherent and actionable.
README/SKILL.md completeness
86
Usable description available
Dependency risk
72
external package install surface, database surface
Install command safety
92
standard package or runtime install path
Permission surface
88
database access
Stars/forks activity
71
816 stars, 65 forks; issue activity unavailable in current metadata
Adoption
88
816 GitHub stars
Financial decision safety
58
Research-only use: do not treat output as financial advice or execute a position without human approval.
Warnings
Method
This report combines public metadata, AI review output, repository freshness, install readiness, OpenAgentSkill events, quality scoring, trust checks, and the agent safety gate. It is not a full source-code security review.
Compare nearby options
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
175K Stars · Audit report
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
85K Stars · Audit report
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K Stars · Audit report