Skill audit report
Use when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing at all. Also use when the user mentions evaluation, eval, benchmarking, testing LLM quality, measuring agent performance, assessing RAG accuracy, or wants to compare prompts/models. This skill is the entry router: it asks diagnostic questions then recommends which sub-skill (local workflow) to use next.
OpenAgentSkill Trust Score
The Trust Score helps an agent decide whether a skill is safe enough to shortlist before installation.
GitHub adoption
INFO76
809 GitHub stars
Stars/forks activity
INFO71
809 stars, 65 forks; issue activity unavailable in current metadata
Recent maintenance
PASS88
1mo since push
License clarity
PASS86
Apache-2.0
README/SKILL.md completeness
PASS86
Metadata includes enough usage and workflow context
Dependency/runtime risk
PASS90
no major dependency risk hints in public metadata
Install availability
PASS92
npx skills add agentscope-ai/OpenJudge --skill meta-eval
Install command safety
INFO68
dynamic command execution, standard package or runtime install path
Permission surface
PASS100
no high-risk permission surface in public metadata
Repository evidence
PASS86
https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/00-meta-eval
Review status
PASS88
AI review data available
Agent Proven outcomes
INFO54
No agent outcome data yet
Checks
Install path
92
npx skills add agentscope-ai/OpenJudge --skill meta-eval
Repository
88
https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/00-meta-eval
License
86
Apache-2.0
Maintenance
88
1mo since push
AI review
88
Approved with no listed issues
README/SKILL.md completeness
86
Usable description available
Dependency risk
90
no major dependency risk hints in public metadata
Install command safety
68
dynamic command execution, standard package or runtime install path
Permission surface
100
no high-risk permission surface in public metadata
Stars/forks activity
71
809 stars, 65 forks; issue activity unavailable in current metadata
Adoption
88
809 GitHub stars
Financial decision safety
58
Research-only use: do not treat output as financial advice or execute a position without human approval.
Warnings
Method
This report combines public metadata, AI review output, repository freshness, install readiness, OpenAgentSkill events, quality scoring, trust checks, and the agent safety gate. It is not a full source-code security review.
Compare nearby options
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
175K Stars · Audit report
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
85K Stars · Audit report
Turn one topic into a narrated Vox-style paper-collage explainer or ad video, from script through captions.
1.8K Stars · Audit report