prompt-regression Eval ====================== Status: failed Score: 69/100 Risk: high Decision: do_not_auto_install Policy: block Reason: Audit score: Risky Install: npx skills add agentscope-ai/OpenJudge --skill prompt-regression Required checks: - PASS Task fit: Task wording matches this skill metadata. - PASS Install path: Install handoff is available. - PASS Install command safety: standard package or runtime install path - WARN Trust score: Good trust signals with a few areas worth checking before rollout. - FAIL Audit score: Risky - FAIL Agent safety gate: This skill should not be selected by an agent without explicit human security review. - PASS License clarity: Apache-2.0 - WARN Permission surface: shell or command execution, database access Warnings: - Trust score: Good trust signals with a few areas worth checking before rollout. - Permission surface: shell or command execution, database access - Audit risk risky exceeds max_risk=medium - High-risk permission hints: Shell or command execution - Potential broker, wallet, exchange, or real-money execution surface; sandbox and explicit approval are required - The full implementation of scripts/pairwise.py is not visible in the excerpt, but the provided portion indicates a well-structured, stdlib-only script with clear usage and exit codes. - The skill relies on the user to correctly generate the comparisons.jsonl file with swapped rows; no automated validation is provided for that input format. - This skill may touch real-money trading, broker, wallet, or exchange operations; use only in a sandbox with explicit approval. - Quality score needs review Validation plan: 1. Inspect repository, README/SKILL.md, license, and recent commits before production use. 2. Install in an isolated workspace or sandbox with no production secrets available. 3. Run the smallest representative task and record files touched, commands run, network access, and outputs. 4. Compare the selected skill against at least one alternative when the eval status is review or failed. 5. Promote only after the agent reports a successful verification result and unresolved warnings are accepted. Do not use when: - teams that need a vendor-supported SLA - production agents without a repository review - The full implementation of scripts/pairwise.py is not visible in the excerpt, but the provided portion indicates a well-structured, stdlib-only script with clear usage and exit codes. - No OpenAgentSkill engagement data yet - Audit risk risky exceeds max_risk=medium - High-risk permission hints: Shell or command execution - Potential broker, wallet, exchange, or real-money execution surface; sandbox and explicit approval are required - The skill relies on the user to correctly generate the comparisons.jsonl file with swapped rows; no automated validation is provided for that input format. URLs: - Skill: https://www.openagentskill.com/skills/agentscope-ai-prompt-regression - Audit: https://www.openagentskill.com/skills/agentscope-ai-prompt-regression/audit - JSON: https://www.openagentskill.com/api/agent/evals?slug=agentscope-ai-prompt-regression