Skill Eval Harness
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
Supply asset profile
Coding and developer agents
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add adewale/skill-eval-harness
Maintenance
fresh
21d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
63
74/100 Quality · 70/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
63 GitHub stars
Repo activity
63 stars, 5 forks
Maintenance
21d since push
License
MIT
Install
npx skills add adewale/skill-eval-harness
Install safety
dynamic command execution, standard package or runtime install path
Permission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Review before production
- Quality score needs review
- Permission surface needs review: secrets or environment access, shell or command execution
- GitHub adoption: 63 GitHub stars
- Stars/forks activity: 63 stars, 5 forks; issue activity unavailable in current metadata
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- Research agents workflows
- Claude Code teams
- builders willing to evaluate younger projects
- Search sources
Suited agents
Install decision
- Command
- npx skills add adewale/skill-eval-harness
- Policy
- block
- Human review
- yes
Trust and risk
- Trust
- 62/100
- Audit
- 79/100
- Risk level
- Needs review
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add adewale/skill-eval-harnessDo not use when
- teams that need a vendor-supported SLA
- high-compliance environments without internal security review
- No major risk signals from current metadata
- High-risk permission hints: Shell or command execution, Secrets or environment access
- Dependency or permission surface needs review
Agent safety v2
39/100 · Avoid automatic install
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
high
Shell or command execution
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Filesystem access
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Secrets or environment access
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
- High-risk permission hints: Shell or command execution, Secrets or environment access
- Dependency or permission surface needs review
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install adewale-skill-eval-harnessAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20Skill%20Eval%20Harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20Skill%20Eval%20Harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/adewale-skill-eval-harness/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use Skill Eval Harness in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20Skill%20Eval%20Harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/adewale-skill-eval-harness/install
Install command: npx skills add adewale/skill-eval-harness
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/adewale-skill-eval-harness/install
LLM text format
/api/skills/adewale-skill-eval-harness/install?format=text
Find alternatives
/api/skills/search?q=Skill%20Eval%20Harness&limit=3
Agent prompt
Use Skill Eval Harness for this task. Review https://www.openagentskill.com/api/skills/adewale-skill-eval-harness/install, then install with: npx skills add adewale/skill-eval-harnessRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/adewale-skill-eval-harness
LLM text
/api/registry/manifest/adewale-skill-eval-harness?format=text
Install alias
/api/registry/install/adewale-skill-eval-harness
Recommend
/api/registry/recommend?task=Use%20Skill%20Eval%20Harness%20in%20an%20agent%20workflow&limit=3
Agent fit
Research agents
Use-case tags
Platforms
Python, Claude Code, OpenAI Agents
Audit report
Needs review · 79/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Companion skill for Research agents
Shortlist this skill and compare it with close alternatives before production adoption.
Role in stack
Companion skill
Primary fit
Research agents
Trust label
Strong shortlist
Install path
Command ready
Use when
- Research agents workflows
- Claude Code teams
- builders willing to evaluate younger projects
Evidence
- recent repository activity
- install command or GitHub repo available
- 74/100 quality profile
- 1 OpenAgentSkill engagement events
review first
- No major risk signals from current metadata
Implementation path
- 1Install it in a sandbox agent and run one Research agents task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Sandbox only
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
CHECK63 GitHub stars
Stars/forks activity
CHECK63 stars, 5 forks; issue activity unavailable in current metadata
Recent maintenance
PASS21d since push
License clarity
PASSMIT
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- Quality score needs review
- Permission surface needs review: secrets or environment access, shell or command execution
- GitHub adoption: 63 GitHub stars
- Stars/forks activity: 63 stars, 5 forks; issue activity unavailable in current metadata
- Dependency/runtime risk: command execution surface, credential or environment access
- Permission surface: secrets or environment access, shell or command execution
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Strong candidate for agent workflows
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Use this skill in these scenarios
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Manage repositories
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Add it to a complete workflow
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
Khazix Skills
A collection of practical, installable AI agent skills for disk cleanup, AI news retrieval, and project management, following the Agent Skills standard.
Awesome Claude Skills
A curated list of resources and tools for enhancing Claude AI workflows.
Claude Scientific Skills
A comprehensive collection of ready-to-use scientific and research skills for AI agents.
Overview
# Skill Eval Harness
[](https://github.com/adewale/skill-eval-harness/actions/workflows/ci.yml) [](LICENSE)
Skill Eval Harness is a Python CLI that measures the **causal lift** of an Agent Skill: it runs the same case, model, and repetition with and without the skill, validates that exact experimental identity, then reports what changed, what passed, and whether the eval leaked its own answer. It reads `evals/shared-benchmark.json`, emits answer-key-safe task rows, grades files under `eval-runs/` locally and deterministically — no model call in the grade path — and writes benchmark reports you can diff across variants.
General eval frameworks (openai/evals, vitest-evals, viteval) score one output against a rubric. This one measures the *difference the skill makes*, and spends its surface area on keeping that difference honest: paired with/without comparison, `tune`/`holdout`/`holdback` split discipline, leakage lint, materialized ablations with provenance gates, and per-model lift. None of those frameworks have them, and they are what make a reported number trustworthy rather than merely green.
## Questions this helps answer
| Question | Command/report to use | |---|---| | Does this skill improve outputs compared with no skill at all? | `prepare` paired `with_skill` / `without_skill` rows, then `benchmark` paired lift and significance. | | Which prompts improved, regressed, saturated, or showed no lift? | `benchmark` `case_flags`, `render-viewer`, and `error-analysis`. | | Is the skill worth its extra tokens or dollars? | `profile-skill`, `token-overhead`, `cost-summary`, and lift-per-dollar summaries. | | Did my latest skill edit introduce a regression? | Re-run the same manifest, inspect `ablation_regressions`, `trend`, and `render-viewer --previous-workspace`. | | Which instruction, checklist, reference, scr
Platform compatibility
Technical details
- Version
- 1.0.0
- License
- MIT
- Last updated
- Aug 18, 2026
- Published
- Jul 29, 2026
Frameworks & tools
Decision snapshot
Companion skill
recent repository activity
Audit
Install review
Install and adoption review
- Security
- 75/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for Skill Eval Harness, ready for a manual X post.
Skill Eval Harness: Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters 63 stars https://www.openagentskill.com/skills/adewale-skill-eval-harness?ref=x
Optional reply with install command
Listing + install path for Skill Eval Harness: https://www.openagentskill.com/skills/adewale-skill-eval-harness?ref=x Install: npx skills add adewale/skill-eval-harness
Listing source
Community indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- adewale
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Community indexed listing is attributed to adewale but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness)
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness)
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness/audit)
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness)Author
adewale
@adewale
Platform fit
Health signals
- GitHub stars
- 63
- Quality score
- 45/100
- Last GitHub push
- Aug 1, 2026
- Framework hints
- 1
- OpenAgentSkill views
- 1
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Sandbox only
- GitHub adoption63 GitHub starsCHECK
- Stars/forks activity63 stars, 5 forks; issue activity unavailable in current metadataCHECK
- Recent maintenance21d since pushPASS
- License clarityMITPASS
- README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
- Dependency/runtime riskcommand execution surface, credential or environment accessCHECK
Related skills
Khazix Skills
A collection of practical, installable AI agent skills for disk cleanup, AI news retrieval, and project management, following the Agent Skills standard.
19.6K StarsAwesome Claude Skills
A curated list of resources and tools for enhancing Claude AI workflows.
65.9K StarsClaude Scientific Skills
A comprehensive collection of ready-to-use scientific and research skills for AI agents.
31.2K Stars