Skill Eval Harness

REVIEW · 62
Community indexed

Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

Verified installs0
Stars63
Version1.0.0
Quality74/100 · Strong
Trust62/100 · Sandbox only
Audit79/100 · Needs review

Supply asset profile

Coding and developer agents

Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.

Browse track

Scenario

GitHub automation

I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.

Agent fit

Claude Code + OpenAI Agents + CLI

Codex, Claude Code, Cursor, CLI, or custom agents.

Install

Ready

npx skills add adewale/skill-eval-harness

Maintenance

fresh

21d since push

Risk

Needs review

Dependency or permission surface needs review

GitHub quality

63

74/100 Quality · 70/100 Trust

Coverage tags

CodingGitHub automationutilityagent-skillskill

Review notes

Dependency or permission surface needs review · Permission surface may require sandboxing

Agent adoption scorecard

Trust, audit, and install readiness at a glance

These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.

Quality

Strong
74

Solid option that is likely worth shortlisting for production workflows.

Trust

Sandbox only
62

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

Audit

Needs review
79

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

OpenAgentSkill Trust Score v5

Human review before install

Run only in a sandbox and compare close alternatives before using it for real work.

PythonCodexClaude CodeCursorOpenAgentSkill CLI

Stars

63 GitHub stars

Repo activity

63 stars, 5 forks

Maintenance

21d since push

License

MIT

Install

npx skills add adewale/skill-eval-harness

Install safety

dynamic command execution, standard package or runtime install path

Permission surface

secrets or environment access, shell or command execution

Agent outcomes

No agent outcome data yet

Docs

Strong README/SKILL.md context

Risk summary

Review before production

  • Quality score needs review
  • Permission surface needs review: secrets or environment access, shell or command execution
  • GitHub adoption: 63 GitHub stars
  • Stars/forks activity: 63 stars, 5 forks; issue activity unavailable in current metadata

Install readiness

Install path available

  • Install path is available
  • Repository evidence is available
  • License is declared
  • No Agent Proven outcome evidence yet

Agent-readable metadata

Machine-readable decision data for this skill.

Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.

Open JSON

Suited tasks

  • Research agents workflows
  • Claude Code teams
  • builders willing to evaluate younger projects
  • Search sources

Suited agents

PythonCodexClaude CodeCursorOpenAgentSkill CLIOpenAI AgentsCLI

Install decision

Command
npx skills add adewale/skill-eval-harness
Policy
block
Human review
yes

Trust and risk

Trust
62/100
Audit
79/100
Risk level
Needs review

Outcome loop

Endpoint
/api/agent/outcome
Event ID
resolve
Outcomes
5

Install command

npx skills add adewale/skill-eval-harness

Do not use when

  • teams that need a vendor-supported SLA
  • high-compliance environments without internal security review
  • No major risk signals from current metadata
  • High-risk permission hints: Shell or command execution, Secrets or environment access
  • Dependency or permission surface needs review

Agent safety v2

39/100 · Avoid automatic install

Blocked for auto-installblock

This skill should not be selected by an agent without explicit human security review.

Do not auto-install. Inspect the source, dependencies, and permission surface first.

Resolve via API

high

Shell or command execution

Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.

medium

Network access

Skill likely fetches remote pages, APIs, repositories, or external services.

medium

Filesystem access

Skill may read or write project files, documents, generated artifacts, or local workspace state.

high

Secrets or environment access

Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.

  • High-risk permission hints: Shell or command execution, Secrets or environment access
  • Dependency or permission surface needs review

Install targets

Install this skill in your agent workflow

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install adewale-skill-eval-harness

Agent resolve plan

Let an agent verify fit before installing.

The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.

Open text plan

Agent should check

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Copy prompt

Task: Use Skill Eval Harness in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20Skill%20Eval%20Harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/adewale-skill-eval-harness/install
Install command: npx skills add adewale/skill-eval-harness
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent handoff

Give an agent the install path, not another directory page.

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

Open install API

Agent prompt

Use Skill Eval Harness for this task. Review https://www.openagentskill.com/api/skills/adewale-skill-eval-harness/install, then install with: npx skills add adewale/skill-eval-harness

Registry metadata

Agent-readable profile for automatic skill selection.

This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.

Open manifest

Agent fit

73/100

Research agents

Platforms

Python, Claude Code, OpenAI Agents

Audit report

Needs review · 79/100

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

View audit reportView eval report

Agent decision cockpit

Companion skill for Research agents

Shortlist this skill and compare it with close alternatives before production adoption.

73
Readiness
Shortlist
Stage

Role in stack

Companion skill

Primary fit

Research agents

Trust label

Strong shortlist

Install path

Command ready

Use when

  • Research agents workflows
  • Claude Code teams
  • builders willing to evaluate younger projects

Evidence

  • recent repository activity
  • install command or GitHub repo available
  • 74/100 quality profile
  • 1 OpenAgentSkill engagement events

review first

  • No major risk signals from current metadata

Implementation path

  1. 1Install it in a sandbox agent and run one Research agents task end to end.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Trust profile

Sandbox only

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

62
OpenAgentSkill Trust Score

GitHub adoption

CHECK

63 GitHub stars

Stars/forks activity

CHECK

63 stars, 5 forks; issue activity unavailable in current metadata

Recent maintenance

PASS

21d since push

License clarity

PASS

MIT

Good signals

  • AI review approved
  • Install path is available
  • Repository evidence is available
  • Recently maintained repository
  • Install command has no obvious high-risk pattern
  • Outcome loop is ready but needs first real agent run

Review before install

  • Quality score needs review
  • Permission surface needs review: secrets or environment access, shell or command execution
  • GitHub adoption: 63 GitHub stars
  • Stars/forks activity: 63 stars, 5 forks; issue activity unavailable in current metadata
  • Dependency/runtime risk: command execution surface, credential or environment access
  • Permission surface: secrets or environment access, shell or command execution
  • No real agent outcome reports yet
  • Human review required before unattended installation

Recommended action

Run only in a sandbox and compare close alternatives before using it for real work.

Quality profile

Strong candidate for agent workflows

Solid option that is likely worth shortlisting for production workflows.

74
GitHub stars
63
Freshness
21d ago
Install ready
Yes
License
MIT

Workflow fit

Use this skill in these scenarios

Workflow fit

Add it to a complete workflow

Alternative shortlist

Compare before you install

Similar skills that may fit this task.

Compare all

Overview

# Skill Eval Harness

[![CI](https://github.com/adewale/skill-eval-harness/actions/workflows/ci.yml/badge.svg)](https://github.com/adewale/skill-eval-harness/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)

Skill Eval Harness is a Python CLI that measures the **causal lift** of an Agent Skill: it runs the same case, model, and repetition with and without the skill, validates that exact experimental identity, then reports what changed, what passed, and whether the eval leaked its own answer. It reads `evals/shared-benchmark.json`, emits answer-key-safe task rows, grades files under `eval-runs/` locally and deterministically — no model call in the grade path — and writes benchmark reports you can diff across variants.

General eval frameworks (openai/evals, vitest-evals, viteval) score one output against a rubric. This one measures the *difference the skill makes*, and spends its surface area on keeping that difference honest: paired with/without comparison, `tune`/`holdout`/`holdback` split discipline, leakage lint, materialized ablations with provenance gates, and per-model lift. None of those frameworks have them, and they are what make a reported number trustworthy rather than merely green.

## Questions this helps answer

| Question | Command/report to use | |---|---| | Does this skill improve outputs compared with no skill at all? | `prepare` paired `with_skill` / `without_skill` rows, then `benchmark` paired lift and significance. | | Which prompts improved, regressed, saturated, or showed no lift? | `benchmark` `case_flags`, `render-viewer`, and `error-analysis`. | | Is the skill worth its extra tokens or dollars? | `profile-skill`, `token-overhead`, `cost-summary`, and lift-per-dollar summaries. | | Did my latest skill edit introduce a regression? | Re-run the same manifest, inspect `ablation_regressions`, `trend`, and `render-viewer --previous-workspace`. | | Which instruction, checklist, reference, scr

Platform compatibility

pythonFULL

Technical details

Version
1.0.0
License
MIT
Last updated
Aug 18, 2026
Published
Jul 29, 2026

Frameworks & tools

Python

Decision snapshot

Companion skill

73
Ready
Shortlist
Stage

recent repository activity

Audit

Install review

Install and adoption review

79
Needs review
Security
75/100
Maintenance
100/100
Install
92/100
Open full auditView eval report

Agent-proven evidence

Agent-proven evidence

Outcome reports after resolve, review, install, and one narrow run.

0
Proven
Needs first agent runAuto-install: review firstLast: Unknown
Success rate
Recent failure
Outcomes
0
Output quality
Failed
0
Not relevant
0
Installs
0
Risk blocked
0
Setup needed
0
Production
0

No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.

Install

Add to agent workflow

Free and open source. Review the report before installing into production agents.

Growth loop

Share kit

X

Scenario-led draft for Skill Eval Harness, ready for a manual X post.

Curator note
Skill Eval Harness: Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters

63 stars

https://www.openagentskill.com/skills/adewale-skill-eval-harness?ref=x
Open X draft
Optional reply with install command
Listing + install path for Skill Eval Harness:
https://www.openagentskill.com/skills/adewale-skill-eval-harness?ref=x

Install: npx skills add adewale/skill-eval-harness

Listing source

Community indexed

Claimable

This listing was indexed from public sources and is not marked official until a maintainer claim is approved.

Creator
adewale
Indexed by
OpenAgentSkill community index

Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.

Claim this skill

Owner claim

Claim this skill listing

This Community indexed listing is attributed to adewale but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.

Creator backlink kit

Add the evidence badges to your README

Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/adewale-skill-eval-harness?metric=listed&label=Listed)](https://www.openagentskill.com/skills/adewale-skill-eval-harness)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/adewale-skill-eval-harness?metric=trust&label=Trust)](https://www.openagentskill.com/skills/adewale-skill-eval-harness)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/adewale-skill-eval-harness?metric=audit&label=Audit)](https://www.openagentskill.com/skills/adewale-skill-eval-harness/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/adewale-skill-eval-harness?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/adewale-skill-eval-harness)

Author

A

adewale

@adewale

Health signals

GitHub stars
63
Quality score
45/100
Last GitHub push
Aug 1, 2026
Framework hints
1
OpenAgentSkill views
1
Install copies
0
Outbound clicks
0

Community signal

Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.

Trust & safety

Sandbox only

62
  • GitHub adoption63 GitHub starsCHECK
  • Stars/forks activity63 stars, 5 forks; issue activity unavailable in current metadataCHECK
  • Recent maintenance21d since pushPASS
  • License clarityMITPASS
  • README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
  • Dependency/runtime riskcommand execution surface, credential or environment accessCHECK