Skill Eval Harness
Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters
Asset-Profil
Coding- und Entwickler-Agents
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Szenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent-Fit
Claude Code + OpenAI Agents + CLI
Geeignet für Codex, Claude Code, Cursor, CLI oder benutzerdefinierte Agents.
Installieren
Bereit
npx skills add adewale/skill-eval-harness
Wartung
Aktuell
22 Tage seit dem letzten Push
Risiko
Prüfung nötig
Dependency or permission surface needs review
GitHub-Qualität
63
74/100 Qualität · 70/100 Vertrauen
Abdeckungs-Tags
Review-Notizen
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent-Adoptionskarte
Vertrauen, Audit und Installationsbereitschaft auf einen Blick
Diese Werte kombinieren öffentliche Repository-Metadaten, OpenAgentSkill-Reviewsignale, Wartungsaktualität und Installationsbereitschaft. Sie helfen bei der Vorauswahl, ersetzen aber keine menschliche Prüfung.
Qualität
StarkSolid option that is likely worth shortlisting for production workflows.
Vertrauen
Nur SandboxNützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
Audit
Prüfung nötigMaschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
OpenAgentSkill Trust Score v5
Menschliche Prüfung vor Installation
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Stars
63 GitHub-Stars
Repository-Aktivität
63 Stars und 5 Forks
Wartung
22 Tage seit dem letzten Push
Lizenz
MIT
Installieren
npx skills add adewale/skill-eval-harness
Installationssicherheit
dynamic command execution, standard package or runtime install path
Berechtigungsfläche
secrets or environment access, shell or command execution
Agent-Ergebnisse
Noch keine Agent-Ergebnisdaten
Dokumentation
Starker README/SKILL.md-Kontext
Risikoübersicht
Vor Produktion prüfen
- Quality score needs review
- Permission surface needs review: secrets or environment access, shell or command execution
- GitHub adoption: 63 GitHub stars
- Stars/forks activity: 63 stars, 5 forks; issue activity unavailable in current metadata
Installationsbereitschaft
Installationspfad verfügbar
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Lizenz ist angegeben
- Noch keine Agent-Proven-Ergebnisbelege
Agent-lesbare Metadaten
Maschinenlesbare Entscheidungsdaten für diesen Skill.
Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.
Geeignete Aufgaben
- Research-Agent-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
- Suchquellen
Geeignete Agents
Installationsentscheidung
- Befehl
- npx skills add adewale/skill-eval-harness
- Richtlinie
- Blockieren
- Menschliche Prüfung
- Ja
Vertrauen und Risiko
- Vertrauen
- 62/100
- Audit
- 79/100
- Risikoebene
- Prüfung nötig
Ergebnis-Loop
- Endpoint
- /api/agent/outcome
- Event-ID
- resolve
- Ergebnisse
- 5
Installationsbefehl
npx skills add adewale/skill-eval-harnessNicht verwenden, wenn
- Teams, die ein vom Anbieter unterstütztes SLA benötigen
- Hochregulierte Umgebungen ohne interne Sicherheitsprüfung
- No major risk signals from current metadata
- Hinweise auf Hochrisiko-Berechtigungen: Shell or command execution, Secrets or environment access
- Dependency or permission surface needs review
Agent-Sicherheit v2
39/100 · Automatische Installation vermeiden
This skill should not be selected by an agent without explicit human security review.
Do not auto-install. Inspect the source, dependencies, and permission surface first.
Hoch
Shell- oder Befehlsausführung
Die Skill-Metadaten verweisen auf Terminal-, CLI-, Shell-, Subprozess- oder Befehlsausführungs-Workflows.
Mittel
Netzwerkzugriff
Die Skill ruft wahrscheinlich Remote-Seiten, APIs, Repositories oder externe Dienste ab.
Mittel
Dateisystemzugriff
Die Skill kann Projektdateien, Dokumente, generierte Artefakte oder den lokalen Arbeitsbereich lesen oder schreiben.
Hoch
Secrets or environment access
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
- Hinweise auf Hochrisiko-Berechtigungen: Shell or command execution, Secrets or environment access
- Dependency or permission surface needs review
Installationsziele
Diesen Skill im Agent-Workflow installieren
Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install adewale-skill-eval-harnessAgent-Auflösungsplan
Lass einen Agent die Eignung vor der Installation prüfen.
Die Resolve API liefert die beste Skill, Alternativen, Sicherheitsrichtlinien, Auditnotizen, Installationsziel und einen direkt nutzbaren Prompt.
JSON öffnen
/api/agent/resolve?task=Use%20Skill%20Eval%20Harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve-Text
/api/agent/resolve?task=Use%20Skill%20Eval%20Harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Installationsübergabe
/api/skills/adewale-skill-eval-harness/install
Agent sollte prüfen
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Prompt kopieren
Task: Use Skill Eval Harness in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20Skill%20Eval%20Harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/adewale-skill-eval-harness/install
Install command: npx skills add adewale/skill-eval-harness
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent-Übergabe
Gib dem Agent den Installationspfad, nicht noch ein Verzeichnis.
Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.
Installationsübergabe
/api/skills/adewale-skill-eval-harness/install
LLM-Textformat
/api/skills/adewale-skill-eval-harness/install?format=text
Alternativen finden
/api/skills/search?q=Skill%20Eval%20Harness&limit=3
Agent-Prompt
Use Skill Eval Harness for this task. Review https://www.openagentskill.com/api/skills/adewale-skill-eval-harness/install, then install with: npx skills add adewale/skill-eval-harnessRegistry-Metadaten
Agent-lesbares Profil für die automatische Skill-Auswahl.
Die Registry API stellt Entscheidungs-, Vertrauens-, Audit-, Use-Case- und Installationssignale ohne UI-Scraping bereit.
Manifest
/api/registry/manifest/adewale-skill-eval-harness
LLM-Text
/api/registry/manifest/adewale-skill-eval-harness?format=text
Installationsalias
/api/registry/install/adewale-skill-eval-harness
Empfehlen
/api/registry/recommend?task=Use%20Skill%20Eval%20Harness%20in%20an%20agent%20workflow&limit=3
Agent-Fit
Recherche-Agents
Plattformen
Python, Claude Code, OpenAI Agents
Audit-Bericht
Prüfung nötig · 79/100
Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
Agent-Entscheidungspanel
Companion skill for Research agents
Shortlist this skill and compare it with close alternatives before production adoption.
Rolle im Stack
Ergänzende Skill
Primäre Eignung
Recherche-Agents
Vertrauenslabel
Starke Shortlist
Installationspfad
Befehl bereit
Verwenden wenn
- Research-Agent-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
Evidenz
- recent repository activity
- install command or GitHub repo available
- Qualitätsprofil 74/100
- 1 OpenAgentSkill-Interaktionen
zuerst prüfen
- No major risk signals from current metadata
Implementierungspfad
- 1Installieren Sie es in einem Sandbox-Agent und führen Sie eine Recherche-Agents-Aufgabe vollständig aus.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Vertrauensprofil
Nur Sandbox
Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
GitHub-Akzeptanz
Prüfen63 GitHub-Stars
Star-/Fork-Aktivität
Prüfen63 Stars und 5 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbar
Aktuelle Wartung
Bestanden22 Tage seit dem letzten Push
Lizenzklarheit
BestandenMIT
Positive Signale
- KI-Prüfung genehmigt
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Kürzlich gewartetes Repository
- Der Installationsbefehl weist kein offensichtliches Hochrisikomuster auf
- Ergebniszyklus ist bereit, benötigt aber den ersten echten Agent-Lauf
Vor Installation prüfen
- Quality score needs review
- Permission surface needs review: secrets or environment access, shell or command execution
- GitHub adoption: 63 GitHub stars
- Stars/forks activity: 63 stars, 5 forks; issue activity unavailable in current metadata
- Dependency/runtime risk: command execution surface, credential or environment access
- Permission surface: secrets or environment access, shell or command execution
- Noch keine echten Agent-Ergebnisberichte
- Vor unbeaufsichtigter Installation ist menschliche Prüfung erforderlich
Empfohlene Aktion
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Qualitätsprofil
Stark Kandidat für Agent-Workflows
Solid option that is likely worth shortlisting for production workflows.
Workflow-Eignung
Diese Skill in diesen Szenarien nutzen
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Manage repositories
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Workflow-Eignung
Zum vollständigen Workflow hinzufügen
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternativen-Shortlist
Vor Installation vergleichen
Similar skills that may fit this task.
Khazix Skills
A collection of practical, installable AI agent skills for disk cleanup, AI news retrieval, and project management, following the Agent Skills standard.
Awesome Claude Skills
A curated list of resources and tools for enhancing Claude AI workflows.
Claude Scientific Skills
A comprehensive collection of ready-to-use scientific and research skills for AI agents.
Übersicht
# Skill Eval Harness
[](https://github.com/adewale/skill-eval-harness/actions/workflows/ci.yml) [](LICENSE)
Skill Eval Harness is a Python CLI that measures the **causal lift** of an Agent Skill: it runs the same case, model, and repetition with and without the skill, validates that exact experimental identity, then reports what changed, what passed, and whether the eval leaked its own answer. It reads `evals/shared-benchmark.json`, emits answer-key-safe task rows, grades files under `eval-runs/` locally and deterministically — no model call in the grade path — and writes benchmark reports you can diff across variants.
General eval frameworks (openai/evals, vitest-evals, viteval) score one output against a rubric. This one measures the *difference the skill makes*, and spends its surface area on keeping that difference honest: paired with/without comparison, `tune`/`holdout`/`holdback` split discipline, leakage lint, materialized ablations with provenance gates, and per-model lift. None of those frameworks have them, and they are what make a reported number trustworthy rather than merely green.
## Questions this helps answer
| Question | Command/report to use | |---|---| | Does this skill improve outputs compared with no skill at all? | `prepare` paired `with_skill` / `without_skill` rows, then `benchmark` paired lift and significance. | | Which prompts improved, regressed, saturated, or showed no lift? | `benchmark` `case_flags`, `render-viewer`, and `error-analysis`. | | Is the skill worth its extra tokens or dollars? | `profile-skill`, `token-overhead`, `cost-summary`, and lift-per-dollar summaries. | | Did my latest skill edit introduce a regression? | Re-run the same manifest, inspect `ablation_regressions`, `trend`, and `render-viewer --previous-workspace`. | | Which instruction, checklist, reference, scr
Plattformkompatibilität
Technische Details
- Version
- 1.0.0
- Lizenz
- MIT
- Letzte Aktualisierung
- 18. Aug. 2026
- Veröffentlicht
- 29. Juli 2026
Frameworks & Tools
Entscheidungsübersicht
Ergänzende Skill
recent repository activity
Audit
Installationsprüfung
Installations- und Adoptionsprüfung
- Sicherheit
- 75/100
- Wartung
- 100/100
- Installieren
- 92/100
Von Agent belegte Evidenz
Von Agent belegte Evidenz
Ergebnisberichte nach Resolve, Prüfung, Installation und einem begrenzten Lauf.
- Erfolgsrate
- —
- Letzter Fehler
- —
- Ergebnisse
- 0
- Ausgabequalität
- —
- Fehlgeschlagen
- 0
- Nicht relevant
- 0
- Installationen
- 0
- Durch Risiko blockiert
- 0
- Einrichtung erforderlich
- 0
- Produktion
- 0
Noch keine Agent-Ergebnisdaten. Der erste Lauf kann Erfolg, Einrichtungsbedarf, Risikoblockaden, Fehler oder Irrelevanz über /api/agent/outcome melden.
Installieren
Zum Agent-Workflow hinzufügen
Kostenlos und Open Source. Bericht vor der Installation in Produktions-Agents prüfen.
Wachstums-Loop
Share-Kit
Szenariobasierter Entwurf für Skill Eval Harness, bereit für einen manuellen X-Post.
Skill Eval Harness: Agent Skill evaluation harness for paired variants, trace artifacts, and runner adapters 63 stars https://www.openagentskill.com/skills/adewale-skill-eval-harness?ref=x
Optionale Antwort mit Installationsbefehl
Listing + install path for Skill Eval Harness: https://www.openagentskill.com/skills/adewale-skill-eval-harness?ref=x Install: npx skills add adewale/skill-eval-harness
Quelle des Eintrags
Community-indexiert
Dieser Eintrag wurde aus öffentlichen Quellen indexiert und ist erst nach Genehmigung eines Maintainer-Anspruchs offiziell.
- Ersteller
- adewale
- Indexiert von
- OpenAgentSkill Community-Index
Die Zuordnung verlinkt auf das öffentliche Repository oder Creator-Profil. Creator können den Eintrag beanspruchen, um Eigentümersignale zu aktualisieren.
Diesen Skill beanspruchenEigentümeranspruch
Diesen Skill-Eintrag beanspruchen
Dieser Community-indexiert-Eintrag wird adewale zugeschrieben, ist aber noch nicht offiziell markiert. Beanspruche ihn, um ein verifiziertes Eigentümersignal hinzuzufügen und künftige Launch-, Installations- und Audit-Updates vertrauenswürdiger zu machen.
Creator-Backlink-Kit
Evidenz-Badges in deine README einfügen
Zeige den kanonischen Eintrag, aktuelle Vertrauens- und Audit-Signale sowie echte Agent-Proven-Evidenz dort, wo Entwickler das Repository bewerten.
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness)
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness)
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness/audit)
[](https://www.openagentskill.com/skills/adewale-skill-eval-harness)Autor
adewale
@adewale
Plattform-Fit
Gesundheitssignale
- GitHub-Stars
- 63
- Qualitätswert
- 45/100
- Letzter GitHub-Push
- 1. Aug. 2026
- Framework-Hinweise
- 1
- OpenAgentSkill-Aufrufe
- 1
- Installationskopien
- 0
- Externe Klicks
- 0
Community-Signal
Teile mit, ob dieser Skill für deinen Agent-Workflow nützlich ist. Zusammengefasstes Feedback verbessert das Ranking im Laufe der Zeit.
Vertrauen & Sicherheit
Nur Sandbox
- GitHub-Akzeptanz63 GitHub-StarsPrüfen
- Star-/Fork-Aktivität63 Stars und 5 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbarPrüfen
- Aktuelle Wartung22 Tage seit dem letzten PushBestanden
- LizenzklarheitMITBestanden
- README/SKILL.md-VollständigkeitMetadaten enthalten ausreichend Nutzungs- und Workflow-KontextBestanden
- Abhängigkeits-/Laufzeitrisikocommand execution surface, credential or environment accessPrüfen
Ähnliche Skills
Khazix Skills
A collection of practical, installable AI agent skills for disk cleanup, AI news retrieval, and project management, following the Agent Skills standard.
19.6K StarsAwesome Claude Skills
A curated list of resources and tools for enhancing Claude AI workflows.
65.9K StarsClaude Scientific Skills
A comprehensive collection of ready-to-use scientific and research skills for AI agents.
31.2K Stars