create-custom-grader

Prüfen · 70
Im Registry indexiert

Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

Verified installs0
Stars187
Version1.0.0
Qualität70/100 · Stark
Vertrauen70/100 · Nur Sandbox
Audit81/100 · Prüfung nötig

Asset-Profil

Recherche und Wissensarbeit

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

Bereich ansehen

Szenario

RAG and knowledge

I need my agent to build a RAG workflow over documents and retrieve reliable context.

Agent-Fit

Claude Code + CLI + Codex

Geeignet für Codex, Claude Code, Cursor, CLI oder benutzerdefinierte Agents.

Installieren

Bereit

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

Wartung

Aktuell

1 Tage seit dem letzten Push

Risiko

Prüfung nötig

Quality score needs review

GitHub-Qualität

187

70/100 Qualität · 78/100 Vertrauen

Abdeckungs-Tags

RechercheRAG and knowledgeautomationagent-skill

Review-Notizen

Quality score needs review · Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata

Agent-Adoptionskarte

Vertrauen, Audit und Installationsbereitschaft auf einen Blick

Diese Werte kombinieren öffentliche Repository-Metadaten, OpenAgentSkill-Reviewsignale, Wartungsaktualität und Installationsbereitschaft. Sie helfen bei der Vorauswahl, ersetzen aber keine menschliche Prüfung.

Qualität

Stark
70

Solid option that is likely worth shortlisting for production workflows.

Vertrauen

Nur Sandbox
70

Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.

Audit

Prüfung nötig
81

Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.

OpenAgentSkill Trust Score v5

Menschliche Prüfung vor Installation

Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

187 GitHub-Stars

Repository-Aktivität

187 Stars und 14 Forks

Wartung

1 Tage seit dem letzten Push

Lizenz

Apache-2.0

Installieren

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

Installationssicherheit

Standard-Paket- oder Laufzeit-Installationspfad

Berechtigungsfläche

shell or command execution, filesystem or document access

Agent-Ergebnisse

Noch keine Agent-Ergebnisdaten

Dokumentation

Starker README/SKILL.md-Kontext

Risikoübersicht

Vor Produktion prüfen

  • Quality score needs review
  • Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata

Installationsbereitschaft

Installationspfad verfügbar

  • Installationspfad ist verfügbar
  • Repository-Belege sind verfügbar
  • Lizenz ist angegeben
  • Noch keine Agent-Proven-Ergebnisbelege

Agent-lesbare Metadaten

Maschinenlesbare Entscheidungsdaten für diesen Skill.

Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.

JSON öffnen

Geeignete Aufgaben

  • Browser automation-Workflows
  • Claude-Code-Teams
  • builders willing to evaluate younger projects
  • Navigate pages

Geeignete Agents

CodexClaude CodeCursorOpenAgentSkill CLICLI

Installationsentscheidung

Befehl
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Richtlinie
Prüfen
Menschliche Prüfung
Ja

Vertrauen und Risiko

Vertrauen
70/100
Audit
81/100
Risikoebene
Prüfung nötig

Ergebnis-Loop

Endpoint
/api/agent/outcome
Event-ID
resolve
Ergebnisse
5

Installationsbefehl

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

Nicht verwenden, wenn

  • Teams, die ein vom Anbieter unterstütztes SLA benötigen
  • Hochregulierte Umgebungen ohne interne Sicherheitsprüfung
  • No OpenAgentSkill engagement data yet
  • Hinweise auf Hochrisiko-Berechtigungen: Shell- oder Befehlsausführung
  • Quality score needs review

Agent-Sicherheit v2

53/100 · Automatische Installation vermeiden

ExperimentellPrüfen

Sparse or mixed signals. Useful for discovery, but not for autonomous installation.

Test manually in an isolated workspace and compare against safer alternatives.

Per API auflösen

Hoch

Shell- oder Befehlsausführung

Die Skill-Metadaten verweisen auf Terminal-, CLI-, Shell-, Subprozess- oder Befehlsausführungs-Workflows.

Mittel

Netzwerkzugriff

Die Skill ruft wahrscheinlich Remote-Seiten, APIs, Repositories oder externe Dienste ab.

Mittel

Dateisystemzugriff

Die Skill kann Projektdateien, Dokumente, generierte Artefakte oder den lokalen Arbeitsbereich lesen oder schreiben.

  • Hinweise auf Hochrisiko-Berechtigungen: Shell- oder Befehlsausführung
  • Quality score needs review

Installationsziele

Diesen Skill im Agent-Workflow installieren

Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install nvidia-create-custom-grader

Agent-Auflösungsplan

Lass einen Agent die Eignung vor der Installation prüfen.

Die Resolve API liefert die beste Skill, Alternativen, Sicherheitsrichtlinien, Auditnotizen, Installationsziel und einen direkt nutzbaren Prompt.

Textplan öffnen

Agent sollte prüfen

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Prompt kopieren

Task: Use create-custom-grader in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install
Install command: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent-Übergabe

Gib dem Agent den Installationspfad, nicht noch ein Verzeichnis.

Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.

Installations-API öffnen

Agent-Prompt

Use create-custom-grader for this task. Review https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install, then install with: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

Registry-Metadaten

Agent-lesbares Profil für die automatische Skill-Auswahl.

Die Registry API stellt Entscheidungs-, Vertrauens-, Audit-, Use-Case- und Installationssignale ohne UI-Scraping bereit.

Manifest öffnen

Agent-Fit

69/100

Browser automation

Plattformen

Claude Code

Audit-Bericht

Prüfung nötig · 81/100

Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.

Audit-Bericht ansehenEval-Bericht ansehen

Agent-Entscheidungspanel

Fallback candidate for Browser automation

Prototype with this skill first; keep a fallback candidate ready.

69
Bereitschaft
Prototyp
Phase

Rolle im Stack

Fallback-Kandidat

Primäre Eignung

Browser automation

Vertrauenslabel

Zuerst prototypisieren

Installationspfad

Befehl bereit

Verwenden wenn

  • Browser automation-Workflows
  • Claude-Code-Teams
  • builders willing to evaluate younger projects

Evidenz

  • recent repository activity
  • install command or GitHub repo available
  • Qualitätsprofil 70/100

zuerst prüfen

  • No OpenAgentSkill engagement data yet

Implementierungspfad

  1. 1Installieren Sie es in einem Sandbox-Agent und führen Sie eine Browser automation-Aufgabe vollständig aus.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Vertrauensprofil

Nur Sandbox

Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.

70
OpenAgentSkill Trust Score

GitHub-Akzeptanz

Info

187 GitHub-Stars

Star-/Fork-Aktivität

Prüfen

187 Stars und 14 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbar

Aktuelle Wartung

Bestanden

1 Tage seit dem letzten Push

Lizenzklarheit

Bestanden

Apache-2.0

Positive Signale

  • KI-Prüfung genehmigt
  • Installationspfad ist verfügbar
  • Repository-Belege sind verfügbar
  • Kürzlich gewartetes Repository
  • Der Installationsbefehl weist kein offensichtliches Hochrisikomuster auf
  • Ergebniszyklus ist bereit, benötigt aber den ersten echten Agent-Lauf

Vor Installation prüfen

  • Quality score needs review
  • Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
  • Noch keine echten Agent-Ergebnisberichte
  • Vor unbeaufsichtigter Installation ist menschliche Prüfung erforderlich

Empfohlene Aktion

Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.

Qualitätsprofil

Stark Kandidat für Agent-Workflows

Solid option that is likely worth shortlisting for production workflows.

70
GitHub-Stars
187
Aktualität
vor 1 Tagen
Installationsbereit
Ja
Lizenz
Apache-2.0

Workflow-Eignung

Diese Skill in diesen Szenarien nutzen

Workflow-Eignung

Zum vollständigen Workflow hinzufügen

Alternativen-Shortlist

Vor Installation vergleichen

Similar skills that may fit this task.

Alle vergleichen

Übersicht

--- name: create-custom-grader description: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. metadata: author: SkillEvaluator Maintainers <maintainers@example.com> ---

# Create Custom Grader

Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.

## Purpose

Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.

## When To Use

Use this skill when the user wants to:

- bring an existing benchmark into SkillEvaluator - turn a rubric into `evals/grader.py` or `evals/grader.sh` - add custom metrics beside the default evaluator metrics - convert task files such as `task.yaml`, `task.json`, pytest checks, or shell verifiers into BYOG or BYOT - prove a team can run its own benchmark through SkillEvaluator

Do not use this skill for ordinary `evals/evals.json` authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.

## Instructions

1. Read the target skill, existing `evals/`, benchmark prompts, fixtures, and any verifier code. 2. Choose `default_plus_custom` when custom metrics should complement default evaluator scoring. 3. Choose `custom_only` only when the user wants the custom grader to own pass/fail semantics. 4. Write or update `evals/grader.py` or `evals/grader.sh`, then validate the Harbor contract.

## Examples

```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom skillevaluator tier3 validate <skill-dir> ```

## Prerequisites

- The target skill directory should contain `SKILL.md`. - The SkillEvaluator CLI should be available as `skillevaluator`. - Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.

## Core Choice

Choose one path before writing files:

| User need | Evaluator shape | | --- | --- | | Existing `evals.json` task plus extra domain checks | Top-level BYOG: `evals/grader.py` or `evals/grader.sh` | | Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus `evals/evals.json` and `evals/files/` | | Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: `evals/harbor/<case>/...` | | User wants only custom reward/pass criteria | `grading.mode: custom_only` | | User wants default evaluator dimensions plus custom metrics | `grading.mode: default_plus_custom` |

Default to `default_plus_custom` unless the user explicitly wants the custom grader to replace the default evaluator metrics.

## Workflow

1. Resolve the target skill and benchmark source. Read the target `SKILL.md`, existing `evals/`, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.

2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as `question` entries. Use the target skill as `expected_skill`. Put each case's required starter files under `evals/files/<case-id>/`, and declare `files: ["evals/files/<case-id>"]` on every corresponding eval entry. Do not omit `files` in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.

3. Scaffold the evaluator contract. For generated tasks: ```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom ``` For shell checks: ```bash skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom ``` For native Harbor tasks: ```bash skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config ```

4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.

5. Validate before running. ```bash skillevaluator validate <skill-dir> --harbor-contract ``` Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.

6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.

## Grader Contract

Python and shell graders run inside the Harbor verifier context. They may read:

- `/logs/agent/trajectory.json` for agent actions and final answer evidence - `/tests/entry.json` for the eval case metadata - `/workspace/input/` for the entry's declared committed fixtures from `evals/files/` - `/solution/` or other task outputs only when the task environment produces them

They must write:

- `/logs/verifier/reward.json` - `/logs/verifier/reward.txt` with a numeric score from `0.0` to `1.0`

Use this reward shape:

```json { "overall": 0.92, "custom_metrics": { "domain_repair": 1.0, "domain_verification": 0.8 }, "details": { "domain_repair": { "score": 1.0, "reason": "The solution repaired the required files." } } } ```

In `default_plus_custom`, default evaluator scoring keeps its `overall` authoritative and adds the grader's `custom_metrics` into reports. In `custom_only`, the grader's `overall` is the pass/fail reward.

Never emit custom metric names that collide with reserved evaluator fields: `security`, `skill_execution`, `skill_efficiency`, `accuracy`, `goal_accuracy`, `behavior_check`, `overall`, `details`, `metrics`, `metric_set`, or `entry_id`.

## Translation Rules

- Convert each rubric item into a deterministic check when possible. - If a rubric item requires judgment, encode observable proxies and explain the limits in `details`. - Keep metrics stable across baseline and with-skill runs. - Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates. - Keep custom metric values clamped to `0.0` through `1.0`. - Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior. - Convert expected trigger/non-trigger metadata into `expected_skill`, `expected_behavior`, negative cases, or custom metrics that inspect trajectory evidence.

## RAPIDS-Style Example

For a benchmark task with `task.yaml`, `code/`, prompt variants, coverage, and a rubric:

1. Copy `code/` into `evals/files/<case-id>/`. 2. Create one or more `evals/evals.json` entries from the prompt variants, and set `files: ["evals/files/<case-id>"]` on each corresponding entry. 3. Set `expected_skill` to the benchmark's target skill. 4. Implement `evals/grader.py` to inspect the agent trajectory and changed workspace files. 5. Emit custom metrics for each rubric criterion, for example `rapids_diagnosis`, `rapids_requirements_repair`, `rapids_repair_safety`, and `rapids_verification`. 6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.

## Limitations

- The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy. - `init-custom-grader` creates scaffolding only; the agent must replace the placeholder scoring logic. - Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.

## Troubleshooting

| Problem | Fix | | --- | --- | | `evals/evals.json` missing | Create entries from the benchmark prompt or run `init-custom-grader` to seed one. | | Custom metrics do not appear | Ensure `reward.json` has numeric values under `custom_metrics` and no reserved-name collisions. | | `custom_only` fails | Write numeric `overall` in `reward.json` or numeric `reward.txt`. | | Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |

## Final Response

When finished, report:

- files created or changed - exact validation and evaluation commands - default evaluator metric results - custom metric results - whether the proof was full E2E or only static/local validation - any benchmark rubric criteria that remain partly judgment-based

Technische Details

Version
1.0.0
Lizenz
Apache-2.0
Letzte Aktualisierung
21. Aug. 2026
Veröffentlicht
20. Aug. 2026

Entscheidungsübersicht

Fallback-Kandidat

69
Bereit
Prototyp
Phase

recent repository activity

Audit

Installationsprüfung

Installations- und Adoptionsprüfung

81
Prüfung nötig
Sicherheit
83/100
Wartung
100/100
Installieren
92/100
Vollständiges Audit öffnenEval-Bericht ansehen

Von Agent belegte Evidenz

Von Agent belegte Evidenz

Ergebnisberichte nach Resolve, Prüfung, Installation und einem begrenzten Lauf.

0
Belegt
Needs first agent runAuto-Installation: zuerst prüfenLetzter: Unbekannt
Erfolgsrate
Letzter Fehler
Ergebnisse
0
Ausgabequalität
Fehlgeschlagen
0
Nicht relevant
0
Installationen
0
Durch Risiko blockiert
0
Einrichtung erforderlich
0
Produktion
0

Noch keine Agent-Ergebnisdaten. Der erste Lauf kann Erfolg, Einrichtungsbedarf, Risikoblockaden, Fehler oder Irrelevanz über /api/agent/outcome melden.

Installieren

Zum Agent-Workflow hinzufügen

Kostenlos und Open Source. Bericht vor der Installation in Produktions-Agents prüfen.

Wachstums-Loop

Share-Kit

X

Szenariobasierter Entwurf für create-custom-grader, bereit für einen manuellen X-Post.

Kuratorenhinweis
create-custom-grader: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check...

187 stars

https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x
X-Entwurf öffnen
Optionale Antwort mit Installationsbefehl
Listing + install path for create-custom-grader:
https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x

Install: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Antwortentwurf öffnen

Quelle des Eintrags

Registry-indexiert

Beanspruchbar

Dieser Eintrag wurde aus öffentlichen Quellen indexiert und ist erst nach Genehmigung eines Maintainer-Anspruchs offiziell.

Ersteller
NVIDIA
Indexiert von
OpenAgentSkill Community-Index

Die Zuordnung verlinkt auf das öffentliche Repository oder Creator-Profil. Creator können den Eintrag beanspruchen, um Eigentümersignale zu aktualisieren.

Diesen Skill beanspruchen

Eigentümeranspruch

Diesen Skill-Eintrag beanspruchen

Dieser Registry-indexiert-Eintrag wird NVIDIA zugeschrieben, ist aber noch nicht offiziell markiert. Beanspruche ihn, um ein verifiziertes Eigentümersignal hinzuzufügen und künftige Launch-, Installations- und Audit-Updates vertrauenswürdiger zu machen.

Creator-Backlink-Kit

Evidenz-Badges in deine README einfügen

Zeige den kanonischen Eintrag, aktuelle Vertrauens- und Audit-Signale sowie echte Agent-Proven-Evidenz dort, wo Entwickler das Repository bewerten.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=listed&label=Listed)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=trust&label=Trust)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=audit&label=Audit)](https://www.openagentskill.com/skills/nvidia-create-custom-grader/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)

Autor

N

NVIDIA

@nvidia

Plattform-Fit

Gesundheitssignale

GitHub-Stars
187
Qualitätswert
39/100
Letzter GitHub-Push
21. Aug. 2026
Framework-Hinweise
Unbekannt
OpenAgentSkill-Aufrufe
0
Installationskopien
0
Externe Klicks
0

Community-Signal

Teile mit, ob dieser Skill für deinen Agent-Workflow nützlich ist. Zusammengefasstes Feedback verbessert das Ranking im Laufe der Zeit.

Vertrauen & Sicherheit

Nur Sandbox

70
  • GitHub-Akzeptanz187 GitHub-StarsInfo
  • Star-/Fork-Aktivität187 Stars und 14 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbarPrüfen
  • Aktuelle Wartung1 Tage seit dem letzten PushBestanden
  • LizenzklarheitApache-2.0Bestanden
  • README/SKILL.md-VollständigkeitMetadaten enthalten ausreichend Nutzungs- und Workflow-KontextBestanden
  • Abhängigkeits-/LaufzeitrisikoBefehlsausführungsflächeInfo