create-custom-grader
Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Asset-Profil
Recherche und Wissensarbeit
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Szenario
RAG and knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Agent-Fit
Claude Code + CLI + Codex
Geeignet für Codex, Claude Code, Cursor, CLI oder benutzerdefinierte Agents.
Installieren
Bereit
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Wartung
Aktuell
1 Tage seit dem letzten Push
Risiko
Prüfung nötig
Quality score needs review
GitHub-Qualität
187
70/100 Qualität · 78/100 Vertrauen
Abdeckungs-Tags
Review-Notizen
Quality score needs review · Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Agent-Adoptionskarte
Vertrauen, Audit und Installationsbereitschaft auf einen Blick
Diese Werte kombinieren öffentliche Repository-Metadaten, OpenAgentSkill-Reviewsignale, Wartungsaktualität und Installationsbereitschaft. Sie helfen bei der Vorauswahl, ersetzen aber keine menschliche Prüfung.
Qualität
StarkSolid option that is likely worth shortlisting for production workflows.
Vertrauen
Nur SandboxNützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
Audit
Prüfung nötigMaschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
OpenAgentSkill Trust Score v5
Menschliche Prüfung vor Installation
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Stars
187 GitHub-Stars
Repository-Aktivität
187 Stars und 14 Forks
Wartung
1 Tage seit dem letzten Push
Lizenz
Apache-2.0
Installieren
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Installationssicherheit
Standard-Paket- oder Laufzeit-Installationspfad
Berechtigungsfläche
shell or command execution, filesystem or document access
Agent-Ergebnisse
Noch keine Agent-Ergebnisdaten
Dokumentation
Starker README/SKILL.md-Kontext
Risikoübersicht
Vor Produktion prüfen
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Installationsbereitschaft
Installationspfad verfügbar
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Lizenz ist angegeben
- Noch keine Agent-Proven-Ergebnisbelege
Agent-lesbare Metadaten
Maschinenlesbare Entscheidungsdaten für diesen Skill.
Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.
Geeignete Aufgaben
- Browser automation-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
- Navigate pages
Geeignete Agents
Installationsentscheidung
- Befehl
- npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
- Richtlinie
- Prüfen
- Menschliche Prüfung
- Ja
Vertrauen und Risiko
- Vertrauen
- 70/100
- Audit
- 81/100
- Risikoebene
- Prüfung nötig
Ergebnis-Loop
- Endpoint
- /api/agent/outcome
- Event-ID
- resolve
- Ergebnisse
- 5
Installationsbefehl
npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderNicht verwenden, wenn
- Teams, die ein vom Anbieter unterstütztes SLA benötigen
- Hochregulierte Umgebungen ohne interne Sicherheitsprüfung
- No OpenAgentSkill engagement data yet
- Hinweise auf Hochrisiko-Berechtigungen: Shell- oder Befehlsausführung
- Quality score needs review
Agent-Sicherheit v2
53/100 · Automatische Installation vermeiden
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
Hoch
Shell- oder Befehlsausführung
Die Skill-Metadaten verweisen auf Terminal-, CLI-, Shell-, Subprozess- oder Befehlsausführungs-Workflows.
Mittel
Netzwerkzugriff
Die Skill ruft wahrscheinlich Remote-Seiten, APIs, Repositories oder externe Dienste ab.
Mittel
Dateisystemzugriff
Die Skill kann Projektdateien, Dokumente, generierte Artefakte oder den lokalen Arbeitsbereich lesen oder schreiben.
- Hinweise auf Hochrisiko-Berechtigungen: Shell- oder Befehlsausführung
- Quality score needs review
Installationsziele
Diesen Skill im Agent-Workflow installieren
Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install nvidia-create-custom-graderAgent-Auflösungsplan
Lass einen Agent die Eignung vor der Installation prüfen.
Die Resolve API liefert die beste Skill, Alternativen, Sicherheitsrichtlinien, Auditnotizen, Installationsziel und einen direkt nutzbaren Prompt.
JSON öffnen
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve-Text
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Installationsübergabe
/api/skills/nvidia-create-custom-grader/install
Agent sollte prüfen
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Prompt kopieren
Task: Use create-custom-grader in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install
Install command: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent-Übergabe
Gib dem Agent den Installationspfad, nicht noch ein Verzeichnis.
Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.
Installationsübergabe
/api/skills/nvidia-create-custom-grader/install
LLM-Textformat
/api/skills/nvidia-create-custom-grader/install?format=text
Alternativen finden
/api/skills/search?q=create-custom-grader&limit=3
Agent-Prompt
Use create-custom-grader for this task. Review https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install, then install with: npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderRegistry-Metadaten
Agent-lesbares Profil für die automatische Skill-Auswahl.
Die Registry API stellt Entscheidungs-, Vertrauens-, Audit-, Use-Case- und Installationssignale ohne UI-Scraping bereit.
Manifest
/api/registry/manifest/nvidia-create-custom-grader
LLM-Text
/api/registry/manifest/nvidia-create-custom-grader?format=text
Installationsalias
/api/registry/install/nvidia-create-custom-grader
Empfehlen
/api/registry/recommend?task=Use%20create-custom-grader%20in%20an%20agent%20workflow&limit=3
Agent-Fit
Browser automation
Plattformen
Claude Code
Audit-Bericht
Prüfung nötig · 81/100
Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
Agent-Entscheidungspanel
Fallback candidate for Browser automation
Prototype with this skill first; keep a fallback candidate ready.
Rolle im Stack
Fallback-Kandidat
Primäre Eignung
Browser automation
Vertrauenslabel
Zuerst prototypisieren
Installationspfad
Befehl bereit
Verwenden wenn
- Browser automation-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
Evidenz
- recent repository activity
- install command or GitHub repo available
- Qualitätsprofil 70/100
zuerst prüfen
- No OpenAgentSkill engagement data yet
Implementierungspfad
- 1Installieren Sie es in einem Sandbox-Agent und führen Sie eine Browser automation-Aufgabe vollständig aus.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Vertrauensprofil
Nur Sandbox
Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
GitHub-Akzeptanz
Info187 GitHub-Stars
Star-/Fork-Aktivität
Prüfen187 Stars und 14 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbar
Aktuelle Wartung
Bestanden1 Tage seit dem letzten Push
Lizenzklarheit
BestandenApache-2.0
Positive Signale
- KI-Prüfung genehmigt
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Kürzlich gewartetes Repository
- Der Installationsbefehl weist kein offensichtliches Hochrisikomuster auf
- Ergebniszyklus ist bereit, benötigt aber den ersten echten Agent-Lauf
Vor Installation prüfen
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
- Noch keine echten Agent-Ergebnisberichte
- Vor unbeaufsichtigter Installation ist menschliche Prüfung erforderlich
Empfohlene Aktion
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Qualitätsprofil
Stark Kandidat für Agent-Workflows
Solid option that is likely worth shortlisting for production workflows.
Workflow-Eignung
Diese Skill in diesen Szenarien nutzen
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
Workflow-Eignung
Zum vollständigen Workflow hinzufügen
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternativen-Shortlist
Vor Installation vergleichen
Similar skills that may fit this task.
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
MoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Cua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
Übersicht
--- name: create-custom-grader description: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. metadata: author: SkillEvaluator Maintainers <maintainers@example.com> ---
# Create Custom Grader
Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.
## Purpose
Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.
## When To Use
Use this skill when the user wants to:
- bring an existing benchmark into SkillEvaluator - turn a rubric into `evals/grader.py` or `evals/grader.sh` - add custom metrics beside the default evaluator metrics - convert task files such as `task.yaml`, `task.json`, pytest checks, or shell verifiers into BYOG or BYOT - prove a team can run its own benchmark through SkillEvaluator
Do not use this skill for ordinary `evals/evals.json` authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.
## Instructions
1. Read the target skill, existing `evals/`, benchmark prompts, fixtures, and any verifier code. 2. Choose `default_plus_custom` when custom metrics should complement default evaluator scoring. 3. Choose `custom_only` only when the user wants the custom grader to own pass/fail semantics. 4. Write or update `evals/grader.py` or `evals/grader.sh`, then validate the Harbor contract.
## Examples
```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom skillevaluator tier3 validate <skill-dir> ```
## Prerequisites
- The target skill directory should contain `SKILL.md`. - The SkillEvaluator CLI should be available as `skillevaluator`. - Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.
## Core Choice
Choose one path before writing files:
| User need | Evaluator shape | | --- | --- | | Existing `evals.json` task plus extra domain checks | Top-level BYOG: `evals/grader.py` or `evals/grader.sh` | | Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus `evals/evals.json` and `evals/files/` | | Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: `evals/harbor/<case>/...` | | User wants only custom reward/pass criteria | `grading.mode: custom_only` | | User wants default evaluator dimensions plus custom metrics | `grading.mode: default_plus_custom` |
Default to `default_plus_custom` unless the user explicitly wants the custom grader to replace the default evaluator metrics.
## Workflow
1. Resolve the target skill and benchmark source. Read the target `SKILL.md`, existing `evals/`, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.
2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as `question` entries. Use the target skill as `expected_skill`. Put each case's required starter files under `evals/files/<case-id>/`, and declare `files: ["evals/files/<case-id>"]` on every corresponding eval entry. Do not omit `files` in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.
3. Scaffold the evaluator contract. For generated tasks: ```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom ``` For shell checks: ```bash skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom ``` For native Harbor tasks: ```bash skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config ```
4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.
5. Validate before running. ```bash skillevaluator validate <skill-dir> --harbor-contract ``` Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.
6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.
## Grader Contract
Python and shell graders run inside the Harbor verifier context. They may read:
- `/logs/agent/trajectory.json` for agent actions and final answer evidence - `/tests/entry.json` for the eval case metadata - `/workspace/input/` for the entry's declared committed fixtures from `evals/files/` - `/solution/` or other task outputs only when the task environment produces them
They must write:
- `/logs/verifier/reward.json` - `/logs/verifier/reward.txt` with a numeric score from `0.0` to `1.0`
Use this reward shape:
```json { "overall": 0.92, "custom_metrics": { "domain_repair": 1.0, "domain_verification": 0.8 }, "details": { "domain_repair": { "score": 1.0, "reason": "The solution repaired the required files." } } } ```
In `default_plus_custom`, default evaluator scoring keeps its `overall` authoritative and adds the grader's `custom_metrics` into reports. In `custom_only`, the grader's `overall` is the pass/fail reward.
Never emit custom metric names that collide with reserved evaluator fields: `security`, `skill_execution`, `skill_efficiency`, `accuracy`, `goal_accuracy`, `behavior_check`, `overall`, `details`, `metrics`, `metric_set`, or `entry_id`.
## Translation Rules
- Convert each rubric item into a deterministic check when possible. - If a rubric item requires judgment, encode observable proxies and explain the limits in `details`. - Keep metrics stable across baseline and with-skill runs. - Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates. - Keep custom metric values clamped to `0.0` through `1.0`. - Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior. - Convert expected trigger/non-trigger metadata into `expected_skill`, `expected_behavior`, negative cases, or custom metrics that inspect trajectory evidence.
## RAPIDS-Style Example
For a benchmark task with `task.yaml`, `code/`, prompt variants, coverage, and a rubric:
1. Copy `code/` into `evals/files/<case-id>/`. 2. Create one or more `evals/evals.json` entries from the prompt variants, and set `files: ["evals/files/<case-id>"]` on each corresponding entry. 3. Set `expected_skill` to the benchmark's target skill. 4. Implement `evals/grader.py` to inspect the agent trajectory and changed workspace files. 5. Emit custom metrics for each rubric criterion, for example `rapids_diagnosis`, `rapids_requirements_repair`, `rapids_repair_safety`, and `rapids_verification`. 6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.
## Limitations
- The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy. - `init-custom-grader` creates scaffolding only; the agent must replace the placeholder scoring logic. - Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.
## Troubleshooting
| Problem | Fix | | --- | --- | | `evals/evals.json` missing | Create entries from the benchmark prompt or run `init-custom-grader` to seed one. | | Custom metrics do not appear | Ensure `reward.json` has numeric values under `custom_metrics` and no reserved-name collisions. | | `custom_only` fails | Write numeric `overall` in `reward.json` or numeric `reward.txt`. | | Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |
## Final Response
When finished, report:
- files created or changed - exact validation and evaluation commands - default evaluator metric results - custom metric results - whether the proof was full E2E or only static/local validation - any benchmark rubric criteria that remain partly judgment-based
Technische Details
- Version
- 1.0.0
- Lizenz
- Apache-2.0
- Letzte Aktualisierung
- 21. Aug. 2026
- Veröffentlicht
- 20. Aug. 2026
Entscheidungsübersicht
Fallback-Kandidat
recent repository activity
Audit
Installationsprüfung
Installations- und Adoptionsprüfung
- Sicherheit
- 83/100
- Wartung
- 100/100
- Installieren
- 92/100
Von Agent belegte Evidenz
Von Agent belegte Evidenz
Ergebnisberichte nach Resolve, Prüfung, Installation und einem begrenzten Lauf.
- Erfolgsrate
- —
- Letzter Fehler
- —
- Ergebnisse
- 0
- Ausgabequalität
- —
- Fehlgeschlagen
- 0
- Nicht relevant
- 0
- Installationen
- 0
- Durch Risiko blockiert
- 0
- Einrichtung erforderlich
- 0
- Produktion
- 0
Noch keine Agent-Ergebnisdaten. Der erste Lauf kann Erfolg, Einrichtungsbedarf, Risikoblockaden, Fehler oder Irrelevanz über /api/agent/outcome melden.
Installieren
Zum Agent-Workflow hinzufügen
Kostenlos und Open Source. Bericht vor der Installation in Produktions-Agents prüfen.
Wachstums-Loop
Share-Kit
Szenariobasierter Entwurf für create-custom-grader, bereit für einen manuellen X-Post.
create-custom-grader: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check... 187 stars https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x
Optionale Antwort mit Installationsbefehl
Listing + install path for create-custom-grader: https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x Install: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Quelle des Eintrags
Registry-indexiert
Dieser Eintrag wurde aus öffentlichen Quellen indexiert und ist erst nach Genehmigung eines Maintainer-Anspruchs offiziell.
- Ersteller
- NVIDIA
- Quelle
- NVIDIA/SkillEvaluator
- Indexiert von
- OpenAgentSkill Community-Index
Die Zuordnung verlinkt auf das öffentliche Repository oder Creator-Profil. Creator können den Eintrag beanspruchen, um Eigentümersignale zu aktualisieren.
Diesen Skill beanspruchenEigentümeranspruch
Diesen Skill-Eintrag beanspruchen
Dieser Registry-indexiert-Eintrag wird NVIDIA zugeschrieben, ist aber noch nicht offiziell markiert. Beanspruche ihn, um ein verifiziertes Eigentümersignal hinzuzufügen und künftige Launch-, Installations- und Audit-Updates vertrauenswürdiger zu machen.
Creator-Backlink-Kit
Evidenz-Badges in deine README einfügen
Zeige den kanonischen Eintrag, aktuelle Vertrauens- und Audit-Signale sowie echte Agent-Proven-Evidenz dort, wo Entwickler das Repository bewerten.
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader/audit)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)Autor
NVIDIA
@nvidia
Tags
Plattform-Fit
Gesundheitssignale
- GitHub-Stars
- 187
- Qualitätswert
- 39/100
- Letzter GitHub-Push
- 21. Aug. 2026
- Framework-Hinweise
- Unbekannt
- OpenAgentSkill-Aufrufe
- 0
- Installationskopien
- 0
- Externe Klicks
- 0
Community-Signal
Teile mit, ob dieser Skill für deinen Agent-Workflow nützlich ist. Zusammengefasstes Feedback verbessert das Ranking im Laufe der Zeit.
Vertrauen & Sicherheit
Nur Sandbox
- GitHub-Akzeptanz187 GitHub-StarsInfo
- Star-/Fork-Aktivität187 Stars und 14 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbarPrüfen
- Aktuelle Wartung1 Tage seit dem letzten PushBestanden
- LizenzklarheitApache-2.0Bestanden
- README/SKILL.md-VollständigkeitMetadaten enthalten ausreichend Nutzungs- und Workflow-KontextBestanden
- Abhängigkeits-/LaufzeitrisikoBefehlsausführungsflächeInfo
Ähnliche Skills
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsCua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
21.4K Stars