create-custom-grader
Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Profil de l’actif
Recherche et travail de connaissance
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scénario
RAG and knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Adéquation Agent
Claude Code + CLI + Codex
Compatible avec Codex, Claude Code, Cursor, CLI ou des Agents personnalisés.
Installer
Prêt
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Maintenance
À jour
2 jours depuis le dernier push
Risque
Revue nécessaire
Quality score needs review
Qualité GitHub
187
70/100 Qualité · 78/100 Confiance
Tags de couverture
Notes de revue
Quality score needs review · Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Carte d’adoption Agent
Confiance, audit et préparation à l’installation en un coup d’œil
Ces scores combinent les métadonnées publiques du dépôt, les signaux de revue OpenAgentSkill, la fraîcheur de maintenance et la préparation à l’installation. Ils servent à présélectionner et ne remplacent pas la revue humaine.
Qualité
SolideSolid option that is likely worth shortlisting for production workflows.
Confiance
Sandbox uniquementCandidate utile avec des signaux de confiance incomplets ou mixtes. Gardez-la dans un espace isolé jusqu’à ce que la boucle de résultats confirme son adéquation.
Audit
Revue nécessaireRevue lisible par machine de la préparation à l’installation, des métadonnées de sécurité, de la maintenance et du risque d’adoption.
Trust Score OpenAgentSkill v5
Revue humaine avant installation
Exécutez uniquement dans un sandbox et comparez les alternatives proches avant usage réel.
Stars
187 stars GitHub
Activité du dépôt
187 stars et 14 forks
Maintenance
2 jours depuis le dernier push
Licence
Apache-2.0
Installer
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Sécurité d’installation
Chemin d’installation standard de package ou runtime
Surface de permissions
shell or command execution, filesystem or document access
Résultats Agent
Pas encore de données de résultats Agent
Documentation
Contexte README/SKILL.md solide
Résumé des risques
Revoir avant production
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Préparation à l’installation
Chemin d’installation disponible
- Le chemin d’installation est disponible
- La preuve du dépôt est disponible
- La licence est déclarée
- Pas encore de preuve de résultat Agent-Proven
Métadonnées lisibles par Agent
Données de décision lisibles par machine pour ce skill.
Utilisez ce bloc ou le JSON intégré pour décider si un Agent doit installer ce skill, choisir une alternative ou demander d’abord une revue humaine.
Tâches adaptées
- workflows Browser automation
- Équipes Claude Code
- builders willing to evaluate younger projects
- Navigate pages
Agents adaptés
Décision d’installation
- Commande
- npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
- Politique
- Revoir
- Revue humaine
- Oui
Confiance et risque
- Confiance
- 70/100
- Audit
- 81/100
- Niveau de risque
- Revue nécessaire
Boucle de résultat
- Endpoint
- /api/agent/outcome
- ID d’événement
- resolve
- Résultats
- 5
Commande d’installation
npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderNe pas utiliser quand
- Équipes qui nécessitent un SLA soutenu par le fournisseur
- Environnements fortement conformes sans revue interne de sécurité
- No OpenAgentSkill engagement data yet
- Indices de permissions à haut risque : exécution shell ou de commande
- Quality score needs review
Sécurité Agent v2
53/100 · Éviter l’installation automatique
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
Élevé
Exécution shell ou de commande
Les métadonnées de la skill font référence à des workflows de terminal, CLI, shell, sous-processus ou exécution de commande.
Moyen
Accès réseau
La skill récupère probablement des pages distantes, API, dépôts ou services externes.
Moyen
Accès au système de fichiers
La skill peut lire ou écrire des fichiers de projet, documents, artefacts générés ou l’état local de l’espace de travail.
- Indices de permissions à haut risque : exécution shell ou de commande
- Quality score needs review
Cibles d’installation
Installer ce skill dans votre workflow Agent
Utilisez le point de terminaison public pour récupérer la commande, la checklist, les prompts et les liens canoniques.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install nvidia-create-custom-graderPlan de résolution Agent
Laissez un Agent vérifier la pertinence avant l’installation.
L’API Resolve renvoie la skill sélectionnée, des alternatives, la politique de sécurité, les notes d’audit, la cible d’installation et un prompt prêt à l’emploi.
Ouvrir JSON
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Texte Resolve
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Relais d’installation
/api/skills/nvidia-create-custom-grader/install
L’Agent doit vérifier
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copier le prompt
Task: Use create-custom-grader in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install
Install command: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Relais Agent
Donnez à l’Agent le chemin d’installation, pas un autre annuaire.
Utilisez le point de terminaison public pour récupérer la commande, la checklist, les prompts et les liens canoniques.
Relais d’installation
/api/skills/nvidia-create-custom-grader/install
Format texte LLM
/api/skills/nvidia-create-custom-grader/install?format=text
Trouver des alternatives
/api/skills/search?q=create-custom-grader&limit=3
Prompt Agent
Use create-custom-grader for this task. Review https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install, then install with: npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderMétadonnées Registry
Profil lisible par Agent pour la sélection automatique de skills.
L’API Registry fournit les signaux de décision, confiance, audit, cas d’usage et installation sans analyser l’interface.
Manifest
/api/registry/manifest/nvidia-create-custom-grader
Texte LLM
/api/registry/manifest/nvidia-create-custom-grader?format=text
Alias d’installation
/api/registry/install/nvidia-create-custom-grader
Recommander
/api/registry/recommend?task=Use%20create-custom-grader%20in%20an%20agent%20workflow&limit=3
Adéquation Agent
Browser automation
Tags de cas d’usage
Plateformes
Claude Code
Rapport d’audit
Revue nécessaire · 81/100
Revue lisible par machine de la préparation à l’installation, des métadonnées de sécurité, de la maintenance et du risque d’adoption.
Panneau de décision Agent
Fallback candidate for Browser automation
Prototype with this skill first; keep a fallback candidate ready.
Rôle dans la pile
Candidate de secours
Pertinence principale
Browser automation
Libellé de confiance
Prototyper d’abord
Chemin d’installation
Commande prête
À utiliser lorsque
- workflows Browser automation
- Équipes Claude Code
- builders willing to evaluate younger projects
Preuves
- recent repository activity
- install command or GitHub repo available
- profil qualité 70/100
revoir d’abord
- No OpenAgentSkill engagement data yet
Chemin d’implémentation
- 1Installez-le dans un Agent en sandbox et exécutez une tâche de Browser automation de bout en bout.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Profil de confiance
Sandbox uniquement
Candidate utile avec des signaux de confiance incomplets ou mixtes. Gardez-la dans un espace isolé jusqu’à ce que la boucle de résultats confirme son adéquation.
Adoption GitHub
Info187 stars GitHub
Activité stars/forks
Vérifier187 stars et 14 forks; l’activité des issues n’est pas disponible dans les métadonnées actuelles
Maintenance récente
Validé2 jours depuis le dernier push
Clarté de licence
ValidéApache-2.0
Signaux positifs
- Revue IA approuvée
- Le chemin d’installation est disponible
- La preuve du dépôt est disponible
- Dépôt maintenu récemment
- La commande d’installation ne présente aucun motif de haut risque évident
- La boucle de résultats est prête mais nécessite la première exécution réelle de l’Agent
Réviser avant installation
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
- Pas encore de rapports de résultats Agent réels
- Une revue humaine est requise avant une installation sans surveillance
Action recommandée
Exécutez uniquement dans un sandbox et comparez les alternatives proches avant usage réel.
Profil qualité
Solide candidat pour les workflows Agent
Solid option that is likely worth shortlisting for production workflows.
Adéquation au workflow
Utilisez cette skill dans ces scénarios
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
Adéquation au workflow
Ajouter à un workflow complet
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Liste d’alternatives
Comparer avant installation
Similar skills that may fit this task.
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
MoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Cua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
Vue d’ensemble
--- name: create-custom-grader description: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. metadata: author: SkillEvaluator Maintainers <maintainers@example.com> ---
# Create Custom Grader
Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.
## Purpose
Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.
## When To Use
Use this skill when the user wants to:
- bring an existing benchmark into SkillEvaluator - turn a rubric into `evals/grader.py` or `evals/grader.sh` - add custom metrics beside the default evaluator metrics - convert task files such as `task.yaml`, `task.json`, pytest checks, or shell verifiers into BYOG or BYOT - prove a team can run its own benchmark through SkillEvaluator
Do not use this skill for ordinary `evals/evals.json` authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.
## Instructions
1. Read the target skill, existing `evals/`, benchmark prompts, fixtures, and any verifier code. 2. Choose `default_plus_custom` when custom metrics should complement default evaluator scoring. 3. Choose `custom_only` only when the user wants the custom grader to own pass/fail semantics. 4. Write or update `evals/grader.py` or `evals/grader.sh`, then validate the Harbor contract.
## Examples
```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom skillevaluator tier3 validate <skill-dir> ```
## Prerequisites
- The target skill directory should contain `SKILL.md`. - The SkillEvaluator CLI should be available as `skillevaluator`. - Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.
## Core Choice
Choose one path before writing files:
| User need | Evaluator shape | | --- | --- | | Existing `evals.json` task plus extra domain checks | Top-level BYOG: `evals/grader.py` or `evals/grader.sh` | | Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus `evals/evals.json` and `evals/files/` | | Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: `evals/harbor/<case>/...` | | User wants only custom reward/pass criteria | `grading.mode: custom_only` | | User wants default evaluator dimensions plus custom metrics | `grading.mode: default_plus_custom` |
Default to `default_plus_custom` unless the user explicitly wants the custom grader to replace the default evaluator metrics.
## Workflow
1. Resolve the target skill and benchmark source. Read the target `SKILL.md`, existing `evals/`, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.
2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as `question` entries. Use the target skill as `expected_skill`. Put each case's required starter files under `evals/files/<case-id>/`, and declare `files: ["evals/files/<case-id>"]` on every corresponding eval entry. Do not omit `files` in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.
3. Scaffold the evaluator contract. For generated tasks: ```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom ``` For shell checks: ```bash skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom ``` For native Harbor tasks: ```bash skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config ```
4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.
5. Validate before running. ```bash skillevaluator validate <skill-dir> --harbor-contract ``` Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.
6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.
## Grader Contract
Python and shell graders run inside the Harbor verifier context. They may read:
- `/logs/agent/trajectory.json` for agent actions and final answer evidence - `/tests/entry.json` for the eval case metadata - `/workspace/input/` for the entry's declared committed fixtures from `evals/files/` - `/solution/` or other task outputs only when the task environment produces them
They must write:
- `/logs/verifier/reward.json` - `/logs/verifier/reward.txt` with a numeric score from `0.0` to `1.0`
Use this reward shape:
```json { "overall": 0.92, "custom_metrics": { "domain_repair": 1.0, "domain_verification": 0.8 }, "details": { "domain_repair": { "score": 1.0, "reason": "The solution repaired the required files." } } } ```
In `default_plus_custom`, default evaluator scoring keeps its `overall` authoritative and adds the grader's `custom_metrics` into reports. In `custom_only`, the grader's `overall` is the pass/fail reward.
Never emit custom metric names that collide with reserved evaluator fields: `security`, `skill_execution`, `skill_efficiency`, `accuracy`, `goal_accuracy`, `behavior_check`, `overall`, `details`, `metrics`, `metric_set`, or `entry_id`.
## Translation Rules
- Convert each rubric item into a deterministic check when possible. - If a rubric item requires judgment, encode observable proxies and explain the limits in `details`. - Keep metrics stable across baseline and with-skill runs. - Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates. - Keep custom metric values clamped to `0.0` through `1.0`. - Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior. - Convert expected trigger/non-trigger metadata into `expected_skill`, `expected_behavior`, negative cases, or custom metrics that inspect trajectory evidence.
## RAPIDS-Style Example
For a benchmark task with `task.yaml`, `code/`, prompt variants, coverage, and a rubric:
1. Copy `code/` into `evals/files/<case-id>/`. 2. Create one or more `evals/evals.json` entries from the prompt variants, and set `files: ["evals/files/<case-id>"]` on each corresponding entry. 3. Set `expected_skill` to the benchmark's target skill. 4. Implement `evals/grader.py` to inspect the agent trajectory and changed workspace files. 5. Emit custom metrics for each rubric criterion, for example `rapids_diagnosis`, `rapids_requirements_repair`, `rapids_repair_safety`, and `rapids_verification`. 6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.
## Limitations
- The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy. - `init-custom-grader` creates scaffolding only; the agent must replace the placeholder scoring logic. - Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.
## Troubleshooting
| Problem | Fix | | --- | --- | | `evals/evals.json` missing | Create entries from the benchmark prompt or run `init-custom-grader` to seed one. | | Custom metrics do not appear | Ensure `reward.json` has numeric values under `custom_metrics` and no reserved-name collisions. | | `custom_only` fails | Write numeric `overall` in `reward.json` or numeric `reward.txt`. | | Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |
## Final Response
When finished, report:
- files created or changed - exact validation and evaluation commands - default evaluator metric results - custom metric results - whether the proof was full E2E or only static/local validation - any benchmark rubric criteria that remain partly judgment-based
Détails techniques
- Version
- 1.0.0
- Licence
- Apache-2.0
- Dernière mise à jour
- 21 août 2026
- Publié
- 20 août 2026
Instantané de décision
Candidate de secours
recent repository activity
Audit
Revue d’installation
Revue d’installation et d’adoption
- Sécurité
- 83/100
- Maintenance
- 100/100
- Installer
- 92/100
Preuves validées par Agent
Preuves validées par Agent
Rapports après resolve, revue, installation et une exécution limitée.
- Taux de réussite
- —
- Échec récent
- —
- Résultats
- 0
- Qualité de sortie
- —
- Échecs
- 0
- Non pertinent
- 0
- Installations
- 0
- Bloqué par le risque
- 0
- Configuration requise
- 0
- Production
- 0
Aucune donnée de résultat Agent pour l’instant. La première exécution peut signaler succès, besoin de configuration, blocage de risque, échec ou non-pertinence via /api/agent/outcome.
Installer
Ajouter au workflow Agent
Gratuit et open source. Examinez le rapport avant l’installation dans des Agents de production.
Boucle de croissance
Kit de partage
Brouillon guidé par scénario pour create-custom-grader, prêt pour une publication manuelle sur X.
create-custom-grader: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check... 187 stars https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x
Réponse facultative avec commande d’installation
Listing + install path for create-custom-grader: https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x Install: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Source de la fiche
Indexé par Registry
Cette fiche a été indexée à partir de sources publiques et n’est pas marquée officielle tant qu’une revendication de mainteneur n’est pas approuvée.
- Créateur
- NVIDIA
- Source
- NVIDIA/SkillEvaluator
- Indexé par
- Index communautaire OpenAgentSkill
L’attribution renvoie au dépôt public ou au profil du créateur. Les créateurs peuvent revendiquer la fiche pour mettre à jour les signaux de propriété.
Revendiquer ce skillRevendication du propriétaire
Revendiquer cette fiche de skill
Cette fiche Indexé par Registry est attribuée à NVIDIA, mais n’est pas encore marquée officielle. Revendiquez-la pour ajouter un signal de propriétaire vérifié et rendre les futures mises à jour de lancement, d’installation et d’audit plus fiables.
Kit de backlinks créateur
Ajoutez les badges de preuve à votre README
Affichez la fiche canonique, les signaux actuels de confiance et d’audit, ainsi que de vraies preuves Agent-Proven là où les développeurs évaluent le dépôt.
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader/audit)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)Auteur
NVIDIA
@nvidia
Tags
Adéquation plateforme
Signaux de santé
- Stars GitHub
- 187
- Score de qualité
- 39/100
- Dernier push GitHub
- 21 août 2026
- Indications de framework
- Inconnu
- Vues OpenAgentSkill
- 0
- Copies d’installation
- 0
- Clics sortants
- 0
Signal de communauté
Indiquez si ce skill semble utile à votre workflow Agent. Les retours agrégés améliorent le classement au fil du temps.
Confiance et sécurité
Sandbox uniquement
- Adoption GitHub187 stars GitHubInfo
- Activité stars/forks187 stars et 14 forks; l’activité des issues n’est pas disponible dans les métadonnées actuellesVérifier
- Maintenance récente2 jours depuis le dernier pushValidé
- Clarté de licenceApache-2.0Validé
- Complétude README/SKILL.mdLes métadonnées incluent suffisamment de contexte d’usage et de workflowValidé
- Risque dépendances/runtimeSurface d’exécution de commandesInfo
Skills associés
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsCua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
21.4K Stars