create-custom-grader
Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Supply asset profile
Research and knowledge work
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
RAG and knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Maintenance
fresh
1d since push
Risk
Needs review
Quality score needs review
GitHub quality
187
70/100 Quality · 78/100 Trust
Coverage tags
Review notes
Quality score needs review · Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
StrongSolid option that is likely worth shortlisting for production workflows.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
187 GitHub stars
Repo activity
187 stars, 14 forks
Maintenance
1d since push
License
Apache-2.0
Install
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Install safety
standard package or runtime install path
Permission surface
shell or command execution, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Review before production
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- Browser automation workflows
- Claude Code teams
- builders willing to evaluate younger projects
- Navigate pages
Suited agents
Install decision
- Command
- npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
- Policy
- review
- Human review
- yes
Trust and risk
- Trust
- 70/100
- Audit
- 81/100
- Risk level
- Needs review
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderDo not use when
- teams that need a vendor-supported SLA
- high-compliance environments without internal security review
- No OpenAgentSkill engagement data yet
- High-risk permission hints: Shell or command execution
- Quality score needs review
Agent safety v2
53/100 · Avoid automatic install
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Shell or command execution
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Filesystem access
Skill may read or write project files, documents, generated artifacts, or local workspace state.
- High-risk permission hints: Shell or command execution
- Quality score needs review
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install nvidia-create-custom-graderAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/nvidia-create-custom-grader/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use create-custom-grader in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install
Install command: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/nvidia-create-custom-grader/install
LLM text format
/api/skills/nvidia-create-custom-grader/install?format=text
Find alternatives
/api/skills/search?q=create-custom-grader&limit=3
Agent prompt
Use create-custom-grader for this task. Review https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install, then install with: npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/nvidia-create-custom-grader
LLM text
/api/registry/manifest/nvidia-create-custom-grader?format=text
Install alias
/api/registry/install/nvidia-create-custom-grader
Recommend
/api/registry/recommend?task=Use%20create-custom-grader%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code
Audit report
Needs review · 81/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Fallback candidate for Browser automation
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Browser automation
Trust label
Prototype first
Install path
Command ready
Use when
- Browser automation workflows
- Claude Code teams
- builders willing to evaluate younger projects
Evidence
- recent repository activity
- install command or GitHub repo available
- 70/100 quality profile
review first
- No OpenAgentSkill engagement data yet
Implementation path
- 1Install it in a sandbox agent and run one Browser automation task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Sandbox only
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
INFO187 GitHub stars
Stars/forks activity
CHECK187 stars, 14 forks; issue activity unavailable in current metadata
Recent maintenance
PASS1d since push
License clarity
PASSApache-2.0
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Strong candidate for agent workflows
Solid option that is likely worth shortlisting for production workflows.
Workflow fit
Use this skill in these scenarios
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
Workflow fit
Add it to a complete workflow
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
MoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Cua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
Overview
--- name: create-custom-grader description: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. metadata: author: SkillEvaluator Maintainers <maintainers@example.com> ---
# Create Custom Grader
Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.
## Purpose
Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.
## When To Use
Use this skill when the user wants to:
- bring an existing benchmark into SkillEvaluator - turn a rubric into `evals/grader.py` or `evals/grader.sh` - add custom metrics beside the default evaluator metrics - convert task files such as `task.yaml`, `task.json`, pytest checks, or shell verifiers into BYOG or BYOT - prove a team can run its own benchmark through SkillEvaluator
Do not use this skill for ordinary `evals/evals.json` authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.
## Instructions
1. Read the target skill, existing `evals/`, benchmark prompts, fixtures, and any verifier code. 2. Choose `default_plus_custom` when custom metrics should complement default evaluator scoring. 3. Choose `custom_only` only when the user wants the custom grader to own pass/fail semantics. 4. Write or update `evals/grader.py` or `evals/grader.sh`, then validate the Harbor contract.
## Examples
```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom skillevaluator tier3 validate <skill-dir> ```
## Prerequisites
- The target skill directory should contain `SKILL.md`. - The SkillEvaluator CLI should be available as `skillevaluator`. - Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.
## Core Choice
Choose one path before writing files:
| User need | Evaluator shape | | --- | --- | | Existing `evals.json` task plus extra domain checks | Top-level BYOG: `evals/grader.py` or `evals/grader.sh` | | Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus `evals/evals.json` and `evals/files/` | | Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: `evals/harbor/<case>/...` | | User wants only custom reward/pass criteria | `grading.mode: custom_only` | | User wants default evaluator dimensions plus custom metrics | `grading.mode: default_plus_custom` |
Default to `default_plus_custom` unless the user explicitly wants the custom grader to replace the default evaluator metrics.
## Workflow
1. Resolve the target skill and benchmark source. Read the target `SKILL.md`, existing `evals/`, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.
2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as `question` entries. Use the target skill as `expected_skill`. Put each case's required starter files under `evals/files/<case-id>/`, and declare `files: ["evals/files/<case-id>"]` on every corresponding eval entry. Do not omit `files` in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.
3. Scaffold the evaluator contract. For generated tasks: ```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom ``` For shell checks: ```bash skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom ``` For native Harbor tasks: ```bash skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config ```
4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.
5. Validate before running. ```bash skillevaluator validate <skill-dir> --harbor-contract ``` Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.
6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.
## Grader Contract
Python and shell graders run inside the Harbor verifier context. They may read:
- `/logs/agent/trajectory.json` for agent actions and final answer evidence - `/tests/entry.json` for the eval case metadata - `/workspace/input/` for the entry's declared committed fixtures from `evals/files/` - `/solution/` or other task outputs only when the task environment produces them
They must write:
- `/logs/verifier/reward.json` - `/logs/verifier/reward.txt` with a numeric score from `0.0` to `1.0`
Use this reward shape:
```json { "overall": 0.92, "custom_metrics": { "domain_repair": 1.0, "domain_verification": 0.8 }, "details": { "domain_repair": { "score": 1.0, "reason": "The solution repaired the required files." } } } ```
In `default_plus_custom`, default evaluator scoring keeps its `overall` authoritative and adds the grader's `custom_metrics` into reports. In `custom_only`, the grader's `overall` is the pass/fail reward.
Never emit custom metric names that collide with reserved evaluator fields: `security`, `skill_execution`, `skill_efficiency`, `accuracy`, `goal_accuracy`, `behavior_check`, `overall`, `details`, `metrics`, `metric_set`, or `entry_id`.
## Translation Rules
- Convert each rubric item into a deterministic check when possible. - If a rubric item requires judgment, encode observable proxies and explain the limits in `details`. - Keep metrics stable across baseline and with-skill runs. - Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates. - Keep custom metric values clamped to `0.0` through `1.0`. - Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior. - Convert expected trigger/non-trigger metadata into `expected_skill`, `expected_behavior`, negative cases, or custom metrics that inspect trajectory evidence.
## RAPIDS-Style Example
For a benchmark task with `task.yaml`, `code/`, prompt variants, coverage, and a rubric:
1. Copy `code/` into `evals/files/<case-id>/`. 2. Create one or more `evals/evals.json` entries from the prompt variants, and set `files: ["evals/files/<case-id>"]` on each corresponding entry. 3. Set `expected_skill` to the benchmark's target skill. 4. Implement `evals/grader.py` to inspect the agent trajectory and changed workspace files. 5. Emit custom metrics for each rubric criterion, for example `rapids_diagnosis`, `rapids_requirements_repair`, `rapids_repair_safety`, and `rapids_verification`. 6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.
## Limitations
- The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy. - `init-custom-grader` creates scaffolding only; the agent must replace the placeholder scoring logic. - Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.
## Troubleshooting
| Problem | Fix | | --- | --- | | `evals/evals.json` missing | Create entries from the benchmark prompt or run `init-custom-grader` to seed one. | | Custom metrics do not appear | Ensure `reward.json` has numeric values under `custom_metrics` and no reserved-name collisions. | | `custom_only` fails | Write numeric `overall` in `reward.json` or numeric `reward.txt`. | | Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |
## Final Response
When finished, report:
- files created or changed - exact validation and evaluation commands - default evaluator metric results - custom metric results - whether the proof was full E2E or only static/local validation - any benchmark rubric criteria that remain partly judgment-based
Technical details
- Version
- 1.0.0
- License
- Apache-2.0
- Last updated
- Aug 21, 2026
- Published
- Aug 20, 2026
Decision snapshot
Fallback candidate
recent repository activity
Audit
Install review
Install and adoption review
- Security
- 83/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for create-custom-grader, ready for a manual X post.
create-custom-grader: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check... 187 stars https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x
Optional reply with install command
Listing + install path for create-custom-grader: https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x Install: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Listing source
Registry indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- NVIDIA
- Source
- NVIDIA/SkillEvaluator
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Registry indexed listing is attributed to NVIDIA but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader/audit)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)Author
NVIDIA
@nvidia
Tags
Platform fit
Health signals
- GitHub stars
- 187
- Quality score
- 39/100
- Last GitHub push
- Aug 21, 2026
- Framework hints
- Unknown
- OpenAgentSkill views
- 0
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Sandbox only
- GitHub adoption187 GitHub starsINFO
- Stars/forks activity187 stars, 14 forks; issue activity unavailable in current metadataCHECK
- Recent maintenance1d since pushPASS
- License clarityApache-2.0PASS
- README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
- Dependency/runtime riskcommand execution surfaceINFO
Related skills
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsCua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
21.4K Stars