create-custom-grader
Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Profil aset
Riset dan pekerjaan pengetahuan
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Skenario
RAG and knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Kecocokan Agent
Claude Code + CLI + Codex
Cocok untuk Codex, Claude Code, Cursor, CLI, atau Agent khusus.
Pasang
Siap
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Pemeliharaan
Terkini
2 hari sejak push
Risiko
Perlu ditinjau
Quality score needs review
Kualitas GitHub
187
70/100 Kualitas · 78/100 Kepercayaan
Tag cakupan
Catatan ulasan
Quality score needs review · Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Kartu adopsi Agent
Kepercayaan, audit, dan kesiapan pemasangan dalam sekali lihat
Skor ini menggabungkan metadata repositori publik, sinyal ulasan OpenAgentSkill, kebaruan pemeliharaan, dan kesiapan pemasangan. Ini adalah sinyal shortlist, bukan pengganti peninjauan manusia.
Kualitas
KuatSolid option that is likely worth shortlisting for production workflows.
Kepercayaan
Hanya sandboxKandidat berguna dengan sinyal kepercayaan yang kurang atau bercampur. Gunakan di ruang kerja terisolasi hingga loop hasil membuktikan kecocokan tugas.
Audit
Perlu ditinjauTinjauan yang dapat dibaca mesin tentang kesiapan pemasangan, metadata keamanan, pemeliharaan, dan risiko adopsi.
Trust Score OpenAgentSkill v5
Tinjauan manusia sebelum pemasangan
Jalankan hanya dalam sandbox dan bandingkan alternatif terdekat sebelum digunakan untuk kerja nyata.
Star
187 star GitHub
Aktivitas repositori
187 star dan 14 fork
Pemeliharaan
2 hari sejak push
Lisensi
Apache-2.0
Pasang
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Keamanan pemasangan
Jalur pemasangan paket atau runtime standar
Cakupan izin
shell or command execution, filesystem or document access
Hasil Agent
Belum ada data hasil Agent
Dokumentasi
Konteks README/SKILL.md kuat
Ringkasan risiko
Tinjau sebelum produksi
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
Kesiapan pemasangan
Jalur pemasangan tersedia
- Jalur pemasangan tersedia
- Bukti repositori tersedia
- Lisensi dinyatakan
- Belum ada bukti hasil Agent-Proven
Metadata yang dapat dibaca Agent
Data keputusan yang dapat dibaca mesin untuk skill ini.
Gunakan blok ini atau JSON tersemat untuk memutuskan apakah Agent perlu memasang skill ini, memilih alternatif, atau meminta tinjauan manusia terlebih dahulu.
Tugas yang sesuai
- alur kerja Browser automation
- Tim Claude Code
- builders willing to evaluate younger projects
- Navigate pages
Agent yang sesuai
Keputusan pemasangan
- Perintah
- npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
- Kebijakan
- Tinjau
- Tinjauan manusia
- Ya
Kepercayaan dan risiko
- Kepercayaan
- 70/100
- Audit
- 81/100
- Tingkat risiko
- Perlu ditinjau
Lingkar hasil
- Endpoint
- /api/agent/outcome
- ID event
- resolve
- Hasil
- 5
Perintah pemasangan
npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderJangan gunakan ketika
- Tim yang membutuhkan SLA dengan dukungan vendor
- Lingkungan berkompliansi tinggi tanpa tinjauan keamanan internal
- No major risk signals from current metadata
- Petunjuk izin berisiko tinggi: eksekusi shell atau perintah
- Quality score needs review
Keamanan Agent v2
53/100 · Hindari pemasangan otomatis
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
Tinggi
Eksekusi shell atau perintah
Metadata skill merujuk terminal, CLI, shell, subprocess, atau alur kerja eksekusi perintah.
Sedang
Akses jaringan
Skill kemungkinan mengambil halaman jarak jauh, API, repositori, atau layanan eksternal.
Sedang
Akses sistem file
Skill dapat membaca atau menulis file proyek, dokumen, artefak yang dihasilkan, atau status workspace lokal.
- Petunjuk izin berisiko tinggi: eksekusi shell atau perintah
- Quality score needs review
Target pemasangan
Pasang skill ini di alur Agent Anda
Gunakan endpoint publik untuk mengambil perintah, checklist keamanan, prompt target, dan tautan kanonis.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install nvidia-create-custom-graderRencana resolusi Agent
Biarkan Agent memverifikasi kecocokan sebelum memasang.
API Resolve mengembalikan skill utama, alternatif, kebijakan keamanan, catatan audit, target pemasangan, dan prompt siap pakai.
Buka JSON
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Teks Resolve
/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Serah-terima pemasangan
/api/skills/nvidia-create-custom-grader/install
Agent harus memeriksa
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Salin prompt
Task: Use create-custom-grader in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install
Install command: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Serah-terima Agent
Berikan jalur pemasangan kepada Agent, bukan direktori lain.
Gunakan endpoint publik untuk mengambil perintah, checklist keamanan, prompt target, dan tautan kanonis.
Serah-terima pemasangan
/api/skills/nvidia-create-custom-grader/install
Format teks LLM
/api/skills/nvidia-create-custom-grader/install?format=text
Cari alternatif
/api/skills/search?q=create-custom-grader&limit=3
Prompt Agent
Use create-custom-grader for this task. Review https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install, then install with: npx skills add NVIDIA/SkillEvaluator --skill create-custom-graderMetadata Registry
Profil yang dapat dibaca Agent untuk pemilihan skill otomatis.
API Registry menyediakan sinyal keputusan, kepercayaan, audit, use case, dan pemasangan tanpa mengikis UI.
Manifest
/api/registry/manifest/nvidia-create-custom-grader
Teks LLM
/api/registry/manifest/nvidia-create-custom-grader?format=text
Alias pemasangan
/api/registry/install/nvidia-create-custom-grader
Rekomendasikan
/api/registry/recommend?task=Use%20create-custom-grader%20in%20an%20agent%20workflow&limit=3
Kecocokan Agent
Browser automation
Platform
Claude Code
Laporan audit
Perlu ditinjau · 81/100
Tinjauan yang dapat dibaca mesin tentang kesiapan pemasangan, metadata keamanan, pemeliharaan, dan risiko adopsi.
Panel keputusan Agent
Fallback candidate for Browser automation
Prototype with this skill first; keep a fallback candidate ready.
Peran di stack
Kandidat cadangan
Kecocokan utama
Browser automation
Label kepercayaan
Buat prototipe dulu
Jalur pemasangan
Perintah siap
Gunakan saat
- alur kerja Browser automation
- Tim Claude Code
- builders willing to evaluate younger projects
Bukti
- recent repository activity
- install command or GitHub repo available
- profil kualitas 70/100
- 2 event interaksi OpenAgentSkill
tinjau dulu
- No major risk signals from current metadata
Jalur implementasi
- 1Pasang di Agent sandbox dan jalankan satu tugas Browser automation dari awal hingga akhir.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Profil kepercayaan
Hanya sandbox
Kandidat berguna dengan sinyal kepercayaan yang kurang atau bercampur. Gunakan di ruang kerja terisolasi hingga loop hasil membuktikan kecocokan tugas.
Adopsi GitHub
Info187 star GitHub
Aktivitas star/fork
Periksa187 star dan 14 fork; aktivitas issue tidak tersedia dalam metadata saat ini
Pemeliharaan terbaru
Lulus2 hari sejak push
Kejelasan lisensi
LulusApache-2.0
Sinyal positif
- Tinjauan AI disetujui
- Jalur pemasangan tersedia
- Bukti repositori tersedia
- Repositori yang baru dipelihara
- Perintah pemasangan tidak memiliki pola berisiko tinggi yang jelas
- Loop hasil siap tetapi membutuhkan eksekusi Agent nyata pertama
Tinjau sebelum memasang
- Quality score needs review
- Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
- Belum ada laporan hasil Agent nyata
- Tinjauan manusia diperlukan sebelum pemasangan tanpa pengawasan
Tindakan yang disarankan
Jalankan hanya dalam sandbox dan bandingkan alternatif terdekat sebelum digunakan untuk kerja nyata.
Profil kualitas
Kuat kandidat untuk alur kerja Agent
Solid option that is likely worth shortlisting for production workflows.
Kecocokan alur kerja
Gunakan skill ini pada skenario berikut
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
Kecocokan alur kerja
Tambahkan ke alur kerja lengkap
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Daftar alternatif
Bandingkan sebelum memasang
Similar skills that may fit this task.
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
MoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Cua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
Ringkasan
--- name: create-custom-grader description: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. metadata: author: SkillEvaluator Maintainers <maintainers@example.com> ---
# Create Custom Grader
Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.
## Purpose
Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.
## When To Use
Use this skill when the user wants to:
- bring an existing benchmark into SkillEvaluator - turn a rubric into `evals/grader.py` or `evals/grader.sh` - add custom metrics beside the default evaluator metrics - convert task files such as `task.yaml`, `task.json`, pytest checks, or shell verifiers into BYOG or BYOT - prove a team can run its own benchmark through SkillEvaluator
Do not use this skill for ordinary `evals/evals.json` authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.
## Instructions
1. Read the target skill, existing `evals/`, benchmark prompts, fixtures, and any verifier code. 2. Choose `default_plus_custom` when custom metrics should complement default evaluator scoring. 3. Choose `custom_only` only when the user wants the custom grader to own pass/fail semantics. 4. Write or update `evals/grader.py` or `evals/grader.sh`, then validate the Harbor contract.
## Examples
```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom skillevaluator tier3 validate <skill-dir> ```
## Prerequisites
- The target skill directory should contain `SKILL.md`. - The SkillEvaluator CLI should be available as `skillevaluator`. - Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.
## Core Choice
Choose one path before writing files:
| User need | Evaluator shape | | --- | --- | | Existing `evals.json` task plus extra domain checks | Top-level BYOG: `evals/grader.py` or `evals/grader.sh` | | Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus `evals/evals.json` and `evals/files/` | | Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: `evals/harbor/<case>/...` | | User wants only custom reward/pass criteria | `grading.mode: custom_only` | | User wants default evaluator dimensions plus custom metrics | `grading.mode: default_plus_custom` |
Default to `default_plus_custom` unless the user explicitly wants the custom grader to replace the default evaluator metrics.
## Workflow
1. Resolve the target skill and benchmark source. Read the target `SKILL.md`, existing `evals/`, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.
2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as `question` entries. Use the target skill as `expected_skill`. Put each case's required starter files under `evals/files/<case-id>/`, and declare `files: ["evals/files/<case-id>"]` on every corresponding eval entry. Do not omit `files` in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.
3. Scaffold the evaluator contract. For generated tasks: ```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom ``` For shell checks: ```bash skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom ``` For native Harbor tasks: ```bash skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config ```
4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.
5. Validate before running. ```bash skillevaluator validate <skill-dir> --harbor-contract ``` Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.
6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.
## Grader Contract
Python and shell graders run inside the Harbor verifier context. They may read:
- `/logs/agent/trajectory.json` for agent actions and final answer evidence - `/tests/entry.json` for the eval case metadata - `/workspace/input/` for the entry's declared committed fixtures from `evals/files/` - `/solution/` or other task outputs only when the task environment produces them
They must write:
- `/logs/verifier/reward.json` - `/logs/verifier/reward.txt` with a numeric score from `0.0` to `1.0`
Use this reward shape:
```json { "overall": 0.92, "custom_metrics": { "domain_repair": 1.0, "domain_verification": 0.8 }, "details": { "domain_repair": { "score": 1.0, "reason": "The solution repaired the required files." } } } ```
In `default_plus_custom`, default evaluator scoring keeps its `overall` authoritative and adds the grader's `custom_metrics` into reports. In `custom_only`, the grader's `overall` is the pass/fail reward.
Never emit custom metric names that collide with reserved evaluator fields: `security`, `skill_execution`, `skill_efficiency`, `accuracy`, `goal_accuracy`, `behavior_check`, `overall`, `details`, `metrics`, `metric_set`, or `entry_id`.
## Translation Rules
- Convert each rubric item into a deterministic check when possible. - If a rubric item requires judgment, encode observable proxies and explain the limits in `details`. - Keep metrics stable across baseline and with-skill runs. - Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates. - Keep custom metric values clamped to `0.0` through `1.0`. - Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior. - Convert expected trigger/non-trigger metadata into `expected_skill`, `expected_behavior`, negative cases, or custom metrics that inspect trajectory evidence.
## RAPIDS-Style Example
For a benchmark task with `task.yaml`, `code/`, prompt variants, coverage, and a rubric:
1. Copy `code/` into `evals/files/<case-id>/`. 2. Create one or more `evals/evals.json` entries from the prompt variants, and set `files: ["evals/files/<case-id>"]` on each corresponding entry. 3. Set `expected_skill` to the benchmark's target skill. 4. Implement `evals/grader.py` to inspect the agent trajectory and changed workspace files. 5. Emit custom metrics for each rubric criterion, for example `rapids_diagnosis`, `rapids_requirements_repair`, `rapids_repair_safety`, and `rapids_verification`. 6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.
## Limitations
- The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy. - `init-custom-grader` creates scaffolding only; the agent must replace the placeholder scoring logic. - Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.
## Troubleshooting
| Problem | Fix | | --- | --- | | `evals/evals.json` missing | Create entries from the benchmark prompt or run `init-custom-grader` to seed one. | | Custom metrics do not appear | Ensure `reward.json` has numeric values under `custom_metrics` and no reserved-name collisions. | | `custom_only` fails | Write numeric `overall` in `reward.json` or numeric `reward.txt`. | | Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |
## Final Response
When finished, report:
- files created or changed - exact validation and evaluation commands - default evaluator metric results - custom metric results - whether the proof was full E2E or only static/local validation - any benchmark rubric criteria that remain partly judgment-based
Detail teknis
- Versi
- 1.0.0
- Lisensi
- Apache-2.0
- Pembaruan terakhir
- 21 Agu 2026
- Diterbitkan
- 20 Agu 2026
Ringkasan keputusan
Kandidat cadangan
recent repository activity
Audit
Tinjauan pemasangan
Tinjauan pemasangan dan adopsi
- Keamanan
- 83/100
- Pemeliharaan
- 100/100
- Pasang
- 92/100
Bukti tervalidasi Agent
Bukti tervalidasi Agent
Laporan hasil setelah resolve, tinjau, pasang, dan satu eksekusi terbatas.
- Tingkat sukses
- —
- Kegagalan terbaru
- —
- Hasil
- 0
- Kualitas output
- —
- Gagal
- 0
- Tidak relevan
- 0
- Pemasangan
- 0
- Diblokir risiko
- 0
- Perlu penyiapan
- 0
- Produksi
- 0
Belum ada data hasil Agent. Eksekusi pertama dapat melaporkan keberhasilan, kebutuhan setup, blok risiko, kegagalan, atau tidak relevan melalui /api/agent/outcome.
Pasang
Tambahkan ke alur Agent
Gratis dan sumber terbuka. Tinjau laporan sebelum memasang pada Agent produksi.
Siklus pertumbuhan
Kit berbagi
Draf berbasis skenario untuk create-custom-grader, siap untuk posting manual di X.
create-custom-grader: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check... 187 stars https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x
Balasan opsional dengan perintah pemasangan
Listing + install path for create-custom-grader: https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x Install: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Sumber listing
Diindeks Registry
Listing ini diindeks dari sumber publik dan belum ditandai resmi hingga klaim pemelihara disetujui.
- Kreator
- NVIDIA
- Sumber
- NVIDIA/SkillEvaluator
- Diindeks oleh
- Indeks komunitas OpenAgentSkill
Atribusi menautkan ke repositori publik atau profil kreator. Kreator dapat mengklaim listing untuk memperbarui sinyal kepemilikan.
Klaim skill iniKlaim pemilik
Klaim listing skill ini
Listing Diindeks Registry ini dikaitkan dengan NVIDIA, tetapi belum ditandai resmi. Klaim untuk menambahkan sinyal pemilik terverifikasi dan membuat pembaruan peluncuran, pemasangan, serta audit berikutnya lebih tepercaya.
Kit backlink kreator
Tambahkan badge bukti ke README Anda
Tampilkan listing kanonis, sinyal kepercayaan dan audit saat ini, serta bukti Agent-Proven nyata di tempat pengembang mengevaluasi repositori.
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader/audit)
[](https://www.openagentskill.com/skills/nvidia-create-custom-grader)Penulis
NVIDIA
@nvidia
Tag
Kecocokan platform
Sinyal kesehatan
- Star GitHub
- 187
- Skor kualitas
- 39/100
- Push GitHub terakhir
- 21 Agu 2026
- Petunjuk framework
- Tidak diketahui
- Tampilan OpenAgentSkill
- 2
- Salinan pemasangan
- 0
- Klik keluar
- 0
Sinyal komunitas
Bagikan apakah skill ini bermanfaat untuk alur kerja Agent Anda. Masukan gabungan meningkatkan peringkat dari waktu ke waktu.
Kepercayaan & keamanan
Hanya sandbox
- Adopsi GitHub187 star GitHubInfo
- Aktivitas star/fork187 star dan 14 fork; aktivitas issue tidak tersedia dalam metadata saat iniPeriksa
- Pemeliharaan terbaru2 hari sejak pushLulus
- Kejelasan lisensiApache-2.0Lulus
- Kelengkapan README/SKILL.mdMetadata memuat konteks penggunaan dan alur kerja yang cukupLulus
- Risiko dependensi/runtimeCakupan eksekusi perintahInfo
Skill terkait
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K StarMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarCua
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
21.4K Star