create-custom-grader

レビュー · 70
Registry に収録

Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

Verified installs0
スター187
バージョン1.0.0
品質70/100 · 強い
信頼70/100 · サンドボックス限定
監査81/100 · 要レビュー

供給アセットの概要

リサーチとナレッジ作業

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

カテゴリを見る

シナリオ

RAG and knowledge

I need my agent to build a RAG workflow over documents and retrieve reliable context.

Agent 適合

Claude Code + CLI + Codex

Codex、Claude Code、Cursor、CLI、またはカスタム Agent に対応します。

インストール

準備完了

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

メンテナンス

新しい

最終プッシュから 1 日

リスク

要レビュー

Quality score needs review

GitHub 品質

187

70/100 品質 · 78/100 信頼

対象タグ

リサーチRAG and knowledge自動化agent-skill

レビュー注記

Quality score needs review · Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata

Agent 導入スコアカード

信頼、監査、インストール準備状況を一目で確認

公開リポジトリのメタデータ、OpenAgentSkill のレビューシグナル、保守の鮮度、インストール準備状況を組み合わせたスコアです。候補選定の目安であり、人によるレビューの代替ではありません。

品質

強い
70

本番ワークフローの候補に値する堅実な選択肢です。

信頼

サンドボックス限定
70

信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。

監査

要レビュー
81

インストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。

OpenAgentSkill Trust Score v5

インストール前に人のレビュー

実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。

CodexClaude CodeCursorOpenAgentSkill CLI

スター

GitHub スター 187

リポジトリ活動

スター 187、フォーク 14

メンテナンス

最終プッシュから 1 日

ライセンス

Apache-2.0

インストール

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

インストール安全性

標準パッケージまたはランタイムのインストールパス

権限範囲

shell or command execution, filesystem or document access

Agent の成果

Agent の成果データはまだありません

ドキュメント

README/SKILL.md の文脈が十分です

リスク概要

本番前にレビュー

  • Quality score needs review
  • Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata

インストール準備状況

インストールパスを利用可能

  • インストールパスを利用できます
  • リポジトリの根拠を利用できます
  • ライセンスが明示されています
  • Agent-Proven の成果エビデンスはまだありません

Agent 可読メタデータ

このスキルの機械可読な判断データ。

このブロックまたは埋め込み JSON を使い、Agent がこのスキルをインストールすべきか、代替を選ぶべきか、先に人のレビューを求めるべきかを判断できます。

JSON を開く

適したタスク

  • Browser automation ワークフロー
  • Claude Code チーム
  • builders willing to evaluate younger projects
  • Navigate pages

適した Agent

CodexClaude CodeCursorOpenAgentSkill CLICLI

インストール判断

コマンド
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
ポリシー
レビュー
人によるレビュー
はい

信頼とリスク

信頼
70/100
監査
81/100
リスクレベル
要レビュー

成果ループ

エンドポイント
/api/agent/outcome
イベント ID
resolve
成果
5

インストールコマンド

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

使わない場合

  • ベンダー提供の SLA が必要なチーム
  • 内部セキュリティレビューのない高コンプライアンス環境
  • OpenAgentSkill の利用フィードバックはまだありません
  • 高リスク権限のヒント: Shell またはコマンド実行
  • Quality score needs review

Agent セーフティ v2

53/100 · 自動インストールを避ける

実験的レビュー

Sparse or mixed signals. Useful for discovery, but not for autonomous installation.

Test manually in an isolated workspace and compare against safer alternatives.

API で解決

Shell またはコマンド実行

Skill メタデータに端末、CLI、Shell、サブプロセス、またはコマンド実行のワークフローが含まれます。

ネットワークアクセス

Skill はリモートページ、API、リポジトリ、外部サービスにアクセスする可能性があります。

ファイルシステムアクセス

Skill はプロジェクトファイル、ドキュメント、生成物、ローカルワークスペース状態を読み書きする可能性があります。

  • 高リスク権限のヒント: Shell またはコマンド実行
  • Quality score needs review

インストール先

Agent ワークフローにこのスキルをインストール

公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install nvidia-create-custom-grader

Agent 解決プラン

インストール前に Agent に適合性を検証させます。

Resolve API は第一候補、代替、安全ポリシー、監査メモ、インストール先、Agent がそのまま使えるプロンプトを返します。

テキストプランを開く

Agent が確認すべきこと

  • Resolve API でタスク適合と代替を確認。
  • 監査・信頼スコアと安全ポリシーの警告を確認。
  • Codex、Claude Code、Cursor、CLI のインストール先互換性を確認。

プロンプトをコピー

Task: Use create-custom-grader in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install
Install command: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 引き継ぎ

別のディレクトリではなく、インストール経路を Agent に渡します。

公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。

Install API を開く

Agent プロンプト

Use create-custom-grader for this task. Review https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install, then install with: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

Registry メタデータ

自動スキル選択用の Agent 可読プロファイル。

Registry API 経由で判断、信頼、監査、ユースケース、インストールのシグナルを提供し、UI をスクレイピングせずに Agent が順位付けできます。

Manifest を開く

Agent 適合

69/100

Browser automation

プラットフォーム

Claude Code

監査レポート

要レビュー · 81/100

インストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。

監査レポートを見る評価レポートを見る

Agent 判断パネル

Fallback candidate for Browser automation

まずこのスキルでプロトタイプを作り、代替候補を用意してください。

69
準備状況
プロトタイプ
段階

スタック内の役割

代替候補

主な適合

Browser automation

信頼ラベル

まずプロトタイプ

インストールパス

コマンド準備済み

使う場面

  • Browser automation ワークフロー
  • Claude Code チーム
  • builders willing to evaluate younger projects

根拠

  • 最近のリポジトリ活動
  • インストールコマンドまたは GitHub リポジトリが利用可能
  • 品質プロファイル 70/100

先にレビュー

  • OpenAgentSkill の利用フィードバックはまだありません

実装パス

  1. 1サンドボックスの Agent にインストールし、Browser automation タスクを一度最初から最後まで実行します。
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

信頼プロファイル

サンドボックス限定

信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。

70
OpenAgentSkill Trust Score

GitHub 採用度

情報

GitHub スター 187

スター/フォーク活動

確認

スター 187、フォーク 14; 現在のメタデータでは Issue 活動を利用できません

最近のメンテナンス

合格

最終プッシュから 1 日

ライセンスの明確さ

合格

Apache-2.0

良いシグナル

  • AI レビュー承認済み
  • インストールパスを利用できます
  • リポジトリの根拠を利用できます
  • 最近保守されたリポジトリ
  • インストールコマンドに明確な高リスクパターンはありません
  • 成果ループは準備済みですが、最初の実行が必要です

インストール前にレビュー

  • Quality score needs review
  • Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
  • 実際の Agent 成果レポートはまだありません
  • 無人インストールの前に人によるレビューが必要です

推奨アクション

実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。

品質プロファイル

強い Agent ワークフロー向けの候補

本番ワークフローの候補に値する堅実な選択肢です。

70
GitHub スター
187
鮮度
1 日前
インストール準備完了
はい
ライセンス
Apache-2.0

ワークフロー適合

このスキルを使うシナリオ

ワークフロー適合

完全なワークフローに追加

代替候補

インストール前に比較

このタスクに適する可能性のある類似スキル。

すべて比較

概要

--- name: create-custom-grader description: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. metadata: author: SkillEvaluator Maintainers <maintainers@example.com> ---

# Create Custom Grader

Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.

## Purpose

Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.

## When To Use

Use this skill when the user wants to:

- bring an existing benchmark into SkillEvaluator - turn a rubric into `evals/grader.py` or `evals/grader.sh` - add custom metrics beside the default evaluator metrics - convert task files such as `task.yaml`, `task.json`, pytest checks, or shell verifiers into BYOG or BYOT - prove a team can run its own benchmark through SkillEvaluator

Do not use this skill for ordinary `evals/evals.json` authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.

## Instructions

1. Read the target skill, existing `evals/`, benchmark prompts, fixtures, and any verifier code. 2. Choose `default_plus_custom` when custom metrics should complement default evaluator scoring. 3. Choose `custom_only` only when the user wants the custom grader to own pass/fail semantics. 4. Write or update `evals/grader.py` or `evals/grader.sh`, then validate the Harbor contract.

## Examples

```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom skillevaluator tier3 validate <skill-dir> ```

## Prerequisites

- The target skill directory should contain `SKILL.md`. - The SkillEvaluator CLI should be available as `skillevaluator`. - Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.

## Core Choice

Choose one path before writing files:

| User need | Evaluator shape | | --- | --- | | Existing `evals.json` task plus extra domain checks | Top-level BYOG: `evals/grader.py` or `evals/grader.sh` | | Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus `evals/evals.json` and `evals/files/` | | Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: `evals/harbor/<case>/...` | | User wants only custom reward/pass criteria | `grading.mode: custom_only` | | User wants default evaluator dimensions plus custom metrics | `grading.mode: default_plus_custom` |

Default to `default_plus_custom` unless the user explicitly wants the custom grader to replace the default evaluator metrics.

## Workflow

1. Resolve the target skill and benchmark source. Read the target `SKILL.md`, existing `evals/`, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.

2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as `question` entries. Use the target skill as `expected_skill`. Put each case's required starter files under `evals/files/<case-id>/`, and declare `files: ["evals/files/<case-id>"]` on every corresponding eval entry. Do not omit `files` in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.

3. Scaffold the evaluator contract. For generated tasks: ```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom ``` For shell checks: ```bash skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom ``` For native Harbor tasks: ```bash skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config ```

4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.

5. Validate before running. ```bash skillevaluator validate <skill-dir> --harbor-contract ``` Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.

6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.

## Grader Contract

Python and shell graders run inside the Harbor verifier context. They may read:

- `/logs/agent/trajectory.json` for agent actions and final answer evidence - `/tests/entry.json` for the eval case metadata - `/workspace/input/` for the entry's declared committed fixtures from `evals/files/` - `/solution/` or other task outputs only when the task environment produces them

They must write:

- `/logs/verifier/reward.json` - `/logs/verifier/reward.txt` with a numeric score from `0.0` to `1.0`

Use this reward shape:

```json { "overall": 0.92, "custom_metrics": { "domain_repair": 1.0, "domain_verification": 0.8 }, "details": { "domain_repair": { "score": 1.0, "reason": "The solution repaired the required files." } } } ```

In `default_plus_custom`, default evaluator scoring keeps its `overall` authoritative and adds the grader's `custom_metrics` into reports. In `custom_only`, the grader's `overall` is the pass/fail reward.

Never emit custom metric names that collide with reserved evaluator fields: `security`, `skill_execution`, `skill_efficiency`, `accuracy`, `goal_accuracy`, `behavior_check`, `overall`, `details`, `metrics`, `metric_set`, or `entry_id`.

## Translation Rules

- Convert each rubric item into a deterministic check when possible. - If a rubric item requires judgment, encode observable proxies and explain the limits in `details`. - Keep metrics stable across baseline and with-skill runs. - Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates. - Keep custom metric values clamped to `0.0` through `1.0`. - Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior. - Convert expected trigger/non-trigger metadata into `expected_skill`, `expected_behavior`, negative cases, or custom metrics that inspect trajectory evidence.

## RAPIDS-Style Example

For a benchmark task with `task.yaml`, `code/`, prompt variants, coverage, and a rubric:

1. Copy `code/` into `evals/files/<case-id>/`. 2. Create one or more `evals/evals.json` entries from the prompt variants, and set `files: ["evals/files/<case-id>"]` on each corresponding entry. 3. Set `expected_skill` to the benchmark's target skill. 4. Implement `evals/grader.py` to inspect the agent trajectory and changed workspace files. 5. Emit custom metrics for each rubric criterion, for example `rapids_diagnosis`, `rapids_requirements_repair`, `rapids_repair_safety`, and `rapids_verification`. 6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.

## Limitations

- The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy. - `init-custom-grader` creates scaffolding only; the agent must replace the placeholder scoring logic. - Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.

## Troubleshooting

| Problem | Fix | | --- | --- | | `evals/evals.json` missing | Create entries from the benchmark prompt or run `init-custom-grader` to seed one. | | Custom metrics do not appear | Ensure `reward.json` has numeric values under `custom_metrics` and no reserved-name collisions. | | `custom_only` fails | Write numeric `overall` in `reward.json` or numeric `reward.txt`. | | Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |

## Final Response

When finished, report:

- files created or changed - exact validation and evaluation commands - default evaluator metric results - custom metric results - whether the proof was full E2E or only static/local validation - any benchmark rubric criteria that remain partly judgment-based

技術詳細

バージョン
1.0.0
ライセンス
Apache-2.0
最終更新
2026年8月21日
公開日
2026年8月20日

判断の要約

代替候補

69
準備完了
プロトタイプ
段階

最近のリポジトリ活動

監査

インストールレビュー

インストールと採用のレビュー

81
要レビュー
セキュリティ
83/100
メンテナンス
100/100
インストール
92/100
完全な監査を開く評価レポートを見る

Agent 実証エビデンス

Agent 実証エビデンス

Resolve、レビュー、インストール、限定実行後の成果レポート。

0
実証済み
Needs first agent run自動インストール: 先にレビュー最新: 不明
成功率
直近の失敗
成果
0
出力品質
失敗
0
非該当
0
インストール数
0
リスクによりブロック
0
設定が必要
0
本番
0

Agent の実行結果はまだありません。最初の実行では /api/agent/outcome を通じて成功、設定要件、リスクによるブロック、失敗、非該当を報告できます。

インストール

Agent ワークフローに追加

無料・オープンソース. 本番 Agent にインストールする前にレポートを確認してください。

成長ループ

共有キット

X

create-custom-grader 用のシナリオベース草案です。X へ手動投稿できます。

キュレーターノート
create-custom-grader: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check...

187 stars

https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x
X 下書きを開く
任意:インストールコマンド付きの返信
Listing + install path for create-custom-grader:
https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x

Install: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
返信の下書きを開く

掲載元

Registry により登録

申請可能

この掲載は公開ソースから登録されており、メンテナー申請が承認されるまで公式として表示されません。

作成者
NVIDIA
インデックス作成者
OpenAgentSkill コミュニティインデックス

帰属は公開リポジトリまたは作成者プロフィールにリンクされています。作成者は掲載を申請して所有権シグナルを更新できます。

このスキルを申請

所有者の申請

このスキル掲載を申請

この Registry により登録 掲載は NVIDIA に帰属していますが、まだ公式として表示されていません。申請すると、確認済み所有者シグナルが追加され、今後の公開、インストール、監査更新の信頼性が高まります。

クリエイター被リンクキット

README にエビデンスバッジを追加

開発者がリポジトリを評価する場所で、正規掲載、現在の信頼・監査シグナル、実際の Agent-Proven エビデンスを表示します。

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=listed&label=Listed)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=trust&label=Trust)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=audit&label=Audit)](https://www.openagentskill.com/skills/nvidia-create-custom-grader/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)

作者

N

NVIDIA

@nvidia

プラットフォーム適合

健全性シグナル

GitHub スター
187
品質スコア
39/100
最終 GitHub プッシュ
2026年8月21日
フレームワークのヒント
不明
OpenAgentSkill 閲覧数
0
インストールコピー数
0
外部クリック
0

コミュニティシグナル

このスキルが Agent ワークフローに役立つかを共有してください。集約されたフィードバックがランキングを改善します。

信頼と安全性

サンドボックス限定

70
  • GitHub 採用度GitHub スター 187情報
  • スター/フォーク活動スター 187、フォーク 14; 現在のメタデータでは Issue 活動を利用できません確認
  • 最近のメンテナンス最終プッシュから 1 日合格
  • ライセンスの明確さApache-2.0合格
  • README/SKILL.md の完全性メタデータには十分な利用・ワークフロー文脈があります合格
  • 依存関係/ランタイムのリスクコマンド実行範囲情報