create-custom-grader

검토 · 70
Registry 색인

Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

Verified installs0
스타187
버전1.0.0
품질70/100 · 강함
신뢰70/100 · 샌드박스 전용
감사81/100 · 검토 필요

공급 자산 프로필

리서치 및 지식 작업

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

트랙 보기

시나리오

RAG and knowledge

I need my agent to build a RAG workflow over documents and retrieve reliable context.

Agent 적합도

Claude Code + CLI + Codex

Codex, Claude Code, Cursor, CLI 또는 맞춤형 Agent에 적합합니다.

설치

준비됨

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

유지보수

최신

마지막 푸시 후 1일

위험

검토 필요

Quality score needs review

GitHub 품질

187

70/100 품질 · 78/100 신뢰

커버리지 태그

리서치RAG and knowledge자동화agent-skill

검토 메모

Quality score needs review · Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata

Agent 채택 스코어카드

신뢰, 감사, 설치 준비 상태를 한눈에 확인하세요

이 점수는 공개 저장소 메타데이터, OpenAgentSkill 검토 신호, 유지보수 최신성, 설치 준비 상태를 결합합니다. 후보 선정 신호일 뿐, 사람의 검토를 대체하지 않습니다.

품질

강함
70

프로덕션 워크플로 후보군에 넣을 만한 탄탄한 선택입니다.

신뢰

샌드박스 전용
70

신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.

감사

검토 필요
81

설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.

OpenAgentSkill 신뢰 점수 v5

설치 전 사람 검토

실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.

CodexClaude CodeCursorOpenAgentSkill CLI

스타

GitHub 스타 187

저장소 활동

스타 187, 포크 14

유지보수

마지막 푸시 후 1일

라이선스

Apache-2.0

설치

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

설치 안전성

표준 패키지 또는 런타임 설치 경로

권한 범위

shell or command execution, filesystem or document access

Agent 결과

아직 Agent 결과 데이터가 없습니다

문서

README/SKILL.md 맥락이 충분합니다

위험 요약

프로덕션 전 검토

  • Quality score needs review
  • Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata

설치 준비 상태

설치 경로 사용 가능

  • 설치 경로를 사용할 수 있습니다
  • 저장소 근거를 사용할 수 있습니다
  • 라이선스가 명시되었습니다
  • 아직 Agent 검증 결과 근거가 없습니다

Agent 읽기용 메타데이터

이 스킬의 기계 판독형 의사결정 데이터.

이 블록 또는 포함된 JSON을 사용해 Agent가 이 스킬을 설치할지, 대안을 고를지, 먼저 사람의 검토를 요청할지 판단할 수 있습니다.

JSON 열기

적합한 작업

  • Browser automation 워크플로
  • Claude Code 팀
  • builders willing to evaluate younger projects
  • Navigate pages

적합한 Agent

CodexClaude CodeCursorOpenAgentSkill CLICLI

설치 결정

명령어
npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
정책
검토
사람 검토

신뢰와 위험

신뢰
70/100
감사
81/100
위험 수준
검토 필요

결과 루프

엔드포인트
/api/agent/outcome
이벤트 ID
resolve
결과
5

설치 명령어

npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

사용하지 말아야 할 경우

  • 벤더 지원 SLA가 필요한 팀
  • 내부 보안 검토가 없는 고규정 준수 환경
  • 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다
  • 고위험 권한 힌트: Shell 또는 명령 실행
  • Quality score needs review

Agent 안전 v2

53/100 · 자동 설치 피하기

실험적검토

Sparse or mixed signals. Useful for discovery, but not for autonomous installation.

Test manually in an isolated workspace and compare against safer alternatives.

API로 해결

높음

Shell 또는 명령 실행

Skill 메타데이터가 터미널, CLI, Shell, 하위 프로세스 또는 명령 실행 워크플로를 참조합니다.

중간

네트워크 접근

Skill은 원격 페이지, API, 저장소 또는 외부 서비스에 접근할 수 있습니다.

중간

파일 시스템 접근

Skill은 프로젝트 파일, 문서, 생성 산출물 또는 로컬 작업 공간 상태를 읽거나 쓸 수 있습니다.

  • 고위험 권한 힌트: Shell 또는 명령 실행
  • Quality score needs review

설치 대상

Agent 워크플로에 이 스킬 설치

공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install nvidia-create-custom-grader

Agent 해결 계획

설치 전에 Agent가 적합성을 검증하게 하세요.

Resolve API는 최우선 스킬, 대안, 안전 정책, 감사 메모, 설치 대상 및 Agent가 페이지를 스크래핑하지 않고 사용할 수 있는 프롬프트를 반환합니다.

텍스트 계획 열기

Agent가 확인할 항목

  • Resolve API에서 작업 적합도와 대안을 확인합니다.
  • 감사 점수, 신뢰 점수 및 안전 정책 경고를 확인합니다.
  • Codex, Claude Code, Cursor 또는 CLI의 설치 대상 호환성을 확인합니다.

프롬프트 복사

Task: Use create-custom-grader in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20create-custom-grader%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install
Install command: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 핸드오프

또 다른 디렉터리 페이지 대신 설치 경로를 Agent에게 제공합니다.

공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.

설치 API 열기

Agent 프롬프트

Use create-custom-grader for this task. Review https://www.openagentskill.com/api/skills/nvidia-create-custom-grader/install, then install with: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

Registry 메타데이터

자동 스킬 선택을 위한 Agent 읽기용 프로필.

Registry API를 통해 동일한 결정, 신뢰, 감사, 사용 사례, 설치 신호를 제공하므로 Agent가 UI를 스크래핑하지 않고도 순위를 매길 수 있습니다.

Manifest 열기

Agent 적합도

69/100

Browser automation

플랫폼

Claude Code

감사 보고서

검토 필요 · 81/100

설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.

감사 보고서 보기평가 보고서 보기

Agent 결정 패널

Fallback candidate for Browser automation

먼저 이 스킬로 프로토타입을 만들고 대체 후보를 준비하세요.

69
준비 상태
프로토타입
단계

스택 내 역할

대체 후보

주요 적합도

Browser automation

신뢰 라벨

먼저 프로토타입

설치 경로

명령어 준비됨

사용 시점

  • Browser automation 워크플로
  • Claude Code 팀
  • builders willing to evaluate younger projects

근거

  • 최근 저장소 활동
  • 설치 명령 또는 GitHub 저장소를 사용할 수 있습니다
  • 품질 프로필 70/100

먼저 검토

  • 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다

구현 경로

  1. 1샌드박스 Agent에 설치하고 Browser automation 작업을 처음부터 끝까지 한 번 실행하세요.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

신뢰 프로필

샌드박스 전용

신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.

70
OpenAgentSkill 신뢰 점수

GitHub 채택도

정보

GitHub 스타 187

스타/포크 활동

확인

스타 187, 포크 14; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다

최근 유지보수

통과

마지막 푸시 후 1일

라이선스 명확성

통과

Apache-2.0

긍정 신호

  • AI 검토 승인됨
  • 설치 경로를 사용할 수 있습니다
  • 저장소 근거를 사용할 수 있습니다
  • 최근 유지보수된 저장소
  • 설치 명령에서 뚜렷한 고위험 패턴이 발견되지 않았습니다
  • 결과 루프는 준비되었지만 첫 실제 Agent 실행이 필요합니다

설치 전 검토

  • Quality score needs review
  • Stars/forks activity: 187 stars, 14 forks; issue activity unavailable in current metadata
  • 아직 실제 Agent 결과 보고서가 없습니다
  • 무인 설치 전에 사람 검토가 필요합니다

권장 작업

실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.

품질 프로필

강함 Agent 워크플로용 후보

프로덕션 워크플로 후보군에 넣을 만한 탄탄한 선택입니다.

70
GitHub 스타
187
최신성
1일 전
설치 준비됨
라이선스
Apache-2.0

워크플로 적합도

이 스킬을 사용할 시나리오

워크플로 적합도

완전한 워크플로에 추가

대안 후보

설치 전 비교

이 작업에 적합할 수 있는 유사 스킬입니다.

모두 비교

개요

--- name: create-custom-grader description: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. metadata: author: SkillEvaluator Maintainers <maintainers@example.com> ---

# Create Custom Grader

Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.

## Purpose

Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.

## When To Use

Use this skill when the user wants to:

- bring an existing benchmark into SkillEvaluator - turn a rubric into `evals/grader.py` or `evals/grader.sh` - add custom metrics beside the default evaluator metrics - convert task files such as `task.yaml`, `task.json`, pytest checks, or shell verifiers into BYOG or BYOT - prove a team can run its own benchmark through SkillEvaluator

Do not use this skill for ordinary `evals/evals.json` authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.

## Instructions

1. Read the target skill, existing `evals/`, benchmark prompts, fixtures, and any verifier code. 2. Choose `default_plus_custom` when custom metrics should complement default evaluator scoring. 3. Choose `custom_only` only when the user wants the custom grader to own pass/fail semantics. 4. Write or update `evals/grader.py` or `evals/grader.sh`, then validate the Harbor contract.

## Examples

```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom skillevaluator tier3 validate <skill-dir> ```

## Prerequisites

- The target skill directory should contain `SKILL.md`. - The SkillEvaluator CLI should be available as `skillevaluator`. - Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.

## Core Choice

Choose one path before writing files:

| User need | Evaluator shape | | --- | --- | | Existing `evals.json` task plus extra domain checks | Top-level BYOG: `evals/grader.py` or `evals/grader.sh` | | Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus `evals/evals.json` and `evals/files/` | | Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: `evals/harbor/<case>/...` | | User wants only custom reward/pass criteria | `grading.mode: custom_only` | | User wants default evaluator dimensions plus custom metrics | `grading.mode: default_plus_custom` |

Default to `default_plus_custom` unless the user explicitly wants the custom grader to replace the default evaluator metrics.

## Workflow

1. Resolve the target skill and benchmark source. Read the target `SKILL.md`, existing `evals/`, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.

2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as `question` entries. Use the target skill as `expected_skill`. Put each case's required starter files under `evals/files/<case-id>/`, and declare `files: ["evals/files/<case-id>"]` on every corresponding eval entry. Do not omit `files` in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.

3. Scaffold the evaluator contract. For generated tasks: ```bash skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom ``` For shell checks: ```bash skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom ``` For native Harbor tasks: ```bash skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config ```

4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.

5. Validate before running. ```bash skillevaluator validate <skill-dir> --harbor-contract ``` Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.

6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.

## Grader Contract

Python and shell graders run inside the Harbor verifier context. They may read:

- `/logs/agent/trajectory.json` for agent actions and final answer evidence - `/tests/entry.json` for the eval case metadata - `/workspace/input/` for the entry's declared committed fixtures from `evals/files/` - `/solution/` or other task outputs only when the task environment produces them

They must write:

- `/logs/verifier/reward.json` - `/logs/verifier/reward.txt` with a numeric score from `0.0` to `1.0`

Use this reward shape:

```json { "overall": 0.92, "custom_metrics": { "domain_repair": 1.0, "domain_verification": 0.8 }, "details": { "domain_repair": { "score": 1.0, "reason": "The solution repaired the required files." } } } ```

In `default_plus_custom`, default evaluator scoring keeps its `overall` authoritative and adds the grader's `custom_metrics` into reports. In `custom_only`, the grader's `overall` is the pass/fail reward.

Never emit custom metric names that collide with reserved evaluator fields: `security`, `skill_execution`, `skill_efficiency`, `accuracy`, `goal_accuracy`, `behavior_check`, `overall`, `details`, `metrics`, `metric_set`, or `entry_id`.

## Translation Rules

- Convert each rubric item into a deterministic check when possible. - If a rubric item requires judgment, encode observable proxies and explain the limits in `details`. - Keep metrics stable across baseline and with-skill runs. - Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates. - Keep custom metric values clamped to `0.0` through `1.0`. - Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior. - Convert expected trigger/non-trigger metadata into `expected_skill`, `expected_behavior`, negative cases, or custom metrics that inspect trajectory evidence.

## RAPIDS-Style Example

For a benchmark task with `task.yaml`, `code/`, prompt variants, coverage, and a rubric:

1. Copy `code/` into `evals/files/<case-id>/`. 2. Create one or more `evals/evals.json` entries from the prompt variants, and set `files: ["evals/files/<case-id>"]` on each corresponding entry. 3. Set `expected_skill` to the benchmark's target skill. 4. Implement `evals/grader.py` to inspect the agent trajectory and changed workspace files. 5. Emit custom metrics for each rubric criterion, for example `rapids_diagnosis`, `rapids_requirements_repair`, `rapids_repair_safety`, and `rapids_verification`. 6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.

## Limitations

- The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy. - `init-custom-grader` creates scaffolding only; the agent must replace the placeholder scoring logic. - Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.

## Troubleshooting

| Problem | Fix | | --- | --- | | `evals/evals.json` missing | Create entries from the benchmark prompt or run `init-custom-grader` to seed one. | | Custom metrics do not appear | Ensure `reward.json` has numeric values under `custom_metrics` and no reserved-name collisions. | | `custom_only` fails | Write numeric `overall` in `reward.json` or numeric `reward.txt`. | | Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |

## Final Response

When finished, report:

- files created or changed - exact validation and evaluation commands - default evaluator metric results - custom metric results - whether the proof was full E2E or only static/local validation - any benchmark rubric criteria that remain partly judgment-based

기술 세부 사항

버전
1.0.0
라이선스
Apache-2.0
최근 업데이트
2026년 8월 21일
게시일
2026년 8월 20일

결정 스냅샷

대체 후보

69
준비됨
프로토타입
단계

최근 저장소 활동

감사

설치 검토

설치 및 채택 검토

81
검토 필요
보안
83/100
유지보수
100/100
설치
92/100
전체 감사 열기평가 보고서 보기

Agent 검증 증거

Agent 검증 증거

Resolve, 검토, 설치 및 한 번의 제한된 실행 후 결과 보고서입니다.

0
검증됨
Needs first agent run자동 설치: 먼저 검토최근: 알 수 없음
성공률
최근 실패
결과
0
출력 품질
실패
0
관련 없음
0
설치
0
위험 차단
0
설정 필요
0
프로덕션
0

아직 Agent 결과 데이터가 없습니다. 첫 실행은 /api/agent/outcome을 통해 성공, 설정 필요, 위험 차단, 실패 또는 비관련 결과를 보고할 수 있습니다.

설치

Agent 워크플로에 추가

무료 오픈 소스. 프로덕션 Agent에 설치하기 전에 보고서를 검토하세요.

성장 루프

공유 키트

X

create-custom-grader용 시나리오 기반 초안입니다. X에 수동으로 게시할 수 있습니다.

큐레이터 노트
create-custom-grader: Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check...

187 stars

https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x
X 초안 열기
선택 사항: 설치 명령이 포함된 답글
Listing + install path for create-custom-grader:
https://www.openagentskill.com/skills/nvidia-create-custom-grader?ref=x

Install: npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader

등록 출처

Registry 색인

소유권 주장 가능

이 등록은 공개 소스에서 색인되었으며 유지보수자 소유권 주장이 승인될 때까지 공식으로 표시되지 않습니다.

제작자
NVIDIA
색인 주체
OpenAgentSkill 커뮤니티 인덱스

귀속은 공개 저장소 또는 제작자 프로필에 연결됩니다. 제작자는 등록을 주장하여 소유권 신호를 업데이트할 수 있습니다.

이 스킬 소유권 주장

소유자 소유권 주장

이 스킬 등록 소유권 주장

이 Registry 색인 등록은 NVIDIA에게 귀속되어 있지만 아직 공식으로 표시되지 않았습니다. 소유권을 주장하면 확인된 소유자 신호가 추가되어 이후 출시, 설치 및 감사 업데이트를 더 신뢰할 수 있습니다.

크리에이터 백링크 키트

README에 증거 배지 추가

개발자가 저장소를 평가하는 위치에 정규 등록, 현재 신뢰 및 감사 신호, 실제 Agent-Proven 증거를 표시합니다.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=listed&label=Listed)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=trust&label=Trust)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=audit&label=Audit)](https://www.openagentskill.com/skills/nvidia-create-custom-grader/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/nvidia-create-custom-grader?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/nvidia-create-custom-grader)

작성자

N

NVIDIA

@nvidia

플랫폼 적합도

상태 신호

GitHub 스타
187
품질 점수
39/100
최근 GitHub 푸시
2026년 8월 21일
프레임워크 힌트
알 수 없음
OpenAgentSkill 조회수
0
설치 명령 복사
0
외부 클릭
0

커뮤니티 신호

이 스킬이 Agent 워크플로에 유용한지 알려 주세요. 집계된 피드백은 시간이 지날수록 순위를 개선합니다.

신뢰와 안전

샌드박스 전용

70
  • GitHub 채택도GitHub 스타 187정보
  • 스타/포크 활동스타 187, 포크 14; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다확인
  • 최근 유지보수마지막 푸시 후 1일통과
  • 라이선스 명확성Apache-2.0통과
  • README/SKILL.md 완성도메타데이터에 충분한 사용 및 워크플로 맥락이 포함되어 있습니다통과
  • 의존성/런타임 위험명령 실행 범위정보