@kennethkhoocy

제작자 · Claude Code

최근 업데이트 · 2026년 8월 24일

llm-gold-bound-failure-check

검토 · 71Registry 색인

Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/va

OpenAgentSkill 신뢰 점수
71/100

샌드박스 전용

품질64/100
감사80/100
스타47
Verified installs0

설치 대상

Codex 설치 프롬프트

Install the "llm-gold-bound-failure-check" agent skill from https://github.com/kennethkhoocy/applied-micro-skills/tree/main/plugins/applied-micro/skills/llm-gold-bound-failure-check. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"kennethkhoocy-llm-gold-bound-failure-check","task":"Install llm-gold-bound-failure-check","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.

공급 자산 프로필

디자인 및 크리에이티브 제작

Design assets, images, video, audio, multimodal media, presentation, and creative production skills.

트랙 보기

시나리오

디자인 및 크리에이티브

I need my agent to produce design assets, UI directions, presentations, or creative media workflows.

Agent 적합도

Claude Code + CLI + Codex

Codex, Claude Code, Cursor, CLI 또는 맞춤형 Agent에 적합합니다.

설치

준비됨

npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

유지보수

최신

오늘 푸시됨

위험

검토 필요

Low GitHub adoption signal

GitHub 품질

47

64/100 품질 · 79/100 신뢰

커버리지 태그

디자인디자인 및 크리에이티브디자인 및 크리에이티브agent-skill

검토 메모

Low GitHub adoption signal · Quality score needs review

Agent 채택 스코어카드

신뢰, 감사, 설치 준비 상태를 한눈에 확인하세요

이 점수는 공개 저장소 메타데이터, OpenAgentSkill 검토 신호, 유지보수 최신성, 설치 준비 상태를 결합합니다. 후보 선정 신호일 뿐, 사람의 검토를 대체하지 않습니다.

품질

유망
64

유용한 후보이지만 채택 전에 대안과 비교하세요.

신뢰

샌드박스 전용
71

신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.

감사

검토 필요
80

설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.

OpenAgentSkill 신뢰 점수 v5

설치 전 사람 검토

실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.

CodexClaude CodeCursorOpenAgentSkill CLI

스타

GitHub 스타 47

저장소 활동

스타 47, 포크 0

유지보수

오늘 푸시됨

라이선스

MIT

설치

npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

설치 안전성

표준 패키지 또는 런타임 설치 경로

권한 범위

파일 시스템 또는 문서 접근

Agent 결과

아직 Agent 결과 데이터가 없습니다

문서

README/SKILL.md 맥락이 충분합니다

위험 요약

프로덕션 전 검토

  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata

설치 준비 상태

설치 경로 사용 가능

  • 설치 경로를 사용할 수 있습니다
  • 저장소 근거를 사용할 수 있습니다
  • 라이선스가 명시되었습니다
  • 아직 Agent 검증 결과 근거가 없습니다

Agent 읽기용 메타데이터

이 스킬의 기계 판독형 의사결정 데이터.

이 블록 또는 포함된 JSON을 사용해 Agent가 이 스킬을 설치할지, 대안을 고를지, 먼저 사람의 검토를 요청할지 판단할 수 있습니다.

View technical data+

적합한 작업

  • Document processing 워크플로
  • Claude Code 팀
  • builders willing to evaluate younger projects
  • Read uploaded files

적합한 Agent

CodexClaude CodeCursorOpenAgentSkill CLICLI

설치 결정

명령어
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
정책
검토
사람 검토

신뢰와 위험

신뢰
71/100
감사
80/100
위험 수준
검토 필요

결과 루프

엔드포인트
/api/agent/outcome
이벤트 ID
resolve
결과
5

설치 명령어

npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

사용하지 말아야 할 경우

  • 벤더 지원 SLA가 필요한 팀
  • production agents without a repository review
  • Low GitHub adoption signal
  • 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다
  • Quality score needs review
가까운 대체 스킬이 아직 색인되지 않았습니다.

Agent 안전 v2

64/100 · 설치 전 검토

권한 메모와 함께 검토됨검토

사용 가능한 후보이지만 Agent는 설치 전에 권한과 감사 메모를 표시해야 합니다.

실제 작업 공간에 설치하기 전에 사람의 승인이 필요합니다.

API로 해결

중간

네트워크 접근

Skill은 원격 페이지, API, 저장소 또는 외부 서비스에 접근할 수 있습니다.

중간

파일 시스템 접근

Skill은 프로젝트 파일, 문서, 생성 산출물 또는 로컬 작업 공간 상태를 읽거나 쓸 수 있습니다.

  • Low GitHub adoption signal

Agent 해결 계획

설치 전에 Agent가 적합성을 검증하게 하세요.

Resolve API는 최우선 스킬, 대안, 안전 정책, 감사 메모, 설치 대상 및 Agent가 페이지를 스크래핑하지 않고 사용할 수 있는 프롬프트를 반환합니다.

텍스트 계획 열기

Agent가 확인할 항목

  • Resolve API에서 작업 적합도와 대안을 확인합니다.
  • 감사 점수, 신뢰 점수 및 안전 정책 경고를 확인합니다.
  • Codex, Claude Code, Cursor 또는 CLI의 설치 대상 호환성을 확인합니다.

프롬프트 복사

Task: Use llm-gold-bound-failure-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 핸드오프

또 다른 디렉터리 페이지 대신 설치 경로를 Agent에게 제공합니다.

공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.

설치 API 열기

Agent 프롬프트

Use llm-gold-bound-failure-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

Registry 메타데이터

자동 스킬 선택을 위한 Agent 읽기용 프로필.

Registry API를 통해 동일한 결정, 신뢰, 감사, 사용 사례, 설치 신호를 제공하므로 Agent가 UI를 스크래핑하지 않고도 순위를 매길 수 있습니다.

Manifest 열기

Agent 적합도

63/100

Document processing

플랫폼

Claude Code

감사 보고서

검토 필요 · 80/100

설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.

감사 보고서 보기평가 보고서 보기

Agent 결정 패널

Fallback candidate for Document processing

먼저 이 스킬로 프로토타입을 만들고 대체 후보를 준비하세요.

63
준비 상태
프로토타입
단계

스택 내 역할

대체 후보

주요 적합도

Document processing

신뢰 라벨

먼저 프로토타입

설치 경로

명령어 준비됨

사용 시점

  • Document processing 워크플로
  • Claude Code 팀
  • builders willing to evaluate younger projects

근거

  • 최근 저장소 활동
  • 설치 명령 또는 GitHub 저장소를 사용할 수 있습니다
  • 품질 프로필 64/100

먼저 검토

  • Low GitHub adoption signal
  • 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다

구현 경로

  1. 1샌드박스 Agent에 설치하고 Document processing 작업을 처음부터 끝까지 한 번 실행하세요.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

신뢰 프로필

샌드박스 전용

신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.

71
OpenAgentSkill 신뢰 점수

GitHub 채택도

확인

GitHub 스타 47

스타/포크 활동

확인

스타 47, 포크 0; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다

최근 유지보수

통과

오늘 푸시됨

라이선스 명확성

통과

MIT

긍정 신호

  • AI 검토 승인됨
  • 설치 경로를 사용할 수 있습니다
  • 저장소 근거를 사용할 수 있습니다
  • 최근 유지보수된 저장소
  • 설치 명령에서 뚜렷한 고위험 패턴이 발견되지 않았습니다
  • 결과 루프는 준비되었지만 첫 실제 Agent 실행이 필요합니다

설치 전 검토

  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
  • 아직 실제 Agent 결과 보고서가 없습니다
  • 무인 설치 전에 사람 검토가 필요합니다

권장 작업

실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.

품질 프로필

유망 Agent 워크플로용 후보

유용한 후보이지만 채택 전에 대안과 비교하세요.

64
GitHub 스타
47
최신성
오늘
설치 준비됨
라이선스
MIT
설치 전 검토: Low GitHub adoption signal

워크플로 적합도

이 스킬을 사용할 시나리오

워크플로 적합도

완전한 워크플로에 추가

개요

--- name: llm-gold-bound-failure-check description: | Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). author: Claude Code version: 1.0.0 date: 2026-07-16 ---

# LLM Gold-Bound Failure Check

## Problem

When an LLM scoring pipeline over-predicts one label, the reflex fix is a prompt clarification ("score positive ONLY when..."). But if the gold standard itself does not separate the texts you want excluded from the texts it labels positive, the revision removes true and false positives together. The pilot fails, the spend is wasted, and — worse — an un-gated adoption would have silently destroyed recall in production.

## Context / Trigger Conditions

- A domain/label shows precision ≪ recall (e.g. P 0.46 / R 0.96) against gold - A prompt edit is proposed to exclude a specific text type (boilerplate, affirmative-program language, non-risk framing) - The label's gold council/inter-rater agreement was already the weakest (κ below ~0.6 is the warning sign that the construct is contested)

## Solution

**Step 0 — the ~$0 check, BEFORE building anything:** read a sample of gold POSITIVES for the weak label and ask: do they contain the feature the revision would exclude? Compare them side-by-side with the false positives.

- Gold positives and false positives are the same kind of text → the failure is **gold-bound**. Stop. No prompt passes a gold-scored gate. The levers are: (a) re-adjudicate the construct with the gold's owners (changes the gold, not the scores), or (b) re-interpret the shipped measure honestly (e.g. "discussion salience" instead of "risk exposure") in downstream analyses. - Gold positives clearly differ from the false positives → a prompt revision is plausible; proceed to a gated pilot.

**Gated pilot design (verified):** 1. Split gold into tune/holdout halves, stratified on the weak label's positives; fixed seed. 2. Draft ONE surgical edit from tune-half errors only — byte-identical elsewhere; verify the diff reverses cleanly. 3. Pre-register the gate on the holdout BEFORE scoring: target-label thresholds (e.g. precision ≥ X AND recall ≥ Y) plus a perturbation tolerance for untouched labels (e.g. within 0.03 F1 / 0.06 κ of a same-serving-rev fresh baseline). 4. Score everything fresh under both prompts (same model revision, same day — this doubles as the drift control). Never write through the production cache layer. 5. Adopt only on a full pass; a REJECT is a valid, cheap outcome.

## Verification

The pilot report shows: the exact prompt diff, tune-vs-holdout metrics for old and new prompts, per-label deltas on untouched sections, and spend. A gold-bound diagnosis is confirmed when the revision moves recall sharply down while precision stays roughly flat.

## Example

Specialist Directors US, 2026-07-16: DEI over-prediction (P 0.46 / R 0.96, council κ 0.24–0.59). A risk-framing-only DEI clause was piloted ($1.17, pre-registered holdout gate). Result: recall 0.895→0.263, precision 0.455 (gate ≥0.60) — REJECT. Reading the tune half showed ~¾ of gold DEI positives were pure affirmative D&I program text, identical in kind to the false positives; the failure was predictable at Step 0. Bonus finding: the DEI-section-only edit left all five other domains within 0.025 F1 / 0.05 κ — single-section prompt edits isolate cleanly, so the perturbation check is a cheap add, not paranoia. Same pattern one week earlier: a cyber classifier pilot gate failure traced to E/D gold contamination (misses were skills-matrix-checkbox-only positives), not model weakness.

## Notes

- Low inter-rater κ on a label is the leading indicator: contested construct → gold-bound failures downstream. - If the pipeline scores all labels in one completion, any post-campaign prompt change forces a full re-score — run this check BEFORE the campaign. - See also: [llm-campaign-drift-gate] for the companion gate on resume boundaries and serving-revision drift (same fresh-baseline discipline). - See also: [annotator-input-parity-check] — run it FIRST. If the model was never shown the document the annotators read, apparent gold-bound failures (e.g. the 2026-07-16 E/D "contamination" reading above) are actually input mismatch: the 2026-07-21 parity audit showed the specialist-director hand labels were pure proxy-statement transcriptions, so checkbox-only positives were recoverable from the right input all along.

기술 세부 사항

버전
1.0.0
라이선스
MIT
최근 업데이트
2026년 8월 24일
게시일
2026년 8월 24일

결정 스냅샷

대체 후보

63
준비됨
프로토타입
단계

최근 저장소 활동

감사

설치 검토

설치 및 채택 검토

80
검토 필요
보안
86/100
유지보수
100/100
설치
92/100
전체 감사 열기평가 보고서 보기

Agent 검증 증거

Agent 검증 증거

Resolve, 검토, 설치 및 한 번의 제한된 실행 후 결과 보고서입니다.

0
검증됨
Needs first agent run자동 설치: 먼저 검토최근: 알 수 없음
성공률
최근 실패
결과
0
출력 품질
실패
0
관련 없음
0
설치
0
위험 차단
0
설정 필요
0
프로덕션
0

아직 Agent 결과 데이터가 없습니다. 첫 실행은 /api/agent/outcome을 통해 성공, 설정 필요, 위험 차단, 실패 또는 비관련 결과를 보고할 수 있습니다.

설치

Agent 워크플로에 추가

무료 오픈 소스. 프로덕션 Agent에 설치하기 전에 보고서를 검토하세요.

성장 루프

공유 키트

X

llm-gold-bound-failure-check용 시나리오 기반 초안입니다. X에 수동으로 게시할 수 있습니다.

큐레이터 노트
llm-gold-bound-failure-check: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on...

47 stars

https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x
X 초안 열기
선택 사항: 설치 명령이 포함된 답글
Listing + install path for llm-gold-bound-failure-check:
https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x

Install: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-...

등록 출처

Registry 색인

소유권 주장 가능

이 등록은 공개 소스에서 색인되었으며 유지보수자 소유권 주장이 승인될 때까지 공식으로 표시되지 않습니다.

제작자
Claude Code
색인 주체
OpenAgentSkill 커뮤니티 인덱스

귀속은 공개 저장소 또는 제작자 프로필에 연결됩니다. 제작자는 등록을 주장하여 소유권 신호를 업데이트할 수 있습니다.

이 스킬 소유권 주장

소유자 소유권 주장

이 스킬 등록 소유권 주장

이 Registry 색인 등록은 Claude Code에게 귀속되어 있지만 아직 공식으로 표시되지 않았습니다. 소유권을 주장하면 확인된 소유자 신호가 추가되어 이후 출시, 설치 및 감사 업데이트를 더 신뢰할 수 있습니다.

크리에이터 백링크 키트

README에 증거 배지 추가

개발자가 저장소를 평가하는 위치에 정규 등록, 현재 신뢰 및 감사 신호, 실제 Agent-Proven 증거를 표시합니다.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=listed&label=Listed)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=trust&label=Trust)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=audit&label=Audit)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)

작성자

C

Claude Code

@claude-code

플랫폼 적합도

상태 신호

GitHub 스타
47
품질 점수
35/100
최근 GitHub 푸시
2026년 8월 24일
프레임워크 힌트
알 수 없음
OpenAgentSkill 조회수
0
설치 명령 복사
0
외부 클릭
0

커뮤니티 신호

이 스킬이 Agent 워크플로에 유용한지 알려 주세요. 집계된 피드백은 시간이 지날수록 순위를 개선합니다.

신뢰와 안전

샌드박스 전용

71
  • GitHub 채택도GitHub 스타 47확인
  • 스타/포크 활동스타 47, 포크 0; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다확인
  • 최근 유지보수오늘 푸시됨통과
  • 라이선스 명확성MIT통과
  • README/SKILL.md 완성도메타데이터에 충분한 사용 및 워크플로 맥락이 포함되어 있습니다통과
  • 의존성/런타임 위험공개 메타데이터에 주요 의존성 위험 힌트가 없습니다통과