제작자 · Claude Code
최근 업데이트 · 2026년 8월 24일
annotator-input-parity-check
Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whos
샌드박스 전용
설치 대상
Codex 설치 프롬프트
Install the "annotator-input-parity-check" agent skill from https://github.com/kennethkhoocy/applied-micro-skills/tree/main/plugins/applied-micro/skills/annotator-input-parity-check. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"kennethkhoocy-annotator-input-parity-check","task":"Install annotator-input-parity-check","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.공급 자산 프로필
리서치 및 지식 작업
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
시나리오
리서치 Agent
I need my agent to research a topic, compare sources, and produce a concise report.
Agent 적합도
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI 또는 맞춤형 Agent에 적합합니다.
설치
준비됨
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
유지보수
최신
오늘 푸시됨
위험
검토 필요
No explicit 'Limitations' section, though the Notes section partially covers boundaries.
GitHub 품질
47
64/100 품질 · 71/100 신뢰
커버리지 태그
검토 메모
No explicit 'Limitations' section, though the Notes section partially covers boundaries. · The skill description is long but well-structured; could be slightly more concise for quick scanning.
Agent 채택 스코어카드
신뢰, 감사, 설치 준비 상태를 한눈에 확인하세요
이 점수는 공개 저장소 메타데이터, OpenAgentSkill 검토 신호, 유지보수 최신성, 설치 준비 상태를 결합합니다. 후보 선정 신호일 뿐, 사람의 검토를 대체하지 않습니다.
품질
유망유용한 후보이지만 채택 전에 대안과 비교하세요.
신뢰
샌드박스 전용신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.
감사
검토 필요설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.
OpenAgentSkill 신뢰 점수 v5
설치 전 사람 검토
실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.
스타
GitHub 스타 47
저장소 활동
스타 47, 포크 0
유지보수
오늘 푸시됨
라이선스
MIT
설치
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
설치 안전성
표준 패키지 또는 런타임 설치 경로
권한 범위
파일 시스템 또는 문서 접근
Agent 결과
아직 Agent 결과 데이터가 없습니다
문서
Usable metadata, review docs
위험 요약
프로덕션 전 검토
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
설치 준비 상태
설치 경로 사용 가능
- 설치 경로를 사용할 수 있습니다
- 저장소 근거를 사용할 수 있습니다
- 라이선스가 명시되었습니다
- 아직 Agent 검증 결과 근거가 없습니다
Agent 읽기용 메타데이터
이 스킬의 기계 판독형 의사결정 데이터.
이 블록 또는 포함된 JSON을 사용해 Agent가 이 스킬을 설치할지, 대안을 고를지, 먼저 사람의 검토를 요청할지 판단할 수 있습니다.
View technical data+
Agent 읽기용 메타데이터
이 스킬의 기계 판독형 의사결정 데이터.
이 블록 또는 포함된 JSON을 사용해 Agent가 이 스킬을 설치할지, 대안을 고를지, 먼저 사람의 검토를 요청할지 판단할 수 있습니다.
적합한 작업
- 리서치 Agent 워크플로
- Claude Code 팀
- builders willing to evaluate younger projects
- 검색 소스
적합한 Agent
설치 결정
- 명령어
- npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
- 정책
- 검토
- 사람 검토
- 예
신뢰와 위험
- 신뢰
- 63/100
- 감사
- 77/100
- 위험 수준
- 검토 필요
결과 루프
- 엔드포인트
- /api/agent/outcome
- 이벤트 ID
- resolve
- 결과
- 5
설치 명령어
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check사용하지 말아야 할 경우
- 벤더 지원 SLA가 필요한 팀
- production agents without a repository review
- Low GitHub adoption signal
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다
Agent 안전 v2
61/100 · 설치 전 검토
사용 가능한 후보이지만 Agent는 설치 전에 권한과 감사 메모를 표시해야 합니다.
실제 작업 공간에 설치하기 전에 사람의 승인이 필요합니다.
중간
네트워크 접근
Skill은 원격 페이지, API, 저장소 또는 외부 서비스에 접근할 수 있습니다.
중간
파일 시스템 접근
Skill은 프로젝트 파일, 문서, 생성 산출물 또는 로컬 작업 공간 상태를 읽거나 쓸 수 있습니다.
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
Agent 해결 계획
설치 전에 Agent가 적합성을 검증하게 하세요.
Resolve API는 최우선 스킬, 대안, 안전 정책, 감사 메모, 설치 대상 및 Agent가 페이지를 스크래핑하지 않고 사용할 수 있는 프롬프트를 반환합니다.
JSON 열기
/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve 텍스트
/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
설치 핸드오프
/api/skills/kennethkhoocy-annotator-input-parity-check/install
Agent가 확인할 항목
- Resolve API에서 작업 적합도와 대안을 확인합니다.
- 감사 점수, 신뢰 점수 및 안전 정책 경고를 확인합니다.
- Codex, Claude Code, Cursor 또는 CLI의 설치 대상 호환성을 확인합니다.
프롬프트 복사
Task: Use annotator-input-parity-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent 핸드오프
또 다른 디렉터리 페이지 대신 설치 경로를 Agent에게 제공합니다.
공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.
설치 핸드오프
/api/skills/kennethkhoocy-annotator-input-parity-check/install
LLM 텍스트 형식
/api/skills/kennethkhoocy-annotator-input-parity-check/install?format=text
대안 찾기
/api/skills/search?q=annotator-input-parity-check&limit=3
Agent 프롬프트
Use annotator-input-parity-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-checkRegistry 메타데이터
자동 스킬 선택을 위한 Agent 읽기용 프로필.
Registry API를 통해 동일한 결정, 신뢰, 감사, 사용 사례, 설치 신호를 제공하므로 Agent가 UI를 스크래핑하지 않고도 순위를 매길 수 있습니다.
Manifest
/api/registry/manifest/kennethkhoocy-annotator-input-parity-check
LLM 텍스트
/api/registry/manifest/kennethkhoocy-annotator-input-parity-check?format=text
설치 별칭
/api/registry/install/kennethkhoocy-annotator-input-parity-check
추천
/api/registry/recommend?task=Use%20annotator-input-parity-check%20in%20an%20agent%20workflow&limit=3
Agent 결정 패널
Fallback candidate for Research agents
먼저 이 스킬로 프로토타입을 만들고 대체 후보를 준비하세요.
스택 내 역할
대체 후보
주요 적합도
리서치 Agent
신뢰 라벨
먼저 프로토타입
설치 경로
명령어 준비됨
사용 시점
- 리서치 Agent 워크플로
- Claude Code 팀
- builders willing to evaluate younger projects
근거
- 최근 저장소 활동
- 설치 명령 또는 GitHub 저장소를 사용할 수 있습니다
- 품질 프로필 64/100
먼저 검토
- Low GitHub adoption signal
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다
구현 경로
- 1샌드박스 Agent에 설치하고 리서치 Agent 작업을 처음부터 끝까지 한 번 실행하세요.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
신뢰 프로필
샌드박스 전용
신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.
GitHub 채택도
확인GitHub 스타 47
스타/포크 활동
확인스타 47, 포크 0; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다
최근 유지보수
통과오늘 푸시됨
라이선스 명확성
통과MIT
긍정 신호
- AI 검토 승인됨
- 설치 경로를 사용할 수 있습니다
- 저장소 근거를 사용할 수 있습니다
- 최근 유지보수된 저장소
- 설치 명령에서 뚜렷한 고위험 패턴이 발견되지 않았습니다
- 결과 루프는 준비되었지만 첫 실제 Agent 실행이 필요합니다
설치 전 검토
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
- 아직 실제 Agent 결과 보고서가 없습니다
- 무인 설치 전에 사람 검토가 필요합니다
권장 작업
실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.
품질 프로필
유망 Agent 워크플로용 후보
유용한 후보이지만 채택 전에 대안과 비교하세요.
워크플로 적합도
이 스킬을 사용할 시나리오
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
워크플로 적합도
완전한 워크플로에 추가
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
대안 후보
설치 전 비교
이 작업에 적합할 수 있는 유사 스킬입니다.
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
Maigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
개요
--- name: annotator-input-parity-check description: | Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read. author: Claude Code version: 1.0.0 date: 2026-07-21 ---
# Annotator Input Parity Check
## Problem
A model built to replicate human labels is fed a different evidence base than the one the annotators used. The mismatch masquerades as a modeling or construct problem: recall collapses on the label subset whose evidence lives only in the annotators' source, audits produce increasingly sophisticated theory ("invisible" positives, construct splits, per-domain reliability gates), and successive model generations inherit the wrong input because each review critiques the lineage from inside the frozen input assumption.
## Context / Trigger Conditions
- Starting any label-replication build (classifier, LLM scorer, extractor) against hand-coded gold. - A validation report says some share of gold positives have "zero signal" in the model's input. - Proposals appear for: construct splits (what the model CAN see vs what the label encodes), human adjudication of "contested" cells, stop rules excluding weak domains, or accepting a permanent accuracy ceiling. - Verified instance (Specialist Directors US, 2026-07-21): three classifier generations (bio-BERT AUC 0.5 → structured RoBERTa "unclassifiable" on 3/5 domains → LLM dossier scorer with E/D construct split + PI adjudication + per-domain stop rules) all read director bios + BoardEx records, while the RA labels were pure transcriptions of PROXY-STATEMENT disclosures (skills matrices + bios, no exogenous data — confirmed in the source paper's methodology, 41 Yale J. Reg. 652, 669-72). The "invisible specialist" mass (43-79% of some domains) was simply the skills-matrix checkbox content the models were never shown. Years of downstream apparatus dissolved once the question "what did the labelers actually read?" was asked.
## Solution
1. Before any design work, write down the annotation protocol as the annotators executed it: source document(s), what they could see, what they could not, whether any exogenous data entered. Get this from the codebook/paper methodology section, not from folklore. If the protocol is unwritten, ask the PI directly: "did labelers consult anything beyond X?" 2. Compare against the model's planned input. Any evidence the annotators had that the model lacks is a hard recall ceiling on exactly the labels that evidence determines — no architecture, prompt, or training fixes it. 3. If a mismatch exists, prefer restoring input parity (give the model the annotators' document) over modeling around the gap. For transcription-style protocols, the task then becomes extraction, not prediction, and validation against the hand labels becomes construct-matched (agreement should be high; disagreement means extraction bugs, not construct philosophy). 4. Only if input parity is impossible (annotators used private knowledge, interviews, paywalled data) is a construct split the honest design — and then the model's output must be named as a DIFFERENT variable, never graded raw against the full gold. 5. When auditing an EXISTING lineage: ask the parity question first, before critiquing rubrics, thresholds, or gold quality. An audit that inherits the input assumption can be internally excellent and still miss the dominant error term.
## Verification
- The protocol-input inventory exists in writing and the model input is a superset of it → recall ceilings from "invisible" labels should disappear; residual disagreement decomposes into extraction errors (fixable) rather than unknowable-label mass. - Quick falsification test for a claimed "unpredictable" label subset: pull 5 such gold positives, open the annotators' source document for each, and check whether the label is visible there. If yes, the problem is input, not construct.
## Notes
- Distinct from [llm-gold-bound-failure-check], which diagnoses gold that fails to SEPARATE classes for a proposed revision; this skill diagnoses model INPUT that omits the annotators' evidence. Run this parity check first — gold-bound analysis of a parity-broken system wastes effort. - The mismatch is self-perpetuating across model generations: each successor inherits the predecessor's feature pipeline, and each audit optimizes within it. Breaking the frame requires asking about the ANNOTATORS, not the model. - Construct splits built on a parity-broken system may still have salvage value for a different question (e.g. record-evident-but-undisclosed expertise is analytically interesting in its own right) — reframe, don't necessarily discard.
기술 세부 사항
- 버전
- 1.0.0
- 라이선스
- MIT
- 최근 업데이트
- 2026년 8월 24일
- 게시일
- 2026년 8월 24일
결정 스냅샷
대체 후보
최근 저장소 활동
Agent 검증 증거
Agent 검증 증거
Resolve, 검토, 설치 및 한 번의 제한된 실행 후 결과 보고서입니다.
- 성공률
- —
- 최근 실패
- —
- 결과
- 0
- 출력 품질
- —
- 실패
- 0
- 관련 없음
- 0
- 설치
- 0
- 위험 차단
- 0
- 설정 필요
- 0
- 프로덕션
- 0
아직 Agent 결과 데이터가 없습니다. 첫 실행은 /api/agent/outcome을 통해 성공, 설정 필요, 위험 차단, 실패 또는 비관련 결과를 보고할 수 있습니다.
성장 루프
공유 키트
annotator-input-parity-check용 시나리오 기반 초안입니다. X에 수동으로 게시할 수 있습니다.
annotator-input-parity-check: Before designing, training, or auditing ANY model that replicates human-annotated labels, aud... 47 stars https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x
선택 사항: 설치 명령이 포함된 답글
Listing + install path for annotator-input-parity-check: https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x Install: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-...
등록 출처
Registry 색인
이 등록은 공개 소스에서 색인되었으며 유지보수자 소유권 주장이 승인될 때까지 공식으로 표시되지 않습니다.
- 제작자
- Claude Code
- 색인 주체
- OpenAgentSkill 커뮤니티 인덱스
귀속은 공개 저장소 또는 제작자 프로필에 연결됩니다. 제작자는 등록을 주장하여 소유권 신호를 업데이트할 수 있습니다.
이 스킬 소유권 주장소유자 소유권 주장
이 스킬 등록 소유권 주장
이 Registry 색인 등록은 Claude Code에게 귀속되어 있지만 아직 공식으로 표시되지 않았습니다. 소유권을 주장하면 확인된 소유자 신호가 추가되어 이후 출시, 설치 및 감사 업데이트를 더 신뢰할 수 있습니다.
크리에이터 백링크 키트
README에 증거 배지 추가
개발자가 저장소를 평가하는 위치에 정규 등록, 현재 신뢰 및 감사 신호, 실제 Agent-Proven 증거를 표시합니다.
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check/audit)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)작성자
Claude Code
@claude-code
플랫폼 적합도
상태 신호
- GitHub 스타
- 47
- 품질 점수
- 35/100
- 최근 GitHub 푸시
- 2026년 8월 24일
- 프레임워크 힌트
- 알 수 없음
- OpenAgentSkill 조회수
- 0
- 설치 명령 복사
- 0
- 외부 클릭
- 0
커뮤니티 신호
이 스킬이 Agent 워크플로에 유용한지 알려 주세요. 집계된 피드백은 시간이 지날수록 순위를 개선합니다.
신뢰와 안전
샌드박스 전용
- GitHub 채택도GitHub 스타 47확인
- 스타/포크 활동스타 47, 포크 0; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다확인
- 최근 유지보수오늘 푸시됨통과
- 라이선스 명확성MIT통과
- README/SKILL.md 완성도공개 메타데이터에 더 충실한 README/SKILL.md 맥락이 필요합니다정보
- 의존성/런타임 위험공개 메타데이터에 주요 의존성 위험 힌트가 없습니다통과
관련 스킬
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K 스타Maigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K 스타Nuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K 스타Infisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K 스타