annotator-input-parity-check
Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whos
Supply asset profile
Research and knowledge work
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Maintenance
fresh
Pushed today
Risk
Needs review
No explicit 'Limitations' section, though the Notes section partially covers boundaries.
GitHub quality
47
64/100 Quality · 71/100 Trust
Coverage tags
Review notes
No explicit 'Limitations' section, though the Notes section partially covers boundaries. · The skill description is long but well-structured; could be slightly more concise for quick scanning.
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
47 GitHub stars
Repo activity
47 stars, 0 forks
Maintenance
Pushed today
License
MIT
Install
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Install safety
standard package or runtime install path
Permission surface
filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Review before production
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- Research agents workflows
- Claude Code teams
- builders willing to evaluate younger projects
- Search sources
Suited agents
Install decision
- Command
- npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
- Policy
- review
- Human review
- yes
Trust and risk
- Trust
- 63/100
- Audit
- 77/100
- Risk level
- Needs review
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-checkDo not use when
- teams that need a vendor-supported SLA
- production agents without a repository review
- Low GitHub adoption signal
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- No OpenAgentSkill engagement data yet
Agent safety v2
61/100 · Review before install
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Filesystem access
Skill may read or write project files, documents, generated artifacts, or local workspace state.
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install kennethkhoocy-annotator-input-parity-checkAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/kennethkhoocy-annotator-input-parity-check/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use annotator-input-parity-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/kennethkhoocy-annotator-input-parity-check/install
LLM text format
/api/skills/kennethkhoocy-annotator-input-parity-check/install?format=text
Find alternatives
/api/skills/search?q=annotator-input-parity-check&limit=3
Agent prompt
Use annotator-input-parity-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-checkRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/kennethkhoocy-annotator-input-parity-check
LLM text
/api/registry/manifest/kennethkhoocy-annotator-input-parity-check?format=text
Install alias
/api/registry/install/kennethkhoocy-annotator-input-parity-check
Recommend
/api/registry/recommend?task=Use%20annotator-input-parity-check%20in%20an%20agent%20workflow&limit=3
Agent fit
Research agents
Use-case tags
Platforms
Claude Code
Audit report
Needs review · 77/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Fallback candidate for Research agents
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
Research agents
Trust label
Prototype first
Install path
Command ready
Use when
- Research agents workflows
- Claude Code teams
- builders willing to evaluate younger projects
Evidence
- recent repository activity
- install command or GitHub repo available
- 64/100 quality profile
review first
- Low GitHub adoption signal
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- No OpenAgentSkill engagement data yet
Implementation path
- 1Install it in a sandbox agent and run one Research agents task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Sandbox only
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
CHECK47 GitHub stars
Stars/forks activity
CHECK47 stars, 0 forks; issue activity unavailable in current metadata
Recent maintenance
PASSPushed today
License clarity
PASSMIT
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Promising candidate for agent workflows
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Use this skill in these scenarios
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Add it to a complete workflow
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
Maigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
Overview
--- name: annotator-input-parity-check description: | Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read. author: Claude Code version: 1.0.0 date: 2026-07-21 ---
# Annotator Input Parity Check
## Problem
A model built to replicate human labels is fed a different evidence base than the one the annotators used. The mismatch masquerades as a modeling or construct problem: recall collapses on the label subset whose evidence lives only in the annotators' source, audits produce increasingly sophisticated theory ("invisible" positives, construct splits, per-domain reliability gates), and successive model generations inherit the wrong input because each review critiques the lineage from inside the frozen input assumption.
## Context / Trigger Conditions
- Starting any label-replication build (classifier, LLM scorer, extractor) against hand-coded gold. - A validation report says some share of gold positives have "zero signal" in the model's input. - Proposals appear for: construct splits (what the model CAN see vs what the label encodes), human adjudication of "contested" cells, stop rules excluding weak domains, or accepting a permanent accuracy ceiling. - Verified instance (Specialist Directors US, 2026-07-21): three classifier generations (bio-BERT AUC 0.5 → structured RoBERTa "unclassifiable" on 3/5 domains → LLM dossier scorer with E/D construct split + PI adjudication + per-domain stop rules) all read director bios + BoardEx records, while the RA labels were pure transcriptions of PROXY-STATEMENT disclosures (skills matrices + bios, no exogenous data — confirmed in the source paper's methodology, 41 Yale J. Reg. 652, 669-72). The "invisible specialist" mass (43-79% of some domains) was simply the skills-matrix checkbox content the models were never shown. Years of downstream apparatus dissolved once the question "what did the labelers actually read?" was asked.
## Solution
1. Before any design work, write down the annotation protocol as the annotators executed it: source document(s), what they could see, what they could not, whether any exogenous data entered. Get this from the codebook/paper methodology section, not from folklore. If the protocol is unwritten, ask the PI directly: "did labelers consult anything beyond X?" 2. Compare against the model's planned input. Any evidence the annotators had that the model lacks is a hard recall ceiling on exactly the labels that evidence determines — no architecture, prompt, or training fixes it. 3. If a mismatch exists, prefer restoring input parity (give the model the annotators' document) over modeling around the gap. For transcription-style protocols, the task then becomes extraction, not prediction, and validation against the hand labels becomes construct-matched (agreement should be high; disagreement means extraction bugs, not construct philosophy). 4. Only if input parity is impossible (annotators used private knowledge, interviews, paywalled data) is a construct split the honest design — and then the model's output must be named as a DIFFERENT variable, never graded raw against the full gold. 5. When auditing an EXISTING lineage: ask the parity question first, before critiquing rubrics, thresholds, or gold quality. An audit that inherits the input assumption can be internally excellent and still miss the dominant error term.
## Verification
- The protocol-input inventory exists in writing and the model input is a superset of it → recall ceilings from "invisible" labels should disappear; residual disagreement decomposes into extraction errors (fixable) rather than unknowable-label mass. - Quick falsification test for a claimed "unpredictable" label subset: pull 5 such gold positives, open the annotators' source document for each, and check whether the label is visible there. If yes, the problem is input, not construct.
## Notes
- Distinct from [llm-gold-bound-failure-check], which diagnoses gold that fails to SEPARATE classes for a proposed revision; this skill diagnoses model INPUT that omits the annotators' evidence. Run this parity check first — gold-bound analysis of a parity-broken system wastes effort. - The mismatch is self-perpetuating across model generations: each successor inherits the predecessor's feature pipeline, and each audit optimizes within it. Breaking the frame requires asking about the ANNOTATORS, not the model. - Construct splits built on a parity-broken system may still have salvage value for a different question (e.g. record-evident-but-undisclosed expertise is analytically interesting in its own right) — reframe, don't necessarily discard.
Technical details
- Version
- 1.0.0
- License
- MIT
- Last updated
- Aug 24, 2026
- Published
- Aug 24, 2026
Decision snapshot
Fallback candidate
recent repository activity
Audit
Install review
Install and adoption review
- Security
- 80/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for annotator-input-parity-check, ready for a manual X post.
annotator-input-parity-check: Before designing, training, or auditing ANY model that replicates human-annotated labels, aud... 47 stars https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x
Optional reply with install command
Listing + install path for annotator-input-parity-check: https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x Install: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-...
Listing source
Registry indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- Claude Code
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Registry indexed listing is attributed to Claude Code but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check/audit)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)Author
Claude Code
@claude-code
Tags
Platform fit
Health signals
- GitHub stars
- 47
- Quality score
- 35/100
- Last GitHub push
- Aug 24, 2026
- Framework hints
- Unknown
- OpenAgentSkill views
- 0
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Sandbox only
- GitHub adoption47 GitHub starsCHECK
- Stars/forks activity47 stars, 0 forks; issue activity unavailable in current metadataCHECK
- Recent maintenancePushed todayPASS
- License clarityMITPASS
- README/SKILL.md completenessPublic metadata needs stronger README/SKILL.md contextINFO
- Dependency/runtime riskno major dependency risk hints in public metadataPASS
Related skills
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K Stars