annotator-input-parity-check

REVIEW · 63
Registry indexed

Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whos

Verified installs0
Stars47
Version1.0.0
Quality64/100 · Promising
Trust63/100 · Sandbox only
Audit77/100 · Needs review

Supply asset profile

Research and knowledge work

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

Browse track

Scenario

Research agents

I need my agent to research a topic, compare sources, and produce a concise report.

Agent fit

Claude Code + CLI + Codex

Codex, Claude Code, Cursor, CLI, or custom agents.

Install

Ready

npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Maintenance

fresh

Pushed today

Risk

Needs review

No explicit 'Limitations' section, though the Notes section partially covers boundaries.

GitHub quality

47

64/100 Quality · 71/100 Trust

Coverage tags

ResearchResearch agentssecurityagent-skill

Review notes

No explicit 'Limitations' section, though the Notes section partially covers boundaries. · The skill description is long but well-structured; could be slightly more concise for quick scanning.

Agent adoption scorecard

Trust, audit, and install readiness at a glance

These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.

Quality

Promising
64

Useful candidate, but compare it with alternatives before adopting.

Trust

Sandbox only
63

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

Audit

Needs review
77

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

OpenAgentSkill Trust Score v5

Human review before install

Run only in a sandbox and compare close alternatives before using it for real work.

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

47 GitHub stars

Repo activity

47 stars, 0 forks

Maintenance

Pushed today

License

MIT

Install

npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Install safety

standard package or runtime install path

Permission surface

filesystem or document access

Agent outcomes

No agent outcome data yet

Docs

Usable metadata, review docs

Risk summary

Review before production

  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars

Install readiness

Install path available

  • Install path is available
  • Repository evidence is available
  • License is declared
  • No Agent Proven outcome evidence yet

Agent-readable metadata

Machine-readable decision data for this skill.

Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.

Open JSON

Suited tasks

  • Research agents workflows
  • Claude Code teams
  • builders willing to evaluate younger projects
  • Search sources

Suited agents

CodexClaude CodeCursorOpenAgentSkill CLICLI

Install decision

Command
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Policy
review
Human review
yes

Trust and risk

Trust
63/100
Audit
77/100
Risk level
Needs review

Outcome loop

Endpoint
/api/agent/outcome
Event ID
resolve
Outcomes
5

Install command

npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Do not use when

  • teams that need a vendor-supported SLA
  • production agents without a repository review
  • Low GitHub adoption signal
  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • No OpenAgentSkill engagement data yet

Agent safety v2

61/100 · Review before install

Reviewed with permission notesreview

Usable candidate, but the agent should surface permission and audit notes before installation.

Require human approval before installing into a real workspace.

Resolve via API

medium

Network access

Skill likely fetches remote pages, APIs, repositories, or external services.

medium

Filesystem access

Skill may read or write project files, documents, generated artifacts, or local workspace state.

  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.

Install targets

Install this skill in your agent workflow

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install kennethkhoocy-annotator-input-parity-check

Agent resolve plan

Let an agent verify fit before installing.

The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.

Open text plan

Agent should check

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Copy prompt

Task: Use annotator-input-parity-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent handoff

Give an agent the install path, not another directory page.

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

Open install API

Agent prompt

Use annotator-input-parity-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Registry metadata

Agent-readable profile for automatic skill selection.

This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.

Open manifest

Agent fit

63/100

Research agents

Platforms

Claude Code

Audit report

Needs review · 77/100

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

View audit reportView eval report

Agent decision cockpit

Fallback candidate for Research agents

Prototype with this skill first; keep a fallback candidate ready.

63
Readiness
Prototype
Stage

Role in stack

Fallback candidate

Primary fit

Research agents

Trust label

Prototype first

Install path

Command ready

Use when

  • Research agents workflows
  • Claude Code teams
  • builders willing to evaluate younger projects

Evidence

  • recent repository activity
  • install command or GitHub repo available
  • 64/100 quality profile

review first

  • Low GitHub adoption signal
  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • No OpenAgentSkill engagement data yet

Implementation path

  1. 1Install it in a sandbox agent and run one Research agents task end to end.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Trust profile

Sandbox only

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

63
OpenAgentSkill Trust Score

GitHub adoption

CHECK

47 GitHub stars

Stars/forks activity

CHECK

47 stars, 0 forks; issue activity unavailable in current metadata

Recent maintenance

PASS

Pushed today

License clarity

PASS

MIT

Good signals

  • AI review approved
  • Install path is available
  • Repository evidence is available
  • Recently maintained repository
  • Install command has no obvious high-risk pattern
  • Outcome loop is ready but needs first real agent run

Review before install

  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
  • No real agent outcome reports yet
  • Human review required before unattended installation

Recommended action

Run only in a sandbox and compare close alternatives before using it for real work.

Quality profile

Promising candidate for agent workflows

Useful candidate, but compare it with alternatives before adopting.

64
GitHub stars
47
Freshness
Today
Install ready
Yes
License
MIT
Review before install: Low GitHub adoption signal · No explicit 'Limitations' section, though the Notes section partially covers boundaries.

Workflow fit

Use this skill in these scenarios

Workflow fit

Add it to a complete workflow

Alternative shortlist

Compare before you install

Similar skills that may fit this task.

Compare all

Overview

--- name: annotator-input-parity-check description: | Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read. author: Claude Code version: 1.0.0 date: 2026-07-21 ---

# Annotator Input Parity Check

## Problem

A model built to replicate human labels is fed a different evidence base than the one the annotators used. The mismatch masquerades as a modeling or construct problem: recall collapses on the label subset whose evidence lives only in the annotators' source, audits produce increasingly sophisticated theory ("invisible" positives, construct splits, per-domain reliability gates), and successive model generations inherit the wrong input because each review critiques the lineage from inside the frozen input assumption.

## Context / Trigger Conditions

- Starting any label-replication build (classifier, LLM scorer, extractor) against hand-coded gold. - A validation report says some share of gold positives have "zero signal" in the model's input. - Proposals appear for: construct splits (what the model CAN see vs what the label encodes), human adjudication of "contested" cells, stop rules excluding weak domains, or accepting a permanent accuracy ceiling. - Verified instance (Specialist Directors US, 2026-07-21): three classifier generations (bio-BERT AUC 0.5 → structured RoBERTa "unclassifiable" on 3/5 domains → LLM dossier scorer with E/D construct split + PI adjudication + per-domain stop rules) all read director bios + BoardEx records, while the RA labels were pure transcriptions of PROXY-STATEMENT disclosures (skills matrices + bios, no exogenous data — confirmed in the source paper's methodology, 41 Yale J. Reg. 652, 669-72). The "invisible specialist" mass (43-79% of some domains) was simply the skills-matrix checkbox content the models were never shown. Years of downstream apparatus dissolved once the question "what did the labelers actually read?" was asked.

## Solution

1. Before any design work, write down the annotation protocol as the annotators executed it: source document(s), what they could see, what they could not, whether any exogenous data entered. Get this from the codebook/paper methodology section, not from folklore. If the protocol is unwritten, ask the PI directly: "did labelers consult anything beyond X?" 2. Compare against the model's planned input. Any evidence the annotators had that the model lacks is a hard recall ceiling on exactly the labels that evidence determines — no architecture, prompt, or training fixes it. 3. If a mismatch exists, prefer restoring input parity (give the model the annotators' document) over modeling around the gap. For transcription-style protocols, the task then becomes extraction, not prediction, and validation against the hand labels becomes construct-matched (agreement should be high; disagreement means extraction bugs, not construct philosophy). 4. Only if input parity is impossible (annotators used private knowledge, interviews, paywalled data) is a construct split the honest design — and then the model's output must be named as a DIFFERENT variable, never graded raw against the full gold. 5. When auditing an EXISTING lineage: ask the parity question first, before critiquing rubrics, thresholds, or gold quality. An audit that inherits the input assumption can be internally excellent and still miss the dominant error term.

## Verification

- The protocol-input inventory exists in writing and the model input is a superset of it → recall ceilings from "invisible" labels should disappear; residual disagreement decomposes into extraction errors (fixable) rather than unknowable-label mass. - Quick falsification test for a claimed "unpredictable" label subset: pull 5 such gold positives, open the annotators' source document for each, and check whether the label is visible there. If yes, the problem is input, not construct.

## Notes

- Distinct from [llm-gold-bound-failure-check], which diagnoses gold that fails to SEPARATE classes for a proposed revision; this skill diagnoses model INPUT that omits the annotators' evidence. Run this parity check first — gold-bound analysis of a parity-broken system wastes effort. - The mismatch is self-perpetuating across model generations: each successor inherits the predecessor's feature pipeline, and each audit optimizes within it. Breaking the frame requires asking about the ANNOTATORS, not the model. - Construct splits built on a parity-broken system may still have salvage value for a different question (e.g. record-evident-but-undisclosed expertise is analytically interesting in its own right) — reframe, don't necessarily discard.

Technical details

Version
1.0.0
License
MIT
Last updated
Aug 24, 2026
Published
Aug 24, 2026

Decision snapshot

Fallback candidate

63
Ready
Prototype
Stage

recent repository activity

Audit

Install review

Install and adoption review

77
Needs review
Security
80/100
Maintenance
100/100
Install
92/100
Open full auditView eval report

Agent-proven evidence

Agent-proven evidence

Outcome reports after resolve, review, install, and one narrow run.

0
Proven
Needs first agent runAuto-install: review firstLast: Unknown
Success rate
Recent failure
Outcomes
0
Output quality
Failed
0
Not relevant
0
Installs
0
Risk blocked
0
Setup needed
0
Production
0

No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.

Install

Add to agent workflow

Free and open source. Review the report before installing into production agents.

Growth loop

Share kit

X

Scenario-led draft for annotator-input-parity-check, ready for a manual X post.

Curator note
annotator-input-parity-check: Before designing, training, or auditing ANY model that replicates human-annotated labels, aud...

47 stars

https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x
Open X draft
Optional reply with install command
Listing + install path for annotator-input-parity-check:
https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x

Install: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-...

Listing source

Registry indexed

Claimable

This listing was indexed from public sources and is not marked official until a maintainer claim is approved.

Indexed by
OpenAgentSkill community index

Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.

Claim this skill

Owner claim

Claim this skill listing

This Registry indexed listing is attributed to Claude Code but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.

Creator backlink kit

Add the evidence badges to your README

Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=listed&label=Listed)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=trust&label=Trust)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=audit&label=Audit)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)

Author

C

Claude Code

@claude-code

Platform fit

Health signals

GitHub stars
47
Quality score
35/100
Last GitHub push
Aug 24, 2026
Framework hints
Unknown
OpenAgentSkill views
0
Install copies
0
Outbound clicks
0

Community signal

Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.

Trust & safety

Sandbox only

63
  • GitHub adoption47 GitHub starsCHECK
  • Stars/forks activity47 stars, 0 forks; issue activity unavailable in current metadataCHECK
  • Recent maintenancePushed todayPASS
  • License clarityMITPASS
  • README/SKILL.md completenessPublic metadata needs stronger README/SKILL.md contextINFO
  • Dependency/runtime riskno major dependency risk hints in public metadataPASS