annotator-input-parity-check

Prüfen · 63
Im Registry indexiert

Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whos

Verified installs0
Stars47
Version1.0.0
Qualität64/100 · Vielversprechend
Vertrauen63/100 · Nur Sandbox
Audit77/100 · Prüfung nötig

Asset-Profil

Recherche und Wissensarbeit

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

Bereich ansehen

Szenario

Recherche-Agents

I need my agent to research a topic, compare sources, and produce a concise report.

Agent-Fit

Claude Code + CLI + Codex

Geeignet für Codex, Claude Code, Cursor, CLI oder benutzerdefinierte Agents.

Installieren

Bereit

npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Wartung

Aktuell

Heute gepusht

Risiko

Prüfung nötig

No explicit 'Limitations' section, though the Notes section partially covers boundaries.

GitHub-Qualität

47

64/100 Qualität · 71/100 Vertrauen

Abdeckungs-Tags

RechercheRecherche-AgentsSicherheitagent-skill

Review-Notizen

No explicit 'Limitations' section, though the Notes section partially covers boundaries. · The skill description is long but well-structured; could be slightly more concise for quick scanning.

Agent-Adoptionskarte

Vertrauen, Audit und Installationsbereitschaft auf einen Blick

Diese Werte kombinieren öffentliche Repository-Metadaten, OpenAgentSkill-Reviewsignale, Wartungsaktualität und Installationsbereitschaft. Sie helfen bei der Vorauswahl, ersetzen aber keine menschliche Prüfung.

Qualität

Vielversprechend
64

Useful candidate, but compare it with alternatives before adopting.

Vertrauen

Nur Sandbox
63

Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.

Audit

Prüfung nötig
77

Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.

OpenAgentSkill Trust Score v5

Menschliche Prüfung vor Installation

Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

47 GitHub-Stars

Repository-Aktivität

47 Stars und 0 Forks

Wartung

Heute gepusht

Lizenz

MIT

Installieren

npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Installationssicherheit

Standard-Paket- oder Laufzeit-Installationspfad

Berechtigungsfläche

Dateisystem- oder Dokumentzugriff

Agent-Ergebnisse

Noch keine Agent-Ergebnisdaten

Dokumentation

Usable metadata, review docs

Risikoübersicht

Vor Produktion prüfen

  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars

Installationsbereitschaft

Installationspfad verfügbar

  • Installationspfad ist verfügbar
  • Repository-Belege sind verfügbar
  • Lizenz ist angegeben
  • Noch keine Agent-Proven-Ergebnisbelege

Agent-lesbare Metadaten

Maschinenlesbare Entscheidungsdaten für diesen Skill.

Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.

JSON öffnen

Geeignete Aufgaben

  • Research-Agent-Workflows
  • Claude-Code-Teams
  • builders willing to evaluate younger projects
  • Suchquellen

Geeignete Agents

CodexClaude CodeCursorOpenAgentSkill CLICLI

Installationsentscheidung

Befehl
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Richtlinie
Prüfen
Menschliche Prüfung
Ja

Vertrauen und Risiko

Vertrauen
63/100
Audit
77/100
Risikoebene
Prüfung nötig

Ergebnis-Loop

Endpoint
/api/agent/outcome
Event-ID
resolve
Ergebnisse
5

Installationsbefehl

npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Nicht verwenden, wenn

  • Teams, die ein vom Anbieter unterstütztes SLA benötigen
  • production agents without a repository review
  • Low GitHub adoption signal
  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • No OpenAgentSkill engagement data yet

Agent-Sicherheit v2

61/100 · Vor Installation prüfen

Mit Berechtigungshinweisen geprüftPrüfen

Nutzbarer Kandidat, aber der Agent sollte Berechtigungs- und Auditnotizen vor der Installation anzeigen.

Vor der Installation in einem echten Arbeitsbereich ist menschliche Freigabe erforderlich.

Per API auflösen

Mittel

Netzwerkzugriff

Die Skill ruft wahrscheinlich Remote-Seiten, APIs, Repositories oder externe Dienste ab.

Mittel

Dateisystemzugriff

Die Skill kann Projektdateien, Dokumente, generierte Artefakte oder den lokalen Arbeitsbereich lesen oder schreiben.

  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.

Installationsziele

Diesen Skill im Agent-Workflow installieren

Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install kennethkhoocy-annotator-input-parity-check

Agent-Auflösungsplan

Lass einen Agent die Eignung vor der Installation prüfen.

Die Resolve API liefert die beste Skill, Alternativen, Sicherheitsrichtlinien, Auditnotizen, Installationsziel und einen direkt nutzbaren Prompt.

Textplan öffnen

Agent sollte prüfen

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Prompt kopieren

Task: Use annotator-input-parity-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent-Übergabe

Gib dem Agent den Installationspfad, nicht noch ein Verzeichnis.

Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.

Installations-API öffnen

Agent-Prompt

Use annotator-input-parity-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check

Registry-Metadaten

Agent-lesbares Profil für die automatische Skill-Auswahl.

Die Registry API stellt Entscheidungs-, Vertrauens-, Audit-, Use-Case- und Installationssignale ohne UI-Scraping bereit.

Manifest öffnen

Agent-Fit

63/100

Recherche-Agents

Plattformen

Claude Code

Audit-Bericht

Prüfung nötig · 77/100

Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.

Audit-Bericht ansehenEval-Bericht ansehen

Agent-Entscheidungspanel

Fallback candidate for Research agents

Prototype with this skill first; keep a fallback candidate ready.

63
Bereitschaft
Prototyp
Phase

Rolle im Stack

Fallback-Kandidat

Primäre Eignung

Recherche-Agents

Vertrauenslabel

Zuerst prototypisieren

Installationspfad

Befehl bereit

Verwenden wenn

  • Research-Agent-Workflows
  • Claude-Code-Teams
  • builders willing to evaluate younger projects

Evidenz

  • recent repository activity
  • install command or GitHub repo available
  • Qualitätsprofil 64/100

zuerst prüfen

  • Low GitHub adoption signal
  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • No OpenAgentSkill engagement data yet

Implementierungspfad

  1. 1Installieren Sie es in einem Sandbox-Agent und führen Sie eine Recherche-Agents-Aufgabe vollständig aus.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Vertrauensprofil

Nur Sandbox

Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.

63
OpenAgentSkill Trust Score

GitHub-Akzeptanz

Prüfen

47 GitHub-Stars

Star-/Fork-Aktivität

Prüfen

47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbar

Aktuelle Wartung

Bestanden

Heute gepusht

Lizenzklarheit

Bestanden

MIT

Positive Signale

  • KI-Prüfung genehmigt
  • Installationspfad ist verfügbar
  • Repository-Belege sind verfügbar
  • Kürzlich gewartetes Repository
  • Der Installationsbefehl weist kein offensichtliches Hochrisikomuster auf
  • Ergebniszyklus ist bereit, benötigt aber den ersten echten Agent-Lauf

Vor Installation prüfen

  • No explicit 'Limitations' section, though the Notes section partially covers boundaries.
  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
  • Noch keine echten Agent-Ergebnisberichte
  • Vor unbeaufsichtigter Installation ist menschliche Prüfung erforderlich

Empfohlene Aktion

Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.

Qualitätsprofil

Vielversprechend Kandidat für Agent-Workflows

Useful candidate, but compare it with alternatives before adopting.

64
GitHub-Stars
47
Aktualität
Heute
Installationsbereit
Ja
Lizenz
MIT
Vor Installation prüfen: Low GitHub adoption signal · No explicit 'Limitations' section, though the Notes section partially covers boundaries.

Workflow-Eignung

Diese Skill in diesen Szenarien nutzen

Workflow-Eignung

Zum vollständigen Workflow hinzufügen

Alternativen-Shortlist

Vor Installation vergleichen

Similar skills that may fit this task.

Alle vergleichen

Übersicht

--- name: annotator-input-parity-check description: | Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read. author: Claude Code version: 1.0.0 date: 2026-07-21 ---

# Annotator Input Parity Check

## Problem

A model built to replicate human labels is fed a different evidence base than the one the annotators used. The mismatch masquerades as a modeling or construct problem: recall collapses on the label subset whose evidence lives only in the annotators' source, audits produce increasingly sophisticated theory ("invisible" positives, construct splits, per-domain reliability gates), and successive model generations inherit the wrong input because each review critiques the lineage from inside the frozen input assumption.

## Context / Trigger Conditions

- Starting any label-replication build (classifier, LLM scorer, extractor) against hand-coded gold. - A validation report says some share of gold positives have "zero signal" in the model's input. - Proposals appear for: construct splits (what the model CAN see vs what the label encodes), human adjudication of "contested" cells, stop rules excluding weak domains, or accepting a permanent accuracy ceiling. - Verified instance (Specialist Directors US, 2026-07-21): three classifier generations (bio-BERT AUC 0.5 → structured RoBERTa "unclassifiable" on 3/5 domains → LLM dossier scorer with E/D construct split + PI adjudication + per-domain stop rules) all read director bios + BoardEx records, while the RA labels were pure transcriptions of PROXY-STATEMENT disclosures (skills matrices + bios, no exogenous data — confirmed in the source paper's methodology, 41 Yale J. Reg. 652, 669-72). The "invisible specialist" mass (43-79% of some domains) was simply the skills-matrix checkbox content the models were never shown. Years of downstream apparatus dissolved once the question "what did the labelers actually read?" was asked.

## Solution

1. Before any design work, write down the annotation protocol as the annotators executed it: source document(s), what they could see, what they could not, whether any exogenous data entered. Get this from the codebook/paper methodology section, not from folklore. If the protocol is unwritten, ask the PI directly: "did labelers consult anything beyond X?" 2. Compare against the model's planned input. Any evidence the annotators had that the model lacks is a hard recall ceiling on exactly the labels that evidence determines — no architecture, prompt, or training fixes it. 3. If a mismatch exists, prefer restoring input parity (give the model the annotators' document) over modeling around the gap. For transcription-style protocols, the task then becomes extraction, not prediction, and validation against the hand labels becomes construct-matched (agreement should be high; disagreement means extraction bugs, not construct philosophy). 4. Only if input parity is impossible (annotators used private knowledge, interviews, paywalled data) is a construct split the honest design — and then the model's output must be named as a DIFFERENT variable, never graded raw against the full gold. 5. When auditing an EXISTING lineage: ask the parity question first, before critiquing rubrics, thresholds, or gold quality. An audit that inherits the input assumption can be internally excellent and still miss the dominant error term.

## Verification

- The protocol-input inventory exists in writing and the model input is a superset of it → recall ceilings from "invisible" labels should disappear; residual disagreement decomposes into extraction errors (fixable) rather than unknowable-label mass. - Quick falsification test for a claimed "unpredictable" label subset: pull 5 such gold positives, open the annotators' source document for each, and check whether the label is visible there. If yes, the problem is input, not construct.

## Notes

- Distinct from [llm-gold-bound-failure-check], which diagnoses gold that fails to SEPARATE classes for a proposed revision; this skill diagnoses model INPUT that omits the annotators' evidence. Run this parity check first — gold-bound analysis of a parity-broken system wastes effort. - The mismatch is self-perpetuating across model generations: each successor inherits the predecessor's feature pipeline, and each audit optimizes within it. Breaking the frame requires asking about the ANNOTATORS, not the model. - Construct splits built on a parity-broken system may still have salvage value for a different question (e.g. record-evident-but-undisclosed expertise is analytically interesting in its own right) — reframe, don't necessarily discard.

Technische Details

Version
1.0.0
Lizenz
MIT
Letzte Aktualisierung
24. Aug. 2026
Veröffentlicht
24. Aug. 2026

Entscheidungsübersicht

Fallback-Kandidat

63
Bereit
Prototyp
Phase

recent repository activity

Audit

Installationsprüfung

Installations- und Adoptionsprüfung

77
Prüfung nötig
Sicherheit
80/100
Wartung
100/100
Installieren
92/100
Vollständiges Audit öffnenEval-Bericht ansehen

Von Agent belegte Evidenz

Von Agent belegte Evidenz

Ergebnisberichte nach Resolve, Prüfung, Installation und einem begrenzten Lauf.

0
Belegt
Needs first agent runAuto-Installation: zuerst prüfenLetzter: Unbekannt
Erfolgsrate
Letzter Fehler
Ergebnisse
0
Ausgabequalität
Fehlgeschlagen
0
Nicht relevant
0
Installationen
0
Durch Risiko blockiert
0
Einrichtung erforderlich
0
Produktion
0

Noch keine Agent-Ergebnisdaten. Der erste Lauf kann Erfolg, Einrichtungsbedarf, Risikoblockaden, Fehler oder Irrelevanz über /api/agent/outcome melden.

Installieren

Zum Agent-Workflow hinzufügen

Kostenlos und Open Source. Bericht vor der Installation in Produktions-Agents prüfen.

Wachstums-Loop

Share-Kit

X

Szenariobasierter Entwurf für annotator-input-parity-check, bereit für einen manuellen X-Post.

Kuratorenhinweis
annotator-input-parity-check: Before designing, training, or auditing ANY model that replicates human-annotated labels, aud...

47 stars

https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x
X-Entwurf öffnen
Optionale Antwort mit Installationsbefehl
Listing + install path for annotator-input-parity-check:
https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x

Install: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-...
Antwortentwurf öffnen

Quelle des Eintrags

Registry-indexiert

Beanspruchbar

Dieser Eintrag wurde aus öffentlichen Quellen indexiert und ist erst nach Genehmigung eines Maintainer-Anspruchs offiziell.

Ersteller
Claude Code
Indexiert von
OpenAgentSkill Community-Index

Die Zuordnung verlinkt auf das öffentliche Repository oder Creator-Profil. Creator können den Eintrag beanspruchen, um Eigentümersignale zu aktualisieren.

Diesen Skill beanspruchen

Eigentümeranspruch

Diesen Skill-Eintrag beanspruchen

Dieser Registry-indexiert-Eintrag wird Claude Code zugeschrieben, ist aber noch nicht offiziell markiert. Beanspruche ihn, um ein verifiziertes Eigentümersignal hinzuzufügen und künftige Launch-, Installations- und Audit-Updates vertrauenswürdiger zu machen.

Creator-Backlink-Kit

Evidenz-Badges in deine README einfügen

Zeige den kanonischen Eintrag, aktuelle Vertrauens- und Audit-Signale sowie echte Agent-Proven-Evidenz dort, wo Entwickler das Repository bewerten.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=listed&label=Listed)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=trust&label=Trust)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=audit&label=Audit)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/kennethkhoocy-annotator-input-parity-check?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)

Autor

C

Claude Code

@claude-code

Plattform-Fit

Gesundheitssignale

GitHub-Stars
47
Qualitätswert
35/100
Letzter GitHub-Push
24. Aug. 2026
Framework-Hinweise
Unbekannt
OpenAgentSkill-Aufrufe
0
Installationskopien
0
Externe Klicks
0

Community-Signal

Teile mit, ob dieser Skill für deinen Agent-Workflow nützlich ist. Zusammengefasstes Feedback verbessert das Ranking im Laufe der Zeit.

Vertrauen & Sicherheit

Nur Sandbox

63
  • GitHub-Akzeptanz47 GitHub-StarsPrüfen
  • Star-/Fork-Aktivität47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbarPrüfen
  • Aktuelle WartungHeute gepushtBestanden
  • LizenzklarheitMITBestanden
  • README/SKILL.md-VollständigkeitÖffentliche Metadaten benötigen mehr README/SKILL.md-KontextInfo
  • Abhängigkeits-/LaufzeitrisikoKeine wesentlichen Abhängigkeitsrisikohinweise in öffentlichen MetadatenBestanden