annotator-input-parity-check
Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whos
Asset-Profil
Recherche und Wissensarbeit
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Szenario
Recherche-Agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent-Fit
Claude Code + CLI + Codex
Geeignet für Codex, Claude Code, Cursor, CLI oder benutzerdefinierte Agents.
Installieren
Bereit
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Wartung
Aktuell
Heute gepusht
Risiko
Prüfung nötig
No explicit 'Limitations' section, though the Notes section partially covers boundaries.
GitHub-Qualität
47
64/100 Qualität · 71/100 Vertrauen
Abdeckungs-Tags
Review-Notizen
No explicit 'Limitations' section, though the Notes section partially covers boundaries. · The skill description is long but well-structured; could be slightly more concise for quick scanning.
Agent-Adoptionskarte
Vertrauen, Audit und Installationsbereitschaft auf einen Blick
Diese Werte kombinieren öffentliche Repository-Metadaten, OpenAgentSkill-Reviewsignale, Wartungsaktualität und Installationsbereitschaft. Sie helfen bei der Vorauswahl, ersetzen aber keine menschliche Prüfung.
Qualität
VielversprechendUseful candidate, but compare it with alternatives before adopting.
Vertrauen
Nur SandboxNützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
Audit
Prüfung nötigMaschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
OpenAgentSkill Trust Score v5
Menschliche Prüfung vor Installation
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Stars
47 GitHub-Stars
Repository-Aktivität
47 Stars und 0 Forks
Wartung
Heute gepusht
Lizenz
MIT
Installieren
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Installationssicherheit
Standard-Paket- oder Laufzeit-Installationspfad
Berechtigungsfläche
Dateisystem- oder Dokumentzugriff
Agent-Ergebnisse
Noch keine Agent-Ergebnisdaten
Dokumentation
Usable metadata, review docs
Risikoübersicht
Vor Produktion prüfen
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
Installationsbereitschaft
Installationspfad verfügbar
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Lizenz ist angegeben
- Noch keine Agent-Proven-Ergebnisbelege
Agent-lesbare Metadaten
Maschinenlesbare Entscheidungsdaten für diesen Skill.
Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.
Geeignete Aufgaben
- Research-Agent-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
- Suchquellen
Geeignete Agents
Installationsentscheidung
- Befehl
- npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
- Richtlinie
- Prüfen
- Menschliche Prüfung
- Ja
Vertrauen und Risiko
- Vertrauen
- 63/100
- Audit
- 77/100
- Risikoebene
- Prüfung nötig
Ergebnis-Loop
- Endpoint
- /api/agent/outcome
- Event-ID
- resolve
- Ergebnisse
- 5
Installationsbefehl
npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-checkNicht verwenden, wenn
- Teams, die ein vom Anbieter unterstütztes SLA benötigen
- production agents without a repository review
- Low GitHub adoption signal
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- No OpenAgentSkill engagement data yet
Agent-Sicherheit v2
61/100 · Vor Installation prüfen
Nutzbarer Kandidat, aber der Agent sollte Berechtigungs- und Auditnotizen vor der Installation anzeigen.
Vor der Installation in einem echten Arbeitsbereich ist menschliche Freigabe erforderlich.
Mittel
Netzwerkzugriff
Die Skill ruft wahrscheinlich Remote-Seiten, APIs, Repositories oder externe Dienste ab.
Mittel
Dateisystemzugriff
Die Skill kann Projektdateien, Dokumente, generierte Artefakte oder den lokalen Arbeitsbereich lesen oder schreiben.
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
Installationsziele
Diesen Skill im Agent-Workflow installieren
Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install kennethkhoocy-annotator-input-parity-checkAgent-Auflösungsplan
Lass einen Agent die Eignung vor der Installation prüfen.
Die Resolve API liefert die beste Skill, Alternativen, Sicherheitsrichtlinien, Auditnotizen, Installationsziel und einen direkt nutzbaren Prompt.
JSON öffnen
/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve-Text
/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Installationsübergabe
/api/skills/kennethkhoocy-annotator-input-parity-check/install
Agent sollte prüfen
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Prompt kopieren
Task: Use annotator-input-parity-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20annotator-input-parity-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent-Übergabe
Gib dem Agent den Installationspfad, nicht noch ein Verzeichnis.
Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.
Installationsübergabe
/api/skills/kennethkhoocy-annotator-input-parity-check/install
LLM-Textformat
/api/skills/kennethkhoocy-annotator-input-parity-check/install?format=text
Alternativen finden
/api/skills/search?q=annotator-input-parity-check&limit=3
Agent-Prompt
Use annotator-input-parity-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-annotator-input-parity-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-checkRegistry-Metadaten
Agent-lesbares Profil für die automatische Skill-Auswahl.
Die Registry API stellt Entscheidungs-, Vertrauens-, Audit-, Use-Case- und Installationssignale ohne UI-Scraping bereit.
Manifest
/api/registry/manifest/kennethkhoocy-annotator-input-parity-check
LLM-Text
/api/registry/manifest/kennethkhoocy-annotator-input-parity-check?format=text
Installationsalias
/api/registry/install/kennethkhoocy-annotator-input-parity-check
Empfehlen
/api/registry/recommend?task=Use%20annotator-input-parity-check%20in%20an%20agent%20workflow&limit=3
Agent-Fit
Recherche-Agents
Plattformen
Claude Code
Audit-Bericht
Prüfung nötig · 77/100
Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
Agent-Entscheidungspanel
Fallback candidate for Research agents
Prototype with this skill first; keep a fallback candidate ready.
Rolle im Stack
Fallback-Kandidat
Primäre Eignung
Recherche-Agents
Vertrauenslabel
Zuerst prototypisieren
Installationspfad
Befehl bereit
Verwenden wenn
- Research-Agent-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
Evidenz
- recent repository activity
- install command or GitHub repo available
- Qualitätsprofil 64/100
zuerst prüfen
- Low GitHub adoption signal
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- No OpenAgentSkill engagement data yet
Implementierungspfad
- 1Installieren Sie es in einem Sandbox-Agent und führen Sie eine Recherche-Agents-Aufgabe vollständig aus.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Vertrauensprofil
Nur Sandbox
Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
GitHub-Akzeptanz
Prüfen47 GitHub-Stars
Star-/Fork-Aktivität
Prüfen47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbar
Aktuelle Wartung
BestandenHeute gepusht
Lizenzklarheit
BestandenMIT
Positive Signale
- KI-Prüfung genehmigt
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Kürzlich gewartetes Repository
- Der Installationsbefehl weist kein offensichtliches Hochrisikomuster auf
- Ergebniszyklus ist bereit, benötigt aber den ersten echten Agent-Lauf
Vor Installation prüfen
- No explicit 'Limitations' section, though the Notes section partially covers boundaries.
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
- Noch keine echten Agent-Ergebnisberichte
- Vor unbeaufsichtigter Installation ist menschliche Prüfung erforderlich
Empfohlene Aktion
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Qualitätsprofil
Vielversprechend Kandidat für Agent-Workflows
Useful candidate, but compare it with alternatives before adopting.
Workflow-Eignung
Diese Skill in diesen Szenarien nutzen
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow-Eignung
Zum vollständigen Workflow hinzufügen
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternativen-Shortlist
Vor Installation vergleichen
Similar skills that may fit this task.
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
Maigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
Übersicht
--- name: annotator-input-parity-check description: | Before designing, training, or auditing ANY model that replicates human-annotated labels, audit the annotation protocol's INPUT — the exact document/evidence the human labelers consulted — and give the model that same input. Use when: (1) designing a classifier/LLM extractor whose target is a hand-coded label set, (2) a label-replication model shows low recall concentrated in a label subset and the diagnosis on offer is "the label's information is not in the features", (3) reviewers propose construct splits (e.g. "designation vs record-evident"), adjudication sittings, or per-domain stop rules to explain residual disagreement with gold, (4) validating an extraction pipeline against labels transcribed from a source document. Symptom of the underlying failure: elaborate theory accumulates to explain why gold is "partially unpredictable" when the model was simply never shown the document the annotators read. author: Claude Code version: 1.0.0 date: 2026-07-21 ---
# Annotator Input Parity Check
## Problem
A model built to replicate human labels is fed a different evidence base than the one the annotators used. The mismatch masquerades as a modeling or construct problem: recall collapses on the label subset whose evidence lives only in the annotators' source, audits produce increasingly sophisticated theory ("invisible" positives, construct splits, per-domain reliability gates), and successive model generations inherit the wrong input because each review critiques the lineage from inside the frozen input assumption.
## Context / Trigger Conditions
- Starting any label-replication build (classifier, LLM scorer, extractor) against hand-coded gold. - A validation report says some share of gold positives have "zero signal" in the model's input. - Proposals appear for: construct splits (what the model CAN see vs what the label encodes), human adjudication of "contested" cells, stop rules excluding weak domains, or accepting a permanent accuracy ceiling. - Verified instance (Specialist Directors US, 2026-07-21): three classifier generations (bio-BERT AUC 0.5 → structured RoBERTa "unclassifiable" on 3/5 domains → LLM dossier scorer with E/D construct split + PI adjudication + per-domain stop rules) all read director bios + BoardEx records, while the RA labels were pure transcriptions of PROXY-STATEMENT disclosures (skills matrices + bios, no exogenous data — confirmed in the source paper's methodology, 41 Yale J. Reg. 652, 669-72). The "invisible specialist" mass (43-79% of some domains) was simply the skills-matrix checkbox content the models were never shown. Years of downstream apparatus dissolved once the question "what did the labelers actually read?" was asked.
## Solution
1. Before any design work, write down the annotation protocol as the annotators executed it: source document(s), what they could see, what they could not, whether any exogenous data entered. Get this from the codebook/paper methodology section, not from folklore. If the protocol is unwritten, ask the PI directly: "did labelers consult anything beyond X?" 2. Compare against the model's planned input. Any evidence the annotators had that the model lacks is a hard recall ceiling on exactly the labels that evidence determines — no architecture, prompt, or training fixes it. 3. If a mismatch exists, prefer restoring input parity (give the model the annotators' document) over modeling around the gap. For transcription-style protocols, the task then becomes extraction, not prediction, and validation against the hand labels becomes construct-matched (agreement should be high; disagreement means extraction bugs, not construct philosophy). 4. Only if input parity is impossible (annotators used private knowledge, interviews, paywalled data) is a construct split the honest design — and then the model's output must be named as a DIFFERENT variable, never graded raw against the full gold. 5. When auditing an EXISTING lineage: ask the parity question first, before critiquing rubrics, thresholds, or gold quality. An audit that inherits the input assumption can be internally excellent and still miss the dominant error term.
## Verification
- The protocol-input inventory exists in writing and the model input is a superset of it → recall ceilings from "invisible" labels should disappear; residual disagreement decomposes into extraction errors (fixable) rather than unknowable-label mass. - Quick falsification test for a claimed "unpredictable" label subset: pull 5 such gold positives, open the annotators' source document for each, and check whether the label is visible there. If yes, the problem is input, not construct.
## Notes
- Distinct from [llm-gold-bound-failure-check], which diagnoses gold that fails to SEPARATE classes for a proposed revision; this skill diagnoses model INPUT that omits the annotators' evidence. Run this parity check first — gold-bound analysis of a parity-broken system wastes effort. - The mismatch is self-perpetuating across model generations: each successor inherits the predecessor's feature pipeline, and each audit optimizes within it. Breaking the frame requires asking about the ANNOTATORS, not the model. - Construct splits built on a parity-broken system may still have salvage value for a different question (e.g. record-evident-but-undisclosed expertise is analytically interesting in its own right) — reframe, don't necessarily discard.
Technische Details
- Version
- 1.0.0
- Lizenz
- MIT
- Letzte Aktualisierung
- 24. Aug. 2026
- Veröffentlicht
- 24. Aug. 2026
Entscheidungsübersicht
Fallback-Kandidat
recent repository activity
Audit
Installationsprüfung
Installations- und Adoptionsprüfung
- Sicherheit
- 80/100
- Wartung
- 100/100
- Installieren
- 92/100
Von Agent belegte Evidenz
Von Agent belegte Evidenz
Ergebnisberichte nach Resolve, Prüfung, Installation und einem begrenzten Lauf.
- Erfolgsrate
- —
- Letzter Fehler
- —
- Ergebnisse
- 0
- Ausgabequalität
- —
- Fehlgeschlagen
- 0
- Nicht relevant
- 0
- Installationen
- 0
- Durch Risiko blockiert
- 0
- Einrichtung erforderlich
- 0
- Produktion
- 0
Noch keine Agent-Ergebnisdaten. Der erste Lauf kann Erfolg, Einrichtungsbedarf, Risikoblockaden, Fehler oder Irrelevanz über /api/agent/outcome melden.
Installieren
Zum Agent-Workflow hinzufügen
Kostenlos und Open Source. Bericht vor der Installation in Produktions-Agents prüfen.
Wachstums-Loop
Share-Kit
Szenariobasierter Entwurf für annotator-input-parity-check, bereit für einen manuellen X-Post.
annotator-input-parity-check: Before designing, training, or auditing ANY model that replicates human-annotated labels, aud... 47 stars https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x
Optionale Antwort mit Installationsbefehl
Listing + install path for annotator-input-parity-check: https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check?ref=x Install: npx skills add kennethkhoocy/applied-micro-skills --skill annotator-input-parity-...
Quelle des Eintrags
Registry-indexiert
Dieser Eintrag wurde aus öffentlichen Quellen indexiert und ist erst nach Genehmigung eines Maintainer-Anspruchs offiziell.
- Ersteller
- Claude Code
- Indexiert von
- OpenAgentSkill Community-Index
Die Zuordnung verlinkt auf das öffentliche Repository oder Creator-Profil. Creator können den Eintrag beanspruchen, um Eigentümersignale zu aktualisieren.
Diesen Skill beanspruchenEigentümeranspruch
Diesen Skill-Eintrag beanspruchen
Dieser Registry-indexiert-Eintrag wird Claude Code zugeschrieben, ist aber noch nicht offiziell markiert. Beanspruche ihn, um ein verifiziertes Eigentümersignal hinzuzufügen und künftige Launch-, Installations- und Audit-Updates vertrauenswürdiger zu machen.
Creator-Backlink-Kit
Evidenz-Badges in deine README einfügen
Zeige den kanonischen Eintrag, aktuelle Vertrauens- und Audit-Signale sowie echte Agent-Proven-Evidenz dort, wo Entwickler das Repository bewerten.
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check/audit)
[](https://www.openagentskill.com/skills/kennethkhoocy-annotator-input-parity-check)Autor
Claude Code
@claude-code
Tags
Plattform-Fit
Gesundheitssignale
- GitHub-Stars
- 47
- Qualitätswert
- 35/100
- Letzter GitHub-Push
- 24. Aug. 2026
- Framework-Hinweise
- Unbekannt
- OpenAgentSkill-Aufrufe
- 0
- Installationskopien
- 0
- Externe Klicks
- 0
Community-Signal
Teile mit, ob dieser Skill für deinen Agent-Workflow nützlich ist. Zusammengefasstes Feedback verbessert das Ranking im Laufe der Zeit.
Vertrauen & Sicherheit
Nur Sandbox
- GitHub-Akzeptanz47 GitHub-StarsPrüfen
- Star-/Fork-Aktivität47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbarPrüfen
- Aktuelle WartungHeute gepushtBestanden
- LizenzklarheitMITBestanden
- README/SKILL.md-VollständigkeitÖffentliche Metadaten benötigen mehr README/SKILL.md-KontextInfo
- Abhängigkeits-/LaufzeitrisikoKeine wesentlichen Abhängigkeitsrisikohinweise in öffentlichen MetadatenBestanden
Ähnliche Skills
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K Stars