@kennethkhoocy

Ersteller · Claude Code

Letzte Aktualisierung · 24. Aug. 2026

llm-gold-bound-failure-check

Prüfen · 71Im Registry indexiert

Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/va

OpenAgentSkill Trust Score
71/100

Nur Sandbox

Qualität64/100
Audit80/100
Stars47
Verified installs0

Installationsziele

Codex-Installationsprompt

Install the "llm-gold-bound-failure-check" agent skill from https://github.com/kennethkhoocy/applied-micro-skills/tree/main/plugins/applied-micro/skills/llm-gold-bound-failure-check. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"kennethkhoocy-llm-gold-bound-failure-check","task":"Install llm-gold-bound-failure-check","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.

Asset-Profil

Design und kreative Produktion

Design assets, images, video, audio, multimodal media, presentation, and creative production skills.

Bereich ansehen

Szenario

Design und Kreativität

I need my agent to produce design assets, UI directions, presentations, or creative media workflows.

Agent-Fit

Claude Code + CLI + Codex

Geeignet für Codex, Claude Code, Cursor, CLI oder benutzerdefinierte Agents.

Installieren

Bereit

npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

Wartung

Aktuell

Heute gepusht

Risiko

Prüfung nötig

Low GitHub adoption signal

GitHub-Qualität

47

64/100 Qualität · 79/100 Vertrauen

Abdeckungs-Tags

DesignDesign und KreativitätDesign und Kreativitätagent-skill

Review-Notizen

Low GitHub adoption signal · Quality score needs review

Agent-Adoptionskarte

Vertrauen, Audit und Installationsbereitschaft auf einen Blick

Diese Werte kombinieren öffentliche Repository-Metadaten, OpenAgentSkill-Reviewsignale, Wartungsaktualität und Installationsbereitschaft. Sie helfen bei der Vorauswahl, ersetzen aber keine menschliche Prüfung.

Qualität

Vielversprechend
64

Useful candidate, but compare it with alternatives before adopting.

Vertrauen

Nur Sandbox
71

Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.

Audit

Prüfung nötig
80

Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.

OpenAgentSkill Trust Score v5

Menschliche Prüfung vor Installation

Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

47 GitHub-Stars

Repository-Aktivität

47 Stars und 0 Forks

Wartung

Heute gepusht

Lizenz

MIT

Installieren

npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

Installationssicherheit

Standard-Paket- oder Laufzeit-Installationspfad

Berechtigungsfläche

Dateisystem- oder Dokumentzugriff

Agent-Ergebnisse

Noch keine Agent-Ergebnisdaten

Dokumentation

Starker README/SKILL.md-Kontext

Risikoübersicht

Vor Produktion prüfen

  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata

Installationsbereitschaft

Installationspfad verfügbar

  • Installationspfad ist verfügbar
  • Repository-Belege sind verfügbar
  • Lizenz ist angegeben
  • Noch keine Agent-Proven-Ergebnisbelege

Agent-lesbare Metadaten

Maschinenlesbare Entscheidungsdaten für diesen Skill.

Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.

View technical data+

Geeignete Aufgaben

  • Document processing-Workflows
  • Claude-Code-Teams
  • builders willing to evaluate younger projects
  • Read uploaded files

Geeignete Agents

CodexClaude CodeCursorOpenAgentSkill CLICLI

Installationsentscheidung

Befehl
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Richtlinie
Prüfen
Menschliche Prüfung
Ja

Vertrauen und Risiko

Vertrauen
71/100
Audit
80/100
Risikoebene
Prüfung nötig

Ergebnis-Loop

Endpoint
/api/agent/outcome
Event-ID
resolve
Ergebnisse
5

Installationsbefehl

npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

Nicht verwenden, wenn

  • Teams, die ein vom Anbieter unterstütztes SLA benötigen
  • production agents without a repository review
  • Low GitHub adoption signal
  • No OpenAgentSkill engagement data yet
  • Quality score needs review

Agent-Sicherheit v2

64/100 · Vor Installation prüfen

Mit Berechtigungshinweisen geprüftPrüfen

Nutzbarer Kandidat, aber der Agent sollte Berechtigungs- und Auditnotizen vor der Installation anzeigen.

Vor der Installation in einem echten Arbeitsbereich ist menschliche Freigabe erforderlich.

Per API auflösen

Mittel

Netzwerkzugriff

Die Skill ruft wahrscheinlich Remote-Seiten, APIs, Repositories oder externe Dienste ab.

Mittel

Dateisystemzugriff

Die Skill kann Projektdateien, Dokumente, generierte Artefakte oder den lokalen Arbeitsbereich lesen oder schreiben.

  • Low GitHub adoption signal

Agent-Auflösungsplan

Lass einen Agent die Eignung vor der Installation prüfen.

Die Resolve API liefert die beste Skill, Alternativen, Sicherheitsrichtlinien, Auditnotizen, Installationsziel und einen direkt nutzbaren Prompt.

Textplan öffnen

Agent sollte prüfen

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Prompt kopieren

Task: Use llm-gold-bound-failure-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent-Übergabe

Gib dem Agent den Installationspfad, nicht noch ein Verzeichnis.

Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.

Installations-API öffnen

Agent-Prompt

Use llm-gold-bound-failure-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check

Registry-Metadaten

Agent-lesbares Profil für die automatische Skill-Auswahl.

Die Registry API stellt Entscheidungs-, Vertrauens-, Audit-, Use-Case- und Installationssignale ohne UI-Scraping bereit.

Manifest öffnen

Agent-Fit

63/100

Document processing

Plattformen

Claude Code

Audit-Bericht

Prüfung nötig · 80/100

Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.

Audit-Bericht ansehenEval-Bericht ansehen

Agent-Entscheidungspanel

Fallback candidate for Document processing

Prototype with this skill first; keep a fallback candidate ready.

63
Bereitschaft
Prototyp
Phase

Rolle im Stack

Fallback-Kandidat

Primäre Eignung

Document processing

Vertrauenslabel

Zuerst prototypisieren

Installationspfad

Befehl bereit

Verwenden wenn

  • Document processing-Workflows
  • Claude-Code-Teams
  • builders willing to evaluate younger projects

Evidenz

  • recent repository activity
  • install command or GitHub repo available
  • Qualitätsprofil 64/100

zuerst prüfen

  • Low GitHub adoption signal
  • No OpenAgentSkill engagement data yet

Implementierungspfad

  1. 1Installieren Sie es in einem Sandbox-Agent und führen Sie eine Document processing-Aufgabe vollständig aus.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Vertrauensprofil

Nur Sandbox

Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.

71
OpenAgentSkill Trust Score

GitHub-Akzeptanz

Prüfen

47 GitHub-Stars

Star-/Fork-Aktivität

Prüfen

47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbar

Aktuelle Wartung

Bestanden

Heute gepusht

Lizenzklarheit

Bestanden

MIT

Positive Signale

  • KI-Prüfung genehmigt
  • Installationspfad ist verfügbar
  • Repository-Belege sind verfügbar
  • Kürzlich gewartetes Repository
  • Der Installationsbefehl weist kein offensichtliches Hochrisikomuster auf
  • Ergebniszyklus ist bereit, benötigt aber den ersten echten Agent-Lauf

Vor Installation prüfen

  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
  • Noch keine echten Agent-Ergebnisberichte
  • Vor unbeaufsichtigter Installation ist menschliche Prüfung erforderlich

Empfohlene Aktion

Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.

Qualitätsprofil

Vielversprechend Kandidat für Agent-Workflows

Useful candidate, but compare it with alternatives before adopting.

64
GitHub-Stars
47
Aktualität
Heute
Installationsbereit
Ja
Lizenz
MIT
Vor Installation prüfen: Low GitHub adoption signal

Workflow-Eignung

Diese Skill in diesen Szenarien nutzen

Workflow-Eignung

Zum vollständigen Workflow hinzufügen

Alternativen-Shortlist

Vor Installation vergleichen

Similar skills that may fit this task.

Alle vergleichen

Übersicht

--- name: llm-gold-bound-failure-check description: | Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). author: Claude Code version: 1.0.0 date: 2026-07-16 ---

# LLM Gold-Bound Failure Check

## Problem

When an LLM scoring pipeline over-predicts one label, the reflex fix is a prompt clarification ("score positive ONLY when..."). But if the gold standard itself does not separate the texts you want excluded from the texts it labels positive, the revision removes true and false positives together. The pilot fails, the spend is wasted, and — worse — an un-gated adoption would have silently destroyed recall in production.

## Context / Trigger Conditions

- A domain/label shows precision ≪ recall (e.g. P 0.46 / R 0.96) against gold - A prompt edit is proposed to exclude a specific text type (boilerplate, affirmative-program language, non-risk framing) - The label's gold council/inter-rater agreement was already the weakest (κ below ~0.6 is the warning sign that the construct is contested)

## Solution

**Step 0 — the ~$0 check, BEFORE building anything:** read a sample of gold POSITIVES for the weak label and ask: do they contain the feature the revision would exclude? Compare them side-by-side with the false positives.

- Gold positives and false positives are the same kind of text → the failure is **gold-bound**. Stop. No prompt passes a gold-scored gate. The levers are: (a) re-adjudicate the construct with the gold's owners (changes the gold, not the scores), or (b) re-interpret the shipped measure honestly (e.g. "discussion salience" instead of "risk exposure") in downstream analyses. - Gold positives clearly differ from the false positives → a prompt revision is plausible; proceed to a gated pilot.

**Gated pilot design (verified):** 1. Split gold into tune/holdout halves, stratified on the weak label's positives; fixed seed. 2. Draft ONE surgical edit from tune-half errors only — byte-identical elsewhere; verify the diff reverses cleanly. 3. Pre-register the gate on the holdout BEFORE scoring: target-label thresholds (e.g. precision ≥ X AND recall ≥ Y) plus a perturbation tolerance for untouched labels (e.g. within 0.03 F1 / 0.06 κ of a same-serving-rev fresh baseline). 4. Score everything fresh under both prompts (same model revision, same day — this doubles as the drift control). Never write through the production cache layer. 5. Adopt only on a full pass; a REJECT is a valid, cheap outcome.

## Verification

The pilot report shows: the exact prompt diff, tune-vs-holdout metrics for old and new prompts, per-label deltas on untouched sections, and spend. A gold-bound diagnosis is confirmed when the revision moves recall sharply down while precision stays roughly flat.

## Example

Specialist Directors US, 2026-07-16: DEI over-prediction (P 0.46 / R 0.96, council κ 0.24–0.59). A risk-framing-only DEI clause was piloted ($1.17, pre-registered holdout gate). Result: recall 0.895→0.263, precision 0.455 (gate ≥0.60) — REJECT. Reading the tune half showed ~¾ of gold DEI positives were pure affirmative D&I program text, identical in kind to the false positives; the failure was predictable at Step 0. Bonus finding: the DEI-section-only edit left all five other domains within 0.025 F1 / 0.05 κ — single-section prompt edits isolate cleanly, so the perturbation check is a cheap add, not paranoia. Same pattern one week earlier: a cyber classifier pilot gate failure traced to E/D gold contamination (misses were skills-matrix-checkbox-only positives), not model weakness.

## Notes

- Low inter-rater κ on a label is the leading indicator: contested construct → gold-bound failures downstream. - If the pipeline scores all labels in one completion, any post-campaign prompt change forces a full re-score — run this check BEFORE the campaign. - See also: [llm-campaign-drift-gate] for the companion gate on resume boundaries and serving-revision drift (same fresh-baseline discipline). - See also: [annotator-input-parity-check] — run it FIRST. If the model was never shown the document the annotators read, apparent gold-bound failures (e.g. the 2026-07-16 E/D "contamination" reading above) are actually input mismatch: the 2026-07-21 parity audit showed the specialist-director hand labels were pure proxy-statement transcriptions, so checkbox-only positives were recoverable from the right input all along.

Technische Details

Version
1.0.0
Lizenz
MIT
Letzte Aktualisierung
24. Aug. 2026
Veröffentlicht
24. Aug. 2026

Entscheidungsübersicht

Fallback-Kandidat

63
Bereit
Prototyp
Phase

recent repository activity

Audit

Installationsprüfung

Installations- und Adoptionsprüfung

80
Prüfung nötig
Sicherheit
86/100
Wartung
100/100
Installieren
92/100
Vollständiges Audit öffnenEval-Bericht ansehen

Von Agent belegte Evidenz

Von Agent belegte Evidenz

Ergebnisberichte nach Resolve, Prüfung, Installation und einem begrenzten Lauf.

0
Belegt
Needs first agent runAuto-Installation: zuerst prüfenLetzter: Unbekannt
Erfolgsrate
Letzter Fehler
Ergebnisse
0
Ausgabequalität
Fehlgeschlagen
0
Nicht relevant
0
Installationen
0
Durch Risiko blockiert
0
Einrichtung erforderlich
0
Produktion
0

Noch keine Agent-Ergebnisdaten. Der erste Lauf kann Erfolg, Einrichtungsbedarf, Risikoblockaden, Fehler oder Irrelevanz über /api/agent/outcome melden.

Installieren

Zum Agent-Workflow hinzufügen

Kostenlos und Open Source. Bericht vor der Installation in Produktions-Agents prüfen.

Wachstums-Loop

Share-Kit

X

Szenariobasierter Entwurf für llm-gold-bound-failure-check, bereit für einen manuellen X-Post.

Kuratorenhinweis
llm-gold-bound-failure-check: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on...

47 stars

https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x
X-Entwurf öffnen
Optionale Antwort mit Installationsbefehl
Listing + install path for llm-gold-bound-failure-check:
https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x

Install: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-...
Antwortentwurf öffnen

Quelle des Eintrags

Registry-indexiert

Beanspruchbar

Dieser Eintrag wurde aus öffentlichen Quellen indexiert und ist erst nach Genehmigung eines Maintainer-Anspruchs offiziell.

Ersteller
Claude Code
Indexiert von
OpenAgentSkill Community-Index

Die Zuordnung verlinkt auf das öffentliche Repository oder Creator-Profil. Creator können den Eintrag beanspruchen, um Eigentümersignale zu aktualisieren.

Diesen Skill beanspruchen

Eigentümeranspruch

Diesen Skill-Eintrag beanspruchen

Dieser Registry-indexiert-Eintrag wird Claude Code zugeschrieben, ist aber noch nicht offiziell markiert. Beanspruche ihn, um ein verifiziertes Eigentümersignal hinzuzufügen und künftige Launch-, Installations- und Audit-Updates vertrauenswürdiger zu machen.

Creator-Backlink-Kit

Evidenz-Badges in deine README einfügen

Zeige den kanonischen Eintrag, aktuelle Vertrauens- und Audit-Signale sowie echte Agent-Proven-Evidenz dort, wo Entwickler das Repository bewerten.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=listed&label=Listed)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=trust&label=Trust)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=audit&label=Audit)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-gold-bound-failure-check?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)

Autor

C

Claude Code

@claude-code

Plattform-Fit

Gesundheitssignale

GitHub-Stars
47
Qualitätswert
35/100
Letzter GitHub-Push
24. Aug. 2026
Framework-Hinweise
Unbekannt
OpenAgentSkill-Aufrufe
0
Installationskopien
0
Externe Klicks
0

Community-Signal

Teile mit, ob dieser Skill für deinen Agent-Workflow nützlich ist. Zusammengefasstes Feedback verbessert das Ranking im Laufe der Zeit.

Vertrauen & Sicherheit

Nur Sandbox

71
  • GitHub-Akzeptanz47 GitHub-StarsPrüfen
  • Star-/Fork-Aktivität47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbarPrüfen
  • Aktuelle WartungHeute gepushtBestanden
  • LizenzklarheitMITBestanden
  • README/SKILL.md-VollständigkeitMetadaten enthalten ausreichend Nutzungs- und Workflow-KontextBestanden
  • Abhängigkeits-/LaufzeitrisikoKeine wesentlichen Abhängigkeitsrisikohinweise in öffentlichen MetadatenBestanden