Ersteller · Claude Code
Letzte Aktualisierung · 24. Aug. 2026
llm-gold-bound-failure-check
Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/va
Nur Sandbox
Installationsziele
Codex-Installationsprompt
Install the "llm-gold-bound-failure-check" agent skill from https://github.com/kennethkhoocy/applied-micro-skills/tree/main/plugins/applied-micro/skills/llm-gold-bound-failure-check. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"kennethkhoocy-llm-gold-bound-failure-check","task":"Install llm-gold-bound-failure-check","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Asset-Profil
Design und kreative Produktion
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Szenario
Design und Kreativität
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Agent-Fit
Claude Code + CLI + Codex
Geeignet für Codex, Claude Code, Cursor, CLI oder benutzerdefinierte Agents.
Installieren
Bereit
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Wartung
Aktuell
Heute gepusht
Risiko
Prüfung nötig
Low GitHub adoption signal
GitHub-Qualität
47
64/100 Qualität · 79/100 Vertrauen
Abdeckungs-Tags
Review-Notizen
Low GitHub adoption signal · Quality score needs review
Agent-Adoptionskarte
Vertrauen, Audit und Installationsbereitschaft auf einen Blick
Diese Werte kombinieren öffentliche Repository-Metadaten, OpenAgentSkill-Reviewsignale, Wartungsaktualität und Installationsbereitschaft. Sie helfen bei der Vorauswahl, ersetzen aber keine menschliche Prüfung.
Qualität
VielversprechendUseful candidate, but compare it with alternatives before adopting.
Vertrauen
Nur SandboxNützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
Audit
Prüfung nötigMaschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
OpenAgentSkill Trust Score v5
Menschliche Prüfung vor Installation
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Stars
47 GitHub-Stars
Repository-Aktivität
47 Stars und 0 Forks
Wartung
Heute gepusht
Lizenz
MIT
Installieren
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Installationssicherheit
Standard-Paket- oder Laufzeit-Installationspfad
Berechtigungsfläche
Dateisystem- oder Dokumentzugriff
Agent-Ergebnisse
Noch keine Agent-Ergebnisdaten
Dokumentation
Starker README/SKILL.md-Kontext
Risikoübersicht
Vor Produktion prüfen
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
Installationsbereitschaft
Installationspfad verfügbar
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Lizenz ist angegeben
- Noch keine Agent-Proven-Ergebnisbelege
Agent-lesbare Metadaten
Maschinenlesbare Entscheidungsdaten für diesen Skill.
Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.
View technical data+
Agent-lesbare Metadaten
Maschinenlesbare Entscheidungsdaten für diesen Skill.
Nutze diesen Block oder das eingebettete JSON, um zu entscheiden, ob ein Agent diesen Skill installieren, eine Alternative wählen oder zuerst menschliche Prüfung anfordern soll.
Geeignete Aufgaben
- Document processing-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
- Read uploaded files
Geeignete Agents
Installationsentscheidung
- Befehl
- npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
- Richtlinie
- Prüfen
- Menschliche Prüfung
- Ja
Vertrauen und Risiko
- Vertrauen
- 71/100
- Audit
- 80/100
- Risikoebene
- Prüfung nötig
Ergebnis-Loop
- Endpoint
- /api/agent/outcome
- Event-ID
- resolve
- Ergebnisse
- 5
Installationsbefehl
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-checkNicht verwenden, wenn
- Teams, die ein vom Anbieter unterstütztes SLA benötigen
- production agents without a repository review
- Low GitHub adoption signal
- No OpenAgentSkill engagement data yet
- Quality score needs review
Alternative
Frontend Design
171.2K Stars
npx skills add anthropics/skills --skill frontend-design
Alternative
Taste Skill: Anti-Slop Frontend
79.7K Stars
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Alternative
Canvas Design
171.2K Stars
npx skills add anthropics/skills --skill canvas-design
Alternative
Anthropic Brand Guidelines
171.2K Stars
npx skills add anthropics/skills --skill brand-guidelines
Agent-Sicherheit v2
64/100 · Vor Installation prüfen
Nutzbarer Kandidat, aber der Agent sollte Berechtigungs- und Auditnotizen vor der Installation anzeigen.
Vor der Installation in einem echten Arbeitsbereich ist menschliche Freigabe erforderlich.
Mittel
Netzwerkzugriff
Die Skill ruft wahrscheinlich Remote-Seiten, APIs, Repositories oder externe Dienste ab.
Mittel
Dateisystemzugriff
Die Skill kann Projektdateien, Dokumente, generierte Artefakte oder den lokalen Arbeitsbereich lesen oder schreiben.
- Low GitHub adoption signal
Agent-Auflösungsplan
Lass einen Agent die Eignung vor der Installation prüfen.
Die Resolve API liefert die beste Skill, Alternativen, Sicherheitsrichtlinien, Auditnotizen, Installationsziel und einen direkt nutzbaren Prompt.
JSON öffnen
/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve-Text
/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Installationsübergabe
/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
Agent sollte prüfen
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Prompt kopieren
Task: Use llm-gold-bound-failure-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent-Übergabe
Gib dem Agent den Installationspfad, nicht noch ein Verzeichnis.
Über den öffentlichen Endpunkt erhältst du Befehl, Sicherheitscheckliste, Ziel-Prompts und kanonische Links.
Installationsübergabe
/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
LLM-Textformat
/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install?format=text
Alternativen finden
/api/skills/search?q=llm-gold-bound-failure-check&limit=3
Agent-Prompt
Use llm-gold-bound-failure-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-checkRegistry-Metadaten
Agent-lesbares Profil für die automatische Skill-Auswahl.
Die Registry API stellt Entscheidungs-, Vertrauens-, Audit-, Use-Case- und Installationssignale ohne UI-Scraping bereit.
Manifest
/api/registry/manifest/kennethkhoocy-llm-gold-bound-failure-check
LLM-Text
/api/registry/manifest/kennethkhoocy-llm-gold-bound-failure-check?format=text
Installationsalias
/api/registry/install/kennethkhoocy-llm-gold-bound-failure-check
Empfehlen
/api/registry/recommend?task=Use%20llm-gold-bound-failure-check%20in%20an%20agent%20workflow&limit=3
Agent-Fit
Document processing
Plattformen
Claude Code
Audit-Bericht
Prüfung nötig · 80/100
Maschinenlesbare Prüfung von Installationsbereitschaft, Sicherheitsmetadaten, Wartung und Akzeptanzrisiko.
Agent-Entscheidungspanel
Fallback candidate for Document processing
Prototype with this skill first; keep a fallback candidate ready.
Rolle im Stack
Fallback-Kandidat
Primäre Eignung
Document processing
Vertrauenslabel
Zuerst prototypisieren
Installationspfad
Befehl bereit
Verwenden wenn
- Document processing-Workflows
- Claude-Code-Teams
- builders willing to evaluate younger projects
Evidenz
- recent repository activity
- install command or GitHub repo available
- Qualitätsprofil 64/100
zuerst prüfen
- Low GitHub adoption signal
- No OpenAgentSkill engagement data yet
Implementierungspfad
- 1Installieren Sie es in einem Sandbox-Agent und führen Sie eine Document processing-Aufgabe vollständig aus.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Vertrauensprofil
Nur Sandbox
Nützlicher Kandidat mit fehlenden oder gemischten Vertrauenssignalen. Bis der Ergebniszyklus die Passung belegt, in einem isolierten Arbeitsbereich verwenden.
GitHub-Akzeptanz
Prüfen47 GitHub-Stars
Star-/Fork-Aktivität
Prüfen47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbar
Aktuelle Wartung
BestandenHeute gepusht
Lizenzklarheit
BestandenMIT
Positive Signale
- KI-Prüfung genehmigt
- Installationspfad ist verfügbar
- Repository-Belege sind verfügbar
- Kürzlich gewartetes Repository
- Der Installationsbefehl weist kein offensichtliches Hochrisikomuster auf
- Ergebniszyklus ist bereit, benötigt aber den ersten echten Agent-Lauf
Vor Installation prüfen
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
- Noch keine echten Agent-Ergebnisberichte
- Vor unbeaufsichtigter Installation ist menschliche Prüfung erforderlich
Empfohlene Aktion
Nur in einer Sandbox ausführen und nahe Alternativen vergleichen, bevor sie produktiv eingesetzt wird.
Qualitätsprofil
Vielversprechend Kandidat für Agent-Workflows
Useful candidate, but compare it with alternatives before adopting.
Workflow-Eignung
Diese Skill in diesen Szenarien nutzen
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Workflow-Eignung
Zum vollständigen Workflow hinzufügen
Design, build, test, and ship interfaces
Frontend and UI
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Alternativen-Shortlist
Vor Installation vergleichen
Similar skills that may fit this task.
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Taste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Canvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
Anthropic Brand Guidelines
Apply Anthropic official brand colors, typography, and visual standards to appropriate Anthropic-related artifacts.
Übersicht
--- name: llm-gold-bound-failure-check description: | Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). author: Claude Code version: 1.0.0 date: 2026-07-16 ---
# LLM Gold-Bound Failure Check
## Problem
When an LLM scoring pipeline over-predicts one label, the reflex fix is a prompt clarification ("score positive ONLY when..."). But if the gold standard itself does not separate the texts you want excluded from the texts it labels positive, the revision removes true and false positives together. The pilot fails, the spend is wasted, and — worse — an un-gated adoption would have silently destroyed recall in production.
## Context / Trigger Conditions
- A domain/label shows precision ≪ recall (e.g. P 0.46 / R 0.96) against gold - A prompt edit is proposed to exclude a specific text type (boilerplate, affirmative-program language, non-risk framing) - The label's gold council/inter-rater agreement was already the weakest (κ below ~0.6 is the warning sign that the construct is contested)
## Solution
**Step 0 — the ~$0 check, BEFORE building anything:** read a sample of gold POSITIVES for the weak label and ask: do they contain the feature the revision would exclude? Compare them side-by-side with the false positives.
- Gold positives and false positives are the same kind of text → the failure is **gold-bound**. Stop. No prompt passes a gold-scored gate. The levers are: (a) re-adjudicate the construct with the gold's owners (changes the gold, not the scores), or (b) re-interpret the shipped measure honestly (e.g. "discussion salience" instead of "risk exposure") in downstream analyses. - Gold positives clearly differ from the false positives → a prompt revision is plausible; proceed to a gated pilot.
**Gated pilot design (verified):** 1. Split gold into tune/holdout halves, stratified on the weak label's positives; fixed seed. 2. Draft ONE surgical edit from tune-half errors only — byte-identical elsewhere; verify the diff reverses cleanly. 3. Pre-register the gate on the holdout BEFORE scoring: target-label thresholds (e.g. precision ≥ X AND recall ≥ Y) plus a perturbation tolerance for untouched labels (e.g. within 0.03 F1 / 0.06 κ of a same-serving-rev fresh baseline). 4. Score everything fresh under both prompts (same model revision, same day — this doubles as the drift control). Never write through the production cache layer. 5. Adopt only on a full pass; a REJECT is a valid, cheap outcome.
## Verification
The pilot report shows: the exact prompt diff, tune-vs-holdout metrics for old and new prompts, per-label deltas on untouched sections, and spend. A gold-bound diagnosis is confirmed when the revision moves recall sharply down while precision stays roughly flat.
## Example
Specialist Directors US, 2026-07-16: DEI over-prediction (P 0.46 / R 0.96, council κ 0.24–0.59). A risk-framing-only DEI clause was piloted ($1.17, pre-registered holdout gate). Result: recall 0.895→0.263, precision 0.455 (gate ≥0.60) — REJECT. Reading the tune half showed ~¾ of gold DEI positives were pure affirmative D&I program text, identical in kind to the false positives; the failure was predictable at Step 0. Bonus finding: the DEI-section-only edit left all five other domains within 0.025 F1 / 0.05 κ — single-section prompt edits isolate cleanly, so the perturbation check is a cheap add, not paranoia. Same pattern one week earlier: a cyber classifier pilot gate failure traced to E/D gold contamination (misses were skills-matrix-checkbox-only positives), not model weakness.
## Notes
- Low inter-rater κ on a label is the leading indicator: contested construct → gold-bound failures downstream. - If the pipeline scores all labels in one completion, any post-campaign prompt change forces a full re-score — run this check BEFORE the campaign. - See also: [llm-campaign-drift-gate] for the companion gate on resume boundaries and serving-revision drift (same fresh-baseline discipline). - See also: [annotator-input-parity-check] — run it FIRST. If the model was never shown the document the annotators read, apparent gold-bound failures (e.g. the 2026-07-16 E/D "contamination" reading above) are actually input mismatch: the 2026-07-21 parity audit showed the specialist-director hand labels were pure proxy-statement transcriptions, so checkbox-only positives were recoverable from the right input all along.
Technische Details
- Version
- 1.0.0
- Lizenz
- MIT
- Letzte Aktualisierung
- 24. Aug. 2026
- Veröffentlicht
- 24. Aug. 2026
Entscheidungsübersicht
Fallback-Kandidat
recent repository activity
Audit
Installationsprüfung
Installations- und Adoptionsprüfung
- Sicherheit
- 86/100
- Wartung
- 100/100
- Installieren
- 92/100
Von Agent belegte Evidenz
Von Agent belegte Evidenz
Ergebnisberichte nach Resolve, Prüfung, Installation und einem begrenzten Lauf.
- Erfolgsrate
- —
- Letzter Fehler
- —
- Ergebnisse
- 0
- Ausgabequalität
- —
- Fehlgeschlagen
- 0
- Nicht relevant
- 0
- Installationen
- 0
- Durch Risiko blockiert
- 0
- Einrichtung erforderlich
- 0
- Produktion
- 0
Noch keine Agent-Ergebnisdaten. Der erste Lauf kann Erfolg, Einrichtungsbedarf, Risikoblockaden, Fehler oder Irrelevanz über /api/agent/outcome melden.
Installieren
Zum Agent-Workflow hinzufügen
Kostenlos und Open Source. Bericht vor der Installation in Produktions-Agents prüfen.
Wachstums-Loop
Share-Kit
Szenariobasierter Entwurf für llm-gold-bound-failure-check, bereit für einen manuellen X-Post.
llm-gold-bound-failure-check: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on... 47 stars https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x
Optionale Antwort mit Installationsbefehl
Listing + install path for llm-gold-bound-failure-check: https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x Install: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-...
Quelle des Eintrags
Registry-indexiert
Dieser Eintrag wurde aus öffentlichen Quellen indexiert und ist erst nach Genehmigung eines Maintainer-Anspruchs offiziell.
- Ersteller
- Claude Code
- Indexiert von
- OpenAgentSkill Community-Index
Die Zuordnung verlinkt auf das öffentliche Repository oder Creator-Profil. Creator können den Eintrag beanspruchen, um Eigentümersignale zu aktualisieren.
Diesen Skill beanspruchenEigentümeranspruch
Diesen Skill-Eintrag beanspruchen
Dieser Registry-indexiert-Eintrag wird Claude Code zugeschrieben, ist aber noch nicht offiziell markiert. Beanspruche ihn, um ein verifiziertes Eigentümersignal hinzuzufügen und künftige Launch-, Installations- und Audit-Updates vertrauenswürdiger zu machen.
Creator-Backlink-Kit
Evidenz-Badges in deine README einfügen
Zeige den kanonischen Eintrag, aktuelle Vertrauens- und Audit-Signale sowie echte Agent-Proven-Evidenz dort, wo Entwickler das Repository bewerten.
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check/audit)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)Autor
Claude Code
@claude-code
Tags
Plattform-Fit
Gesundheitssignale
- GitHub-Stars
- 47
- Qualitätswert
- 35/100
- Letzter GitHub-Push
- 24. Aug. 2026
- Framework-Hinweise
- Unbekannt
- OpenAgentSkill-Aufrufe
- 0
- Installationskopien
- 0
- Externe Klicks
- 0
Community-Signal
Teile mit, ob dieser Skill für deinen Agent-Workflow nützlich ist. Zusammengefasstes Feedback verbessert das Ranking im Laufe der Zeit.
Vertrauen & Sicherheit
Nur Sandbox
- GitHub-Akzeptanz47 GitHub-StarsPrüfen
- Star-/Fork-Aktivität47 Stars und 0 Forks; Issue-Aktivität ist in den aktuellen Metadaten nicht verfügbarPrüfen
- Aktuelle WartungHeute gepushtBestanden
- LizenzklarheitMITBestanden
- README/SKILL.md-VollständigkeitMetadaten enthalten ausreichend Nutzungs- und Workflow-KontextBestanden
- Abhängigkeits-/LaufzeitrisikoKeine wesentlichen Abhängigkeitsrisikohinweise in öffentlichen MetadatenBestanden
Ähnliche Skills
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
171.2K StarsTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
79.7K StarsCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
171.2K StarsAnthropic Brand Guidelines
Apply Anthropic official brand colors, typography, and visual standards to appropriate Anthropic-related artifacts.
171.2K Stars