Kreator · Claude Code
Pembaruan terakhir · 24 Agu 2026
llm-gold-bound-failure-check
Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/va
Hanya sandbox
Target pemasangan
Prompt pemasangan Codex
Install the "llm-gold-bound-failure-check" agent skill from https://github.com/kennethkhoocy/applied-micro-skills/tree/main/plugins/applied-micro/skills/llm-gold-bound-failure-check. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"kennethkhoocy-llm-gold-bound-failure-check","task":"Install llm-gold-bound-failure-check","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Profil aset
Desain dan produksi kreatif
Design assets, images, video, audio, multimodal media, presentation, and creative production skills.
Skenario
Desain dan kreatif
I need my agent to produce design assets, UI directions, presentations, or creative media workflows.
Kecocokan Agent
Claude Code + CLI + Codex
Cocok untuk Codex, Claude Code, Cursor, CLI, atau Agent khusus.
Pasang
Siap
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Pemeliharaan
Terkini
Diperbarui hari ini
Risiko
Perlu ditinjau
Low GitHub adoption signal
Kualitas GitHub
47
64/100 Kualitas · 79/100 Kepercayaan
Tag cakupan
Catatan ulasan
Low GitHub adoption signal · Quality score needs review
Kartu adopsi Agent
Kepercayaan, audit, dan kesiapan pemasangan dalam sekali lihat
Skor ini menggabungkan metadata repositori publik, sinyal ulasan OpenAgentSkill, kebaruan pemeliharaan, dan kesiapan pemasangan. Ini adalah sinyal shortlist, bukan pengganti peninjauan manusia.
Kualitas
MenjanjikanUseful candidate, but compare it with alternatives before adopting.
Kepercayaan
Hanya sandboxKandidat berguna dengan sinyal kepercayaan yang kurang atau bercampur. Gunakan di ruang kerja terisolasi hingga loop hasil membuktikan kecocokan tugas.
Audit
Perlu ditinjauTinjauan yang dapat dibaca mesin tentang kesiapan pemasangan, metadata keamanan, pemeliharaan, dan risiko adopsi.
Trust Score OpenAgentSkill v5
Tinjauan manusia sebelum pemasangan
Jalankan hanya dalam sandbox dan bandingkan alternatif terdekat sebelum digunakan untuk kerja nyata.
Star
47 star GitHub
Aktivitas repositori
47 star dan 0 fork
Pemeliharaan
Diperbarui hari ini
Lisensi
MIT
Pasang
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Keamanan pemasangan
Jalur pemasangan paket atau runtime standar
Cakupan izin
Akses sistem file atau dokumen
Hasil Agent
Belum ada data hasil Agent
Dokumentasi
Konteks README/SKILL.md kuat
Ringkasan risiko
Tinjau sebelum produksi
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
Kesiapan pemasangan
Jalur pemasangan tersedia
- Jalur pemasangan tersedia
- Bukti repositori tersedia
- Lisensi dinyatakan
- Belum ada bukti hasil Agent-Proven
Metadata yang dapat dibaca Agent
Data keputusan yang dapat dibaca mesin untuk skill ini.
Gunakan blok ini atau JSON tersemat untuk memutuskan apakah Agent perlu memasang skill ini, memilih alternatif, atau meminta tinjauan manusia terlebih dahulu.
View technical data+
Metadata yang dapat dibaca Agent
Data keputusan yang dapat dibaca mesin untuk skill ini.
Gunakan blok ini atau JSON tersemat untuk memutuskan apakah Agent perlu memasang skill ini, memilih alternatif, atau meminta tinjauan manusia terlebih dahulu.
Tugas yang sesuai
- alur kerja Document processing
- Tim Claude Code
- builders willing to evaluate younger projects
- Read uploaded files
Agent yang sesuai
Keputusan pemasangan
- Perintah
- npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
- Kebijakan
- Tinjau
- Tinjauan manusia
- Ya
Kepercayaan dan risiko
- Kepercayaan
- 71/100
- Audit
- 80/100
- Tingkat risiko
- Perlu ditinjau
Lingkar hasil
- Endpoint
- /api/agent/outcome
- ID event
- resolve
- Hasil
- 5
Perintah pemasangan
npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-checkJangan gunakan ketika
- Tim yang membutuhkan SLA dengan dukungan vendor
- production agents without a repository review
- Low GitHub adoption signal
- No OpenAgentSkill engagement data yet
- Quality score needs review
Skill alternatif
Frontend Design
171.2K Star
npx skills add anthropics/skills --skill frontend-design
Skill alternatif
Taste Skill: Anti-Slop Frontend
79.7K Star
npx skills add Leonxlnx/taste-skill --skill design-taste-frontend
Skill alternatif
Canvas Design
171.2K Star
npx skills add anthropics/skills --skill canvas-design
Skill alternatif
Anthropic Brand Guidelines
171.2K Star
npx skills add anthropics/skills --skill brand-guidelines
Keamanan Agent v2
64/100 · Tinjau sebelum memasang
Kandidat yang dapat digunakan, tetapi Agent harus menampilkan catatan izin dan audit sebelum memasang.
Memerlukan persetujuan manusia sebelum memasang ke workspace nyata.
Sedang
Akses jaringan
Skill kemungkinan mengambil halaman jarak jauh, API, repositori, atau layanan eksternal.
Sedang
Akses sistem file
Skill dapat membaca atau menulis file proyek, dokumen, artefak yang dihasilkan, atau status workspace lokal.
- Low GitHub adoption signal
Rencana resolusi Agent
Biarkan Agent memverifikasi kecocokan sebelum memasang.
API Resolve mengembalikan skill utama, alternatif, kebijakan keamanan, catatan audit, target pemasangan, dan prompt siap pakai.
Buka JSON
/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Teks Resolve
/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Serah-terima pemasangan
/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
Agent harus memeriksa
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Salin prompt
Task: Use llm-gold-bound-failure-check in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20llm-gold-bound-failure-check%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-check
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Serah-terima Agent
Berikan jalur pemasangan kepada Agent, bukan direktori lain.
Gunakan endpoint publik untuk mengambil perintah, checklist keamanan, prompt target, dan tautan kanonis.
Serah-terima pemasangan
/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install
Format teks LLM
/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install?format=text
Cari alternatif
/api/skills/search?q=llm-gold-bound-failure-check&limit=3
Prompt Agent
Use llm-gold-bound-failure-check for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-llm-gold-bound-failure-check/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-checkMetadata Registry
Profil yang dapat dibaca Agent untuk pemilihan skill otomatis.
API Registry menyediakan sinyal keputusan, kepercayaan, audit, use case, dan pemasangan tanpa mengikis UI.
Manifest
/api/registry/manifest/kennethkhoocy-llm-gold-bound-failure-check
Teks LLM
/api/registry/manifest/kennethkhoocy-llm-gold-bound-failure-check?format=text
Alias pemasangan
/api/registry/install/kennethkhoocy-llm-gold-bound-failure-check
Rekomendasikan
/api/registry/recommend?task=Use%20llm-gold-bound-failure-check%20in%20an%20agent%20workflow&limit=3
Kecocokan Agent
Document processing
Tag use case
Platform
Claude Code
Laporan audit
Perlu ditinjau · 80/100
Tinjauan yang dapat dibaca mesin tentang kesiapan pemasangan, metadata keamanan, pemeliharaan, dan risiko adopsi.
Panel keputusan Agent
Fallback candidate for Document processing
Prototype with this skill first; keep a fallback candidate ready.
Peran di stack
Kandidat cadangan
Kecocokan utama
Document processing
Label kepercayaan
Buat prototipe dulu
Jalur pemasangan
Perintah siap
Gunakan saat
- alur kerja Document processing
- Tim Claude Code
- builders willing to evaluate younger projects
Bukti
- recent repository activity
- install command or GitHub repo available
- profil kualitas 64/100
tinjau dulu
- Low GitHub adoption signal
- No OpenAgentSkill engagement data yet
Jalur implementasi
- 1Pasang di Agent sandbox dan jalankan satu tugas Document processing dari awal hingga akhir.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Profil kepercayaan
Hanya sandbox
Kandidat berguna dengan sinyal kepercayaan yang kurang atau bercampur. Gunakan di ruang kerja terisolasi hingga loop hasil membuktikan kecocokan tugas.
Adopsi GitHub
Periksa47 star GitHub
Aktivitas star/fork
Periksa47 star dan 0 fork; aktivitas issue tidak tersedia dalam metadata saat ini
Pemeliharaan terbaru
LulusDiperbarui hari ini
Kejelasan lisensi
LulusMIT
Sinyal positif
- Tinjauan AI disetujui
- Jalur pemasangan tersedia
- Bukti repositori tersedia
- Repositori yang baru dipelihara
- Perintah pemasangan tidak memiliki pola berisiko tinggi yang jelas
- Loop hasil siap tetapi membutuhkan eksekusi Agent nyata pertama
Tinjau sebelum memasang
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
- Belum ada laporan hasil Agent nyata
- Tinjauan manusia diperlukan sebelum pemasangan tanpa pengawasan
Tindakan yang disarankan
Jalankan hanya dalam sandbox dan bandingkan alternatif terdekat sebelum digunakan untuk kerja nyata.
Profil kualitas
Menjanjikan kandidat untuk alur kerja Agent
Useful candidate, but compare it with alternatives before adopting.
Kecocokan alur kerja
Gunakan skill ini pada skenario berikut
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Kecocokan alur kerja
Tambahkan ke alur kerja lengkap
Design, build, test, and ship interfaces
Frontend and UI
A practical workflow for agents that turn product briefs or Figma designs into polished frontend code, review the result, test it in a browser, and prepare a safe deployment.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Daftar alternatif
Bandingkan sebelum memasang
Similar skills that may fit this task.
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
Taste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
Canvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
Anthropic Brand Guidelines
Apply Anthropic official brand colors, typography, and visual standards to appropriate Anthropic-related artifacts.
Ringkasan
--- name: llm-gold-bound-failure-check description: | Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on prompt revision or model changes. Use when: (1) a scoring pipeline over-predicts a label (precision low, recall high) and a prompt clarification is proposed to tighten it, (2) a pilot/validation gate fails and the fix candidates are prompt edits, (3) inter-rater agreement on the weak label was already low (κ < ~0.6). Core check: if gold POSITIVES share the exact feature the revision would exclude, no prompt can pass a gold-scored gate — recall craters while precision barely moves. Also documents the verified surgical-pilot design (single-section diff, tune/holdout split, pre-registered gate, perturbation check on untouched sections). author: Claude Code version: 1.0.0 date: 2026-07-16 ---
# LLM Gold-Bound Failure Check
## Problem
When an LLM scoring pipeline over-predicts one label, the reflex fix is a prompt clarification ("score positive ONLY when..."). But if the gold standard itself does not separate the texts you want excluded from the texts it labels positive, the revision removes true and false positives together. The pilot fails, the spend is wasted, and — worse — an un-gated adoption would have silently destroyed recall in production.
## Context / Trigger Conditions
- A domain/label shows precision ≪ recall (e.g. P 0.46 / R 0.96) against gold - A prompt edit is proposed to exclude a specific text type (boilerplate, affirmative-program language, non-risk framing) - The label's gold council/inter-rater agreement was already the weakest (κ below ~0.6 is the warning sign that the construct is contested)
## Solution
**Step 0 — the ~$0 check, BEFORE building anything:** read a sample of gold POSITIVES for the weak label and ask: do they contain the feature the revision would exclude? Compare them side-by-side with the false positives.
- Gold positives and false positives are the same kind of text → the failure is **gold-bound**. Stop. No prompt passes a gold-scored gate. The levers are: (a) re-adjudicate the construct with the gold's owners (changes the gold, not the scores), or (b) re-interpret the shipped measure honestly (e.g. "discussion salience" instead of "risk exposure") in downstream analyses. - Gold positives clearly differ from the false positives → a prompt revision is plausible; proceed to a gated pilot.
**Gated pilot design (verified):** 1. Split gold into tune/holdout halves, stratified on the weak label's positives; fixed seed. 2. Draft ONE surgical edit from tune-half errors only — byte-identical elsewhere; verify the diff reverses cleanly. 3. Pre-register the gate on the holdout BEFORE scoring: target-label thresholds (e.g. precision ≥ X AND recall ≥ Y) plus a perturbation tolerance for untouched labels (e.g. within 0.03 F1 / 0.06 κ of a same-serving-rev fresh baseline). 4. Score everything fresh under both prompts (same model revision, same day — this doubles as the drift control). Never write through the production cache layer. 5. Adopt only on a full pass; a REJECT is a valid, cheap outcome.
## Verification
The pilot report shows: the exact prompt diff, tune-vs-holdout metrics for old and new prompts, per-label deltas on untouched sections, and spend. A gold-bound diagnosis is confirmed when the revision moves recall sharply down while precision stays roughly flat.
## Example
Specialist Directors US, 2026-07-16: DEI over-prediction (P 0.46 / R 0.96, council κ 0.24–0.59). A risk-framing-only DEI clause was piloted ($1.17, pre-registered holdout gate). Result: recall 0.895→0.263, precision 0.455 (gate ≥0.60) — REJECT. Reading the tune half showed ~¾ of gold DEI positives were pure affirmative D&I program text, identical in kind to the false positives; the failure was predictable at Step 0. Bonus finding: the DEI-section-only edit left all five other domains within 0.025 F1 / 0.05 κ — single-section prompt edits isolate cleanly, so the perturbation check is a cheap add, not paranoia. Same pattern one week earlier: a cyber classifier pilot gate failure traced to E/D gold contamination (misses were skills-matrix-checkbox-only positives), not model weakness.
## Notes
- Low inter-rater κ on a label is the leading indicator: contested construct → gold-bound failures downstream. - If the pipeline scores all labels in one completion, any post-campaign prompt change forces a full re-score — run this check BEFORE the campaign. - See also: [llm-campaign-drift-gate] for the companion gate on resume boundaries and serving-revision drift (same fresh-baseline discipline). - See also: [annotator-input-parity-check] — run it FIRST. If the model was never shown the document the annotators read, apparent gold-bound failures (e.g. the 2026-07-16 E/D "contamination" reading above) are actually input mismatch: the 2026-07-21 parity audit showed the specialist-director hand labels were pure proxy-statement transcriptions, so checkbox-only positives were recoverable from the right input all along.
Detail teknis
- Versi
- 1.0.0
- Lisensi
- MIT
- Pembaruan terakhir
- 24 Agu 2026
- Diterbitkan
- 24 Agu 2026
Ringkasan keputusan
Kandidat cadangan
recent repository activity
Audit
Tinjauan pemasangan
Tinjauan pemasangan dan adopsi
- Keamanan
- 86/100
- Pemeliharaan
- 100/100
- Pasang
- 92/100
Bukti tervalidasi Agent
Bukti tervalidasi Agent
Laporan hasil setelah resolve, tinjau, pasang, dan satu eksekusi terbatas.
- Tingkat sukses
- —
- Kegagalan terbaru
- —
- Hasil
- 0
- Kualitas output
- —
- Gagal
- 0
- Tidak relevan
- 0
- Pemasangan
- 0
- Diblokir risiko
- 0
- Perlu penyiapan
- 0
- Produksi
- 0
Belum ada data hasil Agent. Eksekusi pertama dapat melaporkan keberhasilan, kebutuhan setup, blok risiko, kegagalan, atau tidak relevan melalui /api/agent/outcome.
Pasang
Tambahkan ke alur Agent
Gratis dan sumber terbuka. Tinjau laporan sebelum memasang pada Agent produksi.
Siklus pertumbuhan
Kit berbagi
Draf berbasis skenario untuk llm-gold-bound-failure-check, siap untuk posting manual di X.
llm-gold-bound-failure-check: Diagnose whether an LLM classifier's validation-gate failure is GOLD-BOUND before spending on... 47 stars https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x
Balasan opsional dengan perintah pemasangan
Listing + install path for llm-gold-bound-failure-check: https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check?ref=x Install: npx skills add kennethkhoocy/applied-micro-skills --skill llm-gold-bound-failure-...
Sumber listing
Diindeks Registry
Listing ini diindeks dari sumber publik dan belum ditandai resmi hingga klaim pemelihara disetujui.
- Kreator
- Claude Code
- Diindeks oleh
- Indeks komunitas OpenAgentSkill
Atribusi menautkan ke repositori publik atau profil kreator. Kreator dapat mengklaim listing untuk memperbarui sinyal kepemilikan.
Klaim skill iniKlaim pemilik
Klaim listing skill ini
Listing Diindeks Registry ini dikaitkan dengan Claude Code, tetapi belum ditandai resmi. Klaim untuk menambahkan sinyal pemilik terverifikasi dan membuat pembaruan peluncuran, pemasangan, serta audit berikutnya lebih tepercaya.
Kit backlink kreator
Tambahkan badge bukti ke README Anda
Tampilkan listing kanonis, sinyal kepercayaan dan audit saat ini, serta bukti Agent-Proven nyata di tempat pengembang mengevaluasi repositori.
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check/audit)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-gold-bound-failure-check)Penulis
Claude Code
@claude-code
Tag
Kecocokan platform
Sinyal kesehatan
- Star GitHub
- 47
- Skor kualitas
- 35/100
- Push GitHub terakhir
- 24 Agu 2026
- Petunjuk framework
- Tidak diketahui
- Tampilan OpenAgentSkill
- 0
- Salinan pemasangan
- 0
- Klik keluar
- 0
Sinyal komunitas
Bagikan apakah skill ini bermanfaat untuk alur kerja Agent Anda. Masukan gabungan meningkatkan peringkat dari waktu ke waktu.
Kepercayaan & keamanan
Hanya sandbox
- Adopsi GitHub47 star GitHubPeriksa
- Aktivitas star/fork47 star dan 0 fork; aktivitas issue tidak tersedia dalam metadata saat iniPeriksa
- Pemeliharaan terbaruDiperbarui hari iniLulus
- Kejelasan lisensiMITLulus
- Kelengkapan README/SKILL.mdMetadata memuat konteks penggunaan dan alur kerja yang cukupLulus
- Risiko dependensi/runtimeTidak ada petunjuk risiko dependensi besar dalam metadata publikLulus
Skill terkait
Frontend Design
Guidance for distinctive, intentional UI design, typography, visual direction, and non-template-like product interfaces.
171.2K StarTaste Skill: Anti-Slop Frontend
Design and implementation guidance for distinctive landing pages, portfolios, product demos, and purposeful redesigns.
79.7K StarCanvas Design
Create original visual art, posters, PNG assets, and PDF documents through a clear design philosophy.
171.2K StarAnthropic Brand Guidelines
Apply Anthropic official brand colors, typography, and visual standards to appropriate Anthropic-related artifacts.
171.2K Star