llm-campaign-drift-gate
Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alias (deepseek-chat, gpt-*-latest, gemini-*-preview, any provider alias without a pinned version). Use when: (1) resuming a paused or credit-exhausted scoring run days after its last chunk,
Supply asset profile
Coding and developer agents
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
Maintenance
fresh
Pushed today
Risk
Needs review
Low GitHub adoption signal
GitHub quality
47
64/100 Quality · 79/100 Trust
Coverage tags
Review notes
Low GitHub adoption signal · Quality score needs review
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
47 GitHub stars
Repo activity
47 stars, 0 forks
Maintenance
Pushed today
License
MIT
Install
npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
Install safety
standard package or runtime install path
Permission surface
no high-risk permission surface in public metadata
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Review before production
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- GitHub automation workflows
- Claude Code teams
- builders willing to evaluate younger projects
- Inspect repository metadata
Suited agents
Install decision
- Command
- npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
- Policy
- review
- Human review
- yes
Trust and risk
- Trust
- 71/100
- Audit
- 81/100
- Risk level
- Needs review
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gateDo not use when
- teams that need a vendor-supported SLA
- production agents without a repository review
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
Alternative
Opencode
200.7K Stars
npx skills add anomalyco/opencode
Alternative
Code Review
168.6K Stars
npx skills add mattpocock/skills --skill code-review
Alternative
Grill With Docs
164.7K Stars
npx skills add mattpocock/skills --skill grill-with-docs
Alternative
To Spec
164.7K Stars
npx skills add mattpocock/skills --skill to-spec
Agent safety v2
69/100 · Review before install
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
- Low GitHub adoption signal
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install kennethkhoocy-llm-campaign-drift-gateAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20llm-campaign-drift-gate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20llm-campaign-drift-gate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/kennethkhoocy-llm-campaign-drift-gate/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use llm-campaign-drift-gate in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20llm-campaign-drift-gate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-llm-campaign-drift-gate/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/kennethkhoocy-llm-campaign-drift-gate/install
LLM text format
/api/skills/kennethkhoocy-llm-campaign-drift-gate/install?format=text
Find alternatives
/api/skills/search?q=llm-campaign-drift-gate&limit=3
Agent prompt
Use llm-campaign-drift-gate for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-llm-campaign-drift-gate/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gateRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/kennethkhoocy-llm-campaign-drift-gate
LLM text
/api/registry/manifest/kennethkhoocy-llm-campaign-drift-gate?format=text
Install alias
/api/registry/install/kennethkhoocy-llm-campaign-drift-gate
Recommend
/api/registry/recommend?task=Use%20llm-campaign-drift-gate%20in%20an%20agent%20workflow&limit=3
Agent fit
GitHub automation
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
Needs review · 81/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Fallback candidate for GitHub automation
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
GitHub automation
Trust label
Prototype first
Install path
Command ready
Use when
- GitHub automation workflows
- Claude Code teams
- builders willing to evaluate younger projects
Evidence
- recent repository activity
- install command or GitHub repo available
- 64/100 quality profile
- 2 OpenAgentSkill engagement events
review first
- Low GitHub adoption signal
Implementation path
- 1Install it in a sandbox agent and run one GitHub automation task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Sandbox only
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
CHECK47 GitHub stars
Stars/forks activity
CHECK47 stars, 0 forks; issue activity unavailable in current metadata
Recent maintenance
PASSPushed today
License clarity
PASSMIT
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Promising candidate for agent workflows
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Use this skill in these scenarios
Manage repositories
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Build and ship code
Coding agents
I need a coding agent that can understand a repository, edit code, and review pull requests.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
Workflow fit
Add it to a complete workflow
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Scrape, clean, and reuse web data
Web data pipeline
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
Opencode
The open source coding agent.
Code Review
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
Grill With Docs
A relentless interview that pressure-tests a plan against the codebase, sharpens domain language, and updates CONTEXT.md and ADRs when decisions become durable.
To Spec
Turn the current conversation and codebase context into a structured implementation spec, then publish it to the configured project issue tracker.
Overview
--- name: llm-campaign-drift-gate description: | Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alias (deepseek-chat, gpt-*-latest, gemini-*-preview, any provider alias without a pinned version). Use when: (1) resuming a paused or credit-exhausted scoring run days after its last chunk, (2) topping up credits to finish a campaign, (3) extending a cached scoring pipeline with new items. Prevents silently splicing two model versions or serving revisions into one measure. Verified 2026-07-16: for $0.30 caught a serving-revision drift WITHIN DeepSeek v4-flash (same alias, same family, litigation scores systematically shifted across a 2-day gap) before an $83 resume spend. author: Claude Code version: 1.1.0 date: 2026-07-16 ---
# LLM Campaign Drift Gate
## Problem
Batch-scoring campaigns (exposure measures, classifiers, extraction runs) call provider aliases that can be silently repointed to a new model at any time. Resuming a half-finished campaign after the alias moves splices two different scorers into one variable, with the version boundary correlated with whatever orders the chunks (time, firm id) — a silent confound. Providers can also RETIRE the old model entirely, making the original campaign uncompletable.
## Context / Trigger Conditions
- Resuming a scoring run more than ~a day after its last paid chunk - "Top up credits and finish the run" requests - Any incremental scoring against an existing response cache - Symptom of a missed gate: a step-change in scores at a resume boundary
## Solution
Before ANY production spend on resume, run a two-part gate (~$0.30–2):
1. **Canary (the decisive check):** sample ~100 already-cached items, re-send their EXACT stored prompts fresh, compare fresh vs cached scores. Gate: ≥97% all-field exact match and no systematic directional shift. Write the comparison in a standalone script — never through the pipeline's cache layer, which would overwrite production entries. 2. **Gold re-validation:** re-score the gold/validation panel fresh and compare agreement metrics to the prior validation (e.g. median F1/κ within ~0.03, no domain dropping >0.10).
Also capture `response.model` on every gate call — pipelines rarely store it, and it is the only direct evidence of a repoint. Check the provider's `/models` endpoint: if the old model id is gone, no rollback exists.
3. **If the canary fails, diagnose BEFORE concluding — two mandatory follow-ups:** - **Date the suspected flip against the provider's changelog** before inferring a model splice. `response.model` on fresh calls identifies today's model only; if the alias already pointed there when the cache was written, there is no family splice and the mismatch needs another explanation. (Verified failure mode: an alias that had served the "new" model for months was misread as a fresh repoint.) - **Fresh-vs-fresh canary** to separate serving drift from temperature-0 nondeterminism: re-score the same items a second time. Drift signature = fresh2-vs-fresh1 agreement high and symmetric while both fresh runs disagree with the cache at a higher rate in the SAME signed direction. Noise signature = fresh-vs-fresh disagrees about as much as fresh-vs-cache, with no directional bias. - Supporting forensic: compare raw-response formatting fingerprints (JSON pretty/compact ratio, key order) between cache and fresh — a heterogeneous or shifted style distribution corroborates a serving change when no model id was recorded.
**Key subtlety (why both checks):** a new model or revision can validate AGAINST GOLD as well as the old one (κ holds or improves) while still disagreeing with the old scores on 10–30% of items, concentrated in borderline-heavy fields. Gold agreement does not license splicing — the gate fails on the canary alone. And alias stability is not serving stability: the same alias serving the same model family can still drift across days via silent serving revisions; a canary-failed resume is a seam either way, and the decision (resume with a documented seam vs re-score the universe) belongs to the budget owner.
## Verification
The gate script logs: fresh `response.model` ids, canary exact-match rate, per-field mismatch counts with signed direction, and the gold-metric deltas. GO only if both checks pass.
## Example
T1 exposure_v2 resume, 2026-07-16: canary returned 71% exact (gate ≥97%) with a litigation-concentrated negative shift, yet holdout median κ improved 0.607→0.644. First interpretation — "alias repointed to a new model family" — was WRONG: the provider changelog showed `deepseek-chat` had served v4-flash since April, months before the campaign. The fresh-vs-fresh follow-up then isolated the true cause: fresh2-vs-fresh1 93% exact/symmetric/litigation 0, both fresh runs vs cache 71–72% with litigation −12 identically — a serving revision within the same model across a 2-day gap, corroborated by a shifted JSON-formatting fingerprint. Total diagnosis cost ~$0.30; the resume-vs-rescore decision went to the budget owner with the seam quantified.
## Notes
- Design campaigns for this failure: per-response content-addressed cache + append-only checkpoint makes "re-score everything under the new model" a clean cache-rotation, not a data loss. - If the cache key embeds the alias string rather than the resolved model, record actual `response.model` in run reports — the cache cannot tell you later which model produced an entry. - One campaign = one model. Budget and schedule so the universe completes within days, or accept that a provider release can force a full re-score. - See also: [llm-gold-bound-failure-check] for the companion pre-campaign check — whether a validation-gate failure is fixable by prompt at all, or bound to the gold construct.
Technical details
- Version
- 1.1.0
- License
- MIT
- Last updated
- Aug 24, 2026
- Published
- Aug 24, 2026
Decision snapshot
Fallback candidate
recent repository activity
Audit
Install review
Install and adoption review
- Security
- 87/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for llm-campaign-drift-gate, ready for a manual X post.
llm-campaign-drift-gate: Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alia... 47 stars https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate?ref=x
Optional reply with install command
Listing + install path for llm-campaign-drift-gate: https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate?ref=x Install: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
Listing source
Registry indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- Claude Code
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Registry indexed listing is attributed to Claude Code but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate/audit)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)Author
Claude Code
@claude-code
Tags
Platform fit
Health signals
- GitHub stars
- 47
- Quality score
- 35/100
- Last GitHub push
- Aug 24, 2026
- Framework hints
- Unknown
- OpenAgentSkill views
- 2
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Sandbox only
- GitHub adoption47 GitHub starsCHECK
- Stars/forks activity47 stars, 0 forks; issue activity unavailable in current metadataCHECK
- Recent maintenancePushed todayPASS
- License clarityMITPASS
- README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
- Dependency/runtime riskno major dependency risk hints in public metadataPASS
Related skills
Opencode
The open source coding agent.
200.7K StarsCode Review
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
168.6K StarsGrill With Docs
A relentless interview that pressure-tests a plan against the codebase, sharpens domain language, and updates CONTEXT.md and ADRs when decisions become durable.
164.7K StarsTo Spec
Turn the current conversation and codebase context into a structured implementation spec, then publish it to the configured project issue tracker.
164.7K Stars