llm-campaign-drift-gate

REVIEW · 71
Registry indexed

Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alias (deepseek-chat, gpt-*-latest, gemini-*-preview, any provider alias without a pinned version). Use when: (1) resuming a paused or credit-exhausted scoring run days after its last chunk,

Verified installs0
Stars47
Version1.1.0
Quality64/100 · Promising
Trust71/100 · Sandbox only
Audit81/100 · Needs review

Supply asset profile

Coding and developer agents

Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.

Browse track

Scenario

GitHub automation

I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.

Agent fit

Claude Code + OpenAI Agents + CLI

Codex, Claude Code, Cursor, CLI, or custom agents.

Install

Ready

npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate

Maintenance

fresh

Pushed today

Risk

Needs review

Low GitHub adoption signal

GitHub quality

47

64/100 Quality · 79/100 Trust

Coverage tags

CodingGitHub automationcoding-agentsagent-skill

Review notes

Low GitHub adoption signal · Quality score needs review

Agent adoption scorecard

Trust, audit, and install readiness at a glance

These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.

Quality

Promising
64

Useful candidate, but compare it with alternatives before adopting.

Trust

Sandbox only
71

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

Audit

Needs review
81

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

OpenAgentSkill Trust Score v5

Human review before install

Run only in a sandbox and compare close alternatives before using it for real work.

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

47 GitHub stars

Repo activity

47 stars, 0 forks

Maintenance

Pushed today

License

MIT

Install

npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate

Install safety

standard package or runtime install path

Permission surface

no high-risk permission surface in public metadata

Agent outcomes

No agent outcome data yet

Docs

Strong README/SKILL.md context

Risk summary

Review before production

  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata

Install readiness

Install path available

  • Install path is available
  • Repository evidence is available
  • License is declared
  • No Agent Proven outcome evidence yet

Agent-readable metadata

Machine-readable decision data for this skill.

Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.

Open JSON

Suited tasks

  • GitHub automation workflows
  • Claude Code teams
  • builders willing to evaluate younger projects
  • Inspect repository metadata

Suited agents

CodexClaude CodeCursorOpenAgentSkill CLIOpenAI AgentsCLI

Install decision

Command
npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
Policy
review
Human review
yes

Trust and risk

Trust
71/100
Audit
81/100
Risk level
Needs review

Outcome loop

Endpoint
/api/agent/outcome
Event ID
resolve
Outcomes
5

Install command

npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate

Do not use when

  • teams that need a vendor-supported SLA
  • production agents without a repository review
  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars

Agent safety v2

69/100 · Review before install

Reviewed with permission notesreview

Usable candidate, but the agent should surface permission and audit notes before installation.

Require human approval before installing into a real workspace.

Resolve via API

medium

Network access

Skill likely fetches remote pages, APIs, repositories, or external services.

  • Low GitHub adoption signal

Install targets

Install this skill in your agent workflow

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install kennethkhoocy-llm-campaign-drift-gate

Agent resolve plan

Let an agent verify fit before installing.

The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.

Open text plan

Agent should check

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Copy prompt

Task: Use llm-campaign-drift-gate in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20llm-campaign-drift-gate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-llm-campaign-drift-gate/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent handoff

Give an agent the install path, not another directory page.

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

Open install API

Agent prompt

Use llm-campaign-drift-gate for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-llm-campaign-drift-gate/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate

Registry metadata

Agent-readable profile for automatic skill selection.

This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.

Open manifest

Agent fit

64/100

GitHub automation

Platforms

Claude Code, OpenAI Agents

Audit report

Needs review · 81/100

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

View audit reportView eval report

Agent decision cockpit

Fallback candidate for GitHub automation

Prototype with this skill first; keep a fallback candidate ready.

64
Readiness
Prototype
Stage

Role in stack

Fallback candidate

Primary fit

GitHub automation

Trust label

Prototype first

Install path

Command ready

Use when

  • GitHub automation workflows
  • Claude Code teams
  • builders willing to evaluate younger projects

Evidence

  • recent repository activity
  • install command or GitHub repo available
  • 64/100 quality profile
  • 2 OpenAgentSkill engagement events

review first

  • Low GitHub adoption signal

Implementation path

  1. 1Install it in a sandbox agent and run one GitHub automation task end to end.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Trust profile

Sandbox only

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

71
OpenAgentSkill Trust Score

GitHub adoption

CHECK

47 GitHub stars

Stars/forks activity

CHECK

47 stars, 0 forks; issue activity unavailable in current metadata

Recent maintenance

PASS

Pushed today

License clarity

PASS

MIT

Good signals

  • AI review approved
  • Install path is available
  • Repository evidence is available
  • Recently maintained repository
  • Install command has no obvious high-risk pattern
  • Outcome loop is ready but needs first real agent run

Review before install

  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 47 GitHub stars
  • Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
  • No real agent outcome reports yet
  • Human review required before unattended installation

Recommended action

Run only in a sandbox and compare close alternatives before using it for real work.

Quality profile

Promising candidate for agent workflows

Useful candidate, but compare it with alternatives before adopting.

64
GitHub stars
47
Freshness
Today
Install ready
Yes
License
MIT
Review before install: Low GitHub adoption signal

Workflow fit

Use this skill in these scenarios

Workflow fit

Add it to a complete workflow

Alternative shortlist

Compare before you install

Similar skills that may fit this task.

Compare all

Overview

--- name: llm-campaign-drift-gate description: | Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alias (deepseek-chat, gpt-*-latest, gemini-*-preview, any provider alias without a pinned version). Use when: (1) resuming a paused or credit-exhausted scoring run days after its last chunk, (2) topping up credits to finish a campaign, (3) extending a cached scoring pipeline with new items. Prevents silently splicing two model versions or serving revisions into one measure. Verified 2026-07-16: for $0.30 caught a serving-revision drift WITHIN DeepSeek v4-flash (same alias, same family, litigation scores systematically shifted across a 2-day gap) before an $83 resume spend. author: Claude Code version: 1.1.0 date: 2026-07-16 ---

# LLM Campaign Drift Gate

## Problem

Batch-scoring campaigns (exposure measures, classifiers, extraction runs) call provider aliases that can be silently repointed to a new model at any time. Resuming a half-finished campaign after the alias moves splices two different scorers into one variable, with the version boundary correlated with whatever orders the chunks (time, firm id) — a silent confound. Providers can also RETIRE the old model entirely, making the original campaign uncompletable.

## Context / Trigger Conditions

- Resuming a scoring run more than ~a day after its last paid chunk - "Top up credits and finish the run" requests - Any incremental scoring against an existing response cache - Symptom of a missed gate: a step-change in scores at a resume boundary

## Solution

Before ANY production spend on resume, run a two-part gate (~$0.30–2):

1. **Canary (the decisive check):** sample ~100 already-cached items, re-send their EXACT stored prompts fresh, compare fresh vs cached scores. Gate: ≥97% all-field exact match and no systematic directional shift. Write the comparison in a standalone script — never through the pipeline's cache layer, which would overwrite production entries. 2. **Gold re-validation:** re-score the gold/validation panel fresh and compare agreement metrics to the prior validation (e.g. median F1/κ within ~0.03, no domain dropping >0.10).

Also capture `response.model` on every gate call — pipelines rarely store it, and it is the only direct evidence of a repoint. Check the provider's `/models` endpoint: if the old model id is gone, no rollback exists.

3. **If the canary fails, diagnose BEFORE concluding — two mandatory follow-ups:** - **Date the suspected flip against the provider's changelog** before inferring a model splice. `response.model` on fresh calls identifies today's model only; if the alias already pointed there when the cache was written, there is no family splice and the mismatch needs another explanation. (Verified failure mode: an alias that had served the "new" model for months was misread as a fresh repoint.) - **Fresh-vs-fresh canary** to separate serving drift from temperature-0 nondeterminism: re-score the same items a second time. Drift signature = fresh2-vs-fresh1 agreement high and symmetric while both fresh runs disagree with the cache at a higher rate in the SAME signed direction. Noise signature = fresh-vs-fresh disagrees about as much as fresh-vs-cache, with no directional bias. - Supporting forensic: compare raw-response formatting fingerprints (JSON pretty/compact ratio, key order) between cache and fresh — a heterogeneous or shifted style distribution corroborates a serving change when no model id was recorded.

**Key subtlety (why both checks):** a new model or revision can validate AGAINST GOLD as well as the old one (κ holds or improves) while still disagreeing with the old scores on 10–30% of items, concentrated in borderline-heavy fields. Gold agreement does not license splicing — the gate fails on the canary alone. And alias stability is not serving stability: the same alias serving the same model family can still drift across days via silent serving revisions; a canary-failed resume is a seam either way, and the decision (resume with a documented seam vs re-score the universe) belongs to the budget owner.

## Verification

The gate script logs: fresh `response.model` ids, canary exact-match rate, per-field mismatch counts with signed direction, and the gold-metric deltas. GO only if both checks pass.

## Example

T1 exposure_v2 resume, 2026-07-16: canary returned 71% exact (gate ≥97%) with a litigation-concentrated negative shift, yet holdout median κ improved 0.607→0.644. First interpretation — "alias repointed to a new model family" — was WRONG: the provider changelog showed `deepseek-chat` had served v4-flash since April, months before the campaign. The fresh-vs-fresh follow-up then isolated the true cause: fresh2-vs-fresh1 93% exact/symmetric/litigation 0, both fresh runs vs cache 71–72% with litigation −12 identically — a serving revision within the same model across a 2-day gap, corroborated by a shifted JSON-formatting fingerprint. Total diagnosis cost ~$0.30; the resume-vs-rescore decision went to the budget owner with the seam quantified.

## Notes

- Design campaigns for this failure: per-response content-addressed cache + append-only checkpoint makes "re-score everything under the new model" a clean cache-rotation, not a data loss. - If the cache key embeds the alias string rather than the resolved model, record actual `response.model` in run reports — the cache cannot tell you later which model produced an entry. - One campaign = one model. Budget and schedule so the universe completes within days, or accept that a provider release can force a full re-score. - See also: [llm-gold-bound-failure-check] for the companion pre-campaign check — whether a validation-gate failure is fixable by prompt at all, or bound to the gold construct.

Technical details

Version
1.1.0
License
MIT
Last updated
Aug 24, 2026
Published
Aug 24, 2026

Decision snapshot

Fallback candidate

64
Ready
Prototype
Stage

recent repository activity

Audit

Install review

Install and adoption review

81
Needs review
Security
87/100
Maintenance
100/100
Install
92/100
Open full auditView eval report

Agent-proven evidence

Agent-proven evidence

Outcome reports after resolve, review, install, and one narrow run.

0
Proven
Needs first agent runAuto-install: review firstLast: Unknown
Success rate
Recent failure
Outcomes
0
Output quality
Failed
0
Not relevant
0
Installs
0
Risk blocked
0
Setup needed
0
Production
0

No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.

Install

Add to agent workflow

Free and open source. Review the report before installing into production agents.

Growth loop

Share kit

X

Scenario-led draft for llm-campaign-drift-gate, ready for a manual X post.

Curator note
llm-campaign-drift-gate: Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alia...

47 stars

https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate?ref=x
Open X draft
Optional reply with install command
Listing + install path for llm-campaign-drift-gate:
https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate?ref=x

Install: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate

Listing source

Registry indexed

Claimable

This listing was indexed from public sources and is not marked official until a maintainer claim is approved.

Indexed by
OpenAgentSkill community index

Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.

Claim this skill

Owner claim

Claim this skill listing

This Registry indexed listing is attributed to Claude Code but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.

Creator backlink kit

Add the evidence badges to your README

Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-campaign-drift-gate?metric=listed&label=Listed)](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-campaign-drift-gate?metric=trust&label=Trust)](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-campaign-drift-gate?metric=audit&label=Audit)](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/kennethkhoocy-llm-campaign-drift-gate?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)

Author

C

Claude Code

@claude-code

Health signals

GitHub stars
47
Quality score
35/100
Last GitHub push
Aug 24, 2026
Framework hints
Unknown
OpenAgentSkill views
2
Install copies
0
Outbound clicks
0

Community signal

Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.

Trust & safety

Sandbox only

71
  • GitHub adoption47 GitHub starsCHECK
  • Stars/forks activity47 stars, 0 forks; issue activity unavailable in current metadataCHECK
  • Recent maintenancePushed todayPASS
  • License clarityMITPASS
  • README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
  • Dependency/runtime riskno major dependency risk hints in public metadataPASS