llm-campaign-drift-gate
Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alias (deepseek-chat, gpt-*-latest, gemini-*-preview, any provider alias without a pinned version). Use when: (1) resuming a paused or credit-exhausted scoring run days after its last chunk,
供给资产档案
编程与开发 Agent
代码审查、仓库分析、测试、CI、GitHub、DevOps 与开发工作流 Skill。
场景
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
适配 Agent
Claude Code + OpenAI Agents + CLI
适用于 Codex、Claude Code、Cursor、CLI 或自定义 Agent。
安装
就绪
npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
维护状态
新鲜
今天有推送
风险
需审查
Low GitHub adoption signal
GitHub 质量
47
64/100 质量 · 79/100 信任
覆盖标签
审查说明
Low GitHub adoption signal · Quality score needs review
Agent 采用评分卡
一眼查看信任、审计与安装准备度
这些分数综合公开仓库元数据、OpenAgentSkill 审查信号、维护新鲜度与安装准备度。它用于候选筛选,不替代人工审查。
质量
有潜力有用的候选项,但采用前应与替代方案比较。
信任
仅限沙盒有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。
审计
需审查对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。
OpenAgentSkill 信任评分 v5
安装前需人工审查
仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。
Stars
47 个 GitHub Stars
仓库活跃度
47 个 Star,0 个 Fork
维护状态
今天有推送
许可证
MIT
安装
npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
安装安全性
标准软件包或运行时安装路径
权限范围
公开元数据中未发现高风险权限范围
Agent 结果
暂未有 Agent 结果数据
文档
README/SKILL.md 上下文充分
风险摘要
生产前审查
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
安装准备度
安装路径可用
- 安装路径可用
- 仓库证据可用
- 已声明许可证
- 暂无 Agent 验证结果证据
Agent 可读元数据
这个 Skill 的机器可读决策数据。
使用此区块或内嵌 JSON 判断 Agent 是否应安装该 Skill、选择替代方案,或先请求人工审查。
适用任务
- GitHub automation 工作流
- Claude Code 团队
- builders willing to evaluate younger projects
- Inspect repository metadata
适用 Agent
安装决策
- 命令
- npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
- 策略
- 审查
- 人工审查
- 是
信任与风险
- 信任
- 71/100
- 审计
- 81/100
- 风险级别
- 需审查
结果闭环
- 端点
- /api/agent/outcome
- 事件 ID
- resolve
- 结果
- 5
安装命令
npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate不适用场景
- 需要厂商支持 SLA 的团队
- production agents without a repository review
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
替代 Skill
Opencode
200.7K Stars
npx skills add anomalyco/opencode
替代 Skill
Code Review
168.6K Stars
npx skills add mattpocock/skills --skill code-review
替代 Skill
Grill With Docs
164.7K Stars
npx skills add mattpocock/skills --skill grill-with-docs
替代 Skill
To Spec
164.7K Stars
npx skills add mattpocock/skills --skill to-spec
Agent 安全 v2
69/100 · 安装前审查
可用候选,但 Agent 在安装前应展示权限与审计说明。
在真实工作区安装前需要人工批准。
中
网络访问
Skill 可能访问远程页面、API、仓库或外部服务。
- Low GitHub adoption signal
安装目标
在你的 Agent 工作流中安装此 Skill
通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install kennethkhoocy-llm-campaign-drift-gateAgent 解析计划
让 Agent 在安装前验证匹配度。
Resolve API 返回首选 Skill、替代方案、安全策略、审计说明、安装目标和可直接执行的提示词,无需抓取此页面。
打开 JSON
/api/agent/resolve?task=Use%20llm-campaign-drift-gate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve 文本
/api/agent/resolve?task=Use%20llm-campaign-drift-gate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
安装交接
/api/skills/kennethkhoocy-llm-campaign-drift-gate/install
Agent 应检查
- 从 Resolve API 检查任务匹配与替代方案。
- 检查审计评分、信任评分和安全策略警告。
- 检查 Codex、Claude Code、Cursor 或 CLI 的安装目标兼容性。
复制提示词
Task: Use llm-campaign-drift-gate in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20llm-campaign-drift-gate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/kennethkhoocy-llm-campaign-drift-gate/install
Install command: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent 交接
把安装路径交给 Agent,而不是再给一个目录页。
通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。
安装交接
/api/skills/kennethkhoocy-llm-campaign-drift-gate/install
LLM 文本格式
/api/skills/kennethkhoocy-llm-campaign-drift-gate/install?format=text
寻找替代方案
/api/skills/search?q=llm-campaign-drift-gate&limit=3
Agent 提示词
Use llm-campaign-drift-gate for this task. Review https://www.openagentskill.com/api/skills/kennethkhoocy-llm-campaign-drift-gate/install, then install with: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gateRegistry 元数据
用于自动选择 Skill 的 Agent 可读档案。
本页通过 Registry API 提供相同的决策、信任、审计、场景和安装信号,让 Agent 无需抓取界面即可排序。
Manifest
/api/registry/manifest/kennethkhoocy-llm-campaign-drift-gate
LLM 文本
/api/registry/manifest/kennethkhoocy-llm-campaign-drift-gate?format=text
安装别名
/api/registry/install/kennethkhoocy-llm-campaign-drift-gate
推荐
/api/registry/recommend?task=Use%20llm-campaign-drift-gate%20in%20an%20agent%20workflow&limit=3
适配 Agent
GitHub automation
平台
Claude Code, OpenAI Agents
Agent 决策面板
Fallback candidate for GitHub automation
先用此 Skill 做原型验证,并保留备选方案。
栈中角色
备选候选
主要匹配
GitHub automation
信任标签
先做原型验证
安装路径
命令已就绪
适用场景
- GitHub automation 工作流
- Claude Code 团队
- builders willing to evaluate younger projects
证据
- 仓库近期活跃
- 已提供安装命令或 GitHub 仓库
- 64/100 质量档案
- 2 个 OpenAgentSkill 交互事件
先审查
- Low GitHub adoption signal
实施路径
- 1在沙盒 Agent 中安装它,并端到端完成一次GitHub automation任务。
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
信任档案
仅限沙盒
有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。
GitHub 采用度
检查47 个 GitHub Stars
Star/Fork 活跃度
检查47 个 Star,0 个 Fork; 当前元数据中没有议题活跃度信息
近期维护
通过今天有推送
许可证清晰度
通过MIT
积极信号
- AI 审查已通过
- 安装路径可用
- 仓库证据可用
- 近期维护的仓库
- 安装命令未发现明显高风险模式
- 结果闭环已就绪,但需要首次真实 Agent 运行
安装前审查
- Low GitHub adoption signal
- Quality score needs review
- GitHub adoption: 47 GitHub stars
- Stars/forks activity: 47 stars, 0 forks; issue activity unavailable in current metadata
- 暂未有真实 Agent 结果报告
- 无人值守安装前需要人工审查
建议操作
仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。
质量档案
有潜力 适用于 Agent 工作流的候选
有用的候选项,但采用前应与替代方案比较。
工作流匹配
在这些场景使用此 Skill
Manage repositories
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Build and ship code
Coding agents
I need a coding agent that can understand a repository, edit code, and review pull requests.
Parse messy files
Document processing
I need my agent to read PDFs, extract tables, and turn documents into structured data.
工作流匹配
加入完整工作流
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Scrape, clean, and reuse web data
Web data pipeline
A practical workflow for agents that crawl public pages, extract clean content, normalize data, and hand it to downstream research or RAG workflows.
替代方案短名单
安装前对比
可能适合该任务的相近 Skill。
Opencode
The open source coding agent.
Code Review
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
Grill With Docs
A relentless interview that pressure-tests a plan against the codebase, sharpens domain language, and updates CONTEXT.md and ADRs when decisions become durable.
To Spec
Turn the current conversation and codebase context into a structured implementation spec, then publish it to the configured project issue tracker.
概览
--- name: llm-campaign-drift-gate description: | Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alias (deepseek-chat, gpt-*-latest, gemini-*-preview, any provider alias without a pinned version). Use when: (1) resuming a paused or credit-exhausted scoring run days after its last chunk, (2) topping up credits to finish a campaign, (3) extending a cached scoring pipeline with new items. Prevents silently splicing two model versions or serving revisions into one measure. Verified 2026-07-16: for $0.30 caught a serving-revision drift WITHIN DeepSeek v4-flash (same alias, same family, litigation scores systematically shifted across a 2-day gap) before an $83 resume spend. author: Claude Code version: 1.1.0 date: 2026-07-16 ---
# LLM Campaign Drift Gate
## Problem
Batch-scoring campaigns (exposure measures, classifiers, extraction runs) call provider aliases that can be silently repointed to a new model at any time. Resuming a half-finished campaign after the alias moves splices two different scorers into one variable, with the version boundary correlated with whatever orders the chunks (time, firm id) — a silent confound. Providers can also RETIRE the old model entirely, making the original campaign uncompletable.
## Context / Trigger Conditions
- Resuming a scoring run more than ~a day after its last paid chunk - "Top up credits and finish the run" requests - Any incremental scoring against an existing response cache - Symptom of a missed gate: a step-change in scores at a resume boundary
## Solution
Before ANY production spend on resume, run a two-part gate (~$0.30–2):
1. **Canary (the decisive check):** sample ~100 already-cached items, re-send their EXACT stored prompts fresh, compare fresh vs cached scores. Gate: ≥97% all-field exact match and no systematic directional shift. Write the comparison in a standalone script — never through the pipeline's cache layer, which would overwrite production entries. 2. **Gold re-validation:** re-score the gold/validation panel fresh and compare agreement metrics to the prior validation (e.g. median F1/κ within ~0.03, no domain dropping >0.10).
Also capture `response.model` on every gate call — pipelines rarely store it, and it is the only direct evidence of a repoint. Check the provider's `/models` endpoint: if the old model id is gone, no rollback exists.
3. **If the canary fails, diagnose BEFORE concluding — two mandatory follow-ups:** - **Date the suspected flip against the provider's changelog** before inferring a model splice. `response.model` on fresh calls identifies today's model only; if the alias already pointed there when the cache was written, there is no family splice and the mismatch needs another explanation. (Verified failure mode: an alias that had served the "new" model for months was misread as a fresh repoint.) - **Fresh-vs-fresh canary** to separate serving drift from temperature-0 nondeterminism: re-score the same items a second time. Drift signature = fresh2-vs-fresh1 agreement high and symmetric while both fresh runs disagree with the cache at a higher rate in the SAME signed direction. Noise signature = fresh-vs-fresh disagrees about as much as fresh-vs-cache, with no directional bias. - Supporting forensic: compare raw-response formatting fingerprints (JSON pretty/compact ratio, key order) between cache and fresh — a heterogeneous or shifted style distribution corroborates a serving change when no model id was recorded.
**Key subtlety (why both checks):** a new model or revision can validate AGAINST GOLD as well as the old one (κ holds or improves) while still disagreeing with the old scores on 10–30% of items, concentrated in borderline-heavy fields. Gold agreement does not license splicing — the gate fails on the canary alone. And alias stability is not serving stability: the same alias serving the same model family can still drift across days via silent serving revisions; a canary-failed resume is a seam either way, and the decision (resume with a documented seam vs re-score the universe) belongs to the budget owner.
## Verification
The gate script logs: fresh `response.model` ids, canary exact-match rate, per-field mismatch counts with signed direction, and the gold-metric deltas. GO only if both checks pass.
## Example
T1 exposure_v2 resume, 2026-07-16: canary returned 71% exact (gate ≥97%) with a litigation-concentrated negative shift, yet holdout median κ improved 0.607→0.644. First interpretation — "alias repointed to a new model family" — was WRONG: the provider changelog showed `deepseek-chat` had served v4-flash since April, months before the campaign. The fresh-vs-fresh follow-up then isolated the true cause: fresh2-vs-fresh1 93% exact/symmetric/litigation 0, both fresh runs vs cache 71–72% with litigation −12 identically — a serving revision within the same model across a 2-day gap, corroborated by a shifted JSON-formatting fingerprint. Total diagnosis cost ~$0.30; the resume-vs-rescore decision went to the budget owner with the seam quantified.
## Notes
- Design campaigns for this failure: per-response content-addressed cache + append-only checkpoint makes "re-score everything under the new model" a clean cache-rotation, not a data loss. - If the cache key embeds the alias string rather than the resolved model, record actual `response.model` in run reports — the cache cannot tell you later which model produced an entry. - One campaign = one model. Budget and schedule so the universe completes within days, or accept that a provider release can force a full re-score. - See also: [llm-gold-bound-failure-check] for the companion pre-campaign check — whether a validation-gate failure is fixable by prompt at all, or bound to the gold construct.
技术详情
- 版本
- 1.1.0
- 许可证
- MIT
- 最近更新
- 2026年8月24日
- 发布时间
- 2026年8月24日
决策摘要
备选候选
仓库近期活跃
Agent 验证证据
Agent 验证证据
来自解析、审查、安装和一次小范围运行后的结果报告。
- 成功率
- —
- 近期失败
- —
- 结果
- 0
- 输出质量
- —
- 失败
- 0
- 不相关
- 0
- 安装次数
- 0
- 风险拦截
- 0
- 需要配置
- 0
- 生产环境
- 0
暂时没有 Agent 结果数据。首次 Agent 执行可以通过 /api/agent/outcome 报告成功、需要设置、风险拦截、失败或不相关。
增长闭环
分享工具包
为 llm-campaign-drift-gate 准备的场景化草稿,可手动发布到 X。
llm-campaign-drift-gate: Gate resumption of any multi-day LLM batch-scoring campaign that calls an unpinned model alia... 47 stars https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate?ref=x
可选:带安装命令的回复
Listing + install path for llm-campaign-drift-gate: https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate?ref=x Install: npx skills add kennethkhoocy/applied-micro-skills --skill llm-campaign-drift-gate
收录来源
Registry 收录
此列表来自公开来源,维护者认领获批前不会标记为官方。
- 创作者
- Claude Code
- 收录方
- OpenAgentSkill 社区索引
归属链接指向公开仓库或创作者主页。创作者可认领列表以更新所有权信号。
认领此 Skill所有者认领
认领此 Skill 页面
这条 Registry 收录 列表归属于 Claude Code,但尚未标记为官方。认领后可增加已验证所有者信号,使后续发布、安装和审计更新更值得信赖。
创作者外链工具包
将证据徽章加入你的 README
在开发者评估仓库的位置展示规范页面、当前信任与审计信号,以及真实的 Agent 验证证据。
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate/audit)
[](https://www.openagentskill.com/skills/kennethkhoocy-llm-campaign-drift-gate)作者
Claude Code
@claude-code
健康信号
- GitHub Stars
- 47
- 质量评分
- 35/100
- 最近 GitHub 推送
- 2026年8月24日
- 框架提示
- 未知
- OpenAgentSkill 浏览量
- 2
- 复制安装命令
- 0
- 跳转点击
- 0
社区信号
告诉我们这个 Skill 是否对你的 Agent 工作流有帮助。汇总反馈会持续改善排序。
信任与安全
仅限沙盒
- GitHub 采用度47 个 GitHub Stars检查
- Star/Fork 活跃度47 个 Star,0 个 Fork; 当前元数据中没有议题活跃度信息检查
- 近期维护今天有推送通过
- 许可证清晰度MIT通过
- README/SKILL.md 完整度元数据包含足够的用法与工作流上下文通过
- 依赖与运行时风险公开元数据中未发现主要依赖风险提示通过
相关 Skill
Opencode
The open source coding agent.
200.7K StarsCode Review
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
168.6K StarsGrill With Docs
A relentless interview that pressure-tests a plan against the codebase, sharpens domain language, and updates CONTEXT.md and ADRs when decisions become durable.
164.7K StarsTo Spec
Turn the current conversation and codebase context into a structured implementation spec, then publish it to the configured project issue tracker.
164.7K Stars