Creator · wanshuiyin
Last updated · Sep 2, 2026
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
Creator · wanshuiyin
Last updated · Sep 2, 2026
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
Creator · wanshuiyin
Last updated · Sep 2, 2026
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
Creator · wanshuiyin
Last updated · Sep 2, 2026
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
Review then install
Install targets
Codex install prompt
Install the "ablation-planner" agent skill from https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/ablation-planner. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wanshuiyin-ablation-planner","task":"Install ablation-planner","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Maintenance
fresh
11d since push
Risk
Safe to try
Quality score needs review
GitHub quality
16K
89/100 Quality · 84/100 Trust
Coverage tags
Review notes
Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Review then installGood shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
Audit
Safe to tryA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Use as the primary candidate after human or sandbox review.
Stars
16K GitHub stars
Repo activity
16K stars, 1.4K forks
Maintenance
11d since push
License
MIT
Install
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
Agent should check
Copy prompt
Task: Use ablation-planner in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install
Install command: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
LLM text format
/api/skills/wanshuiyin-ablation-planner/install?format=text
Find alternatives
/api/skills/search?q=ablation-planner&limit=3
Agent prompt
Use ablation-planner for this task. Review https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install, then install with: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wanshuiyin-ablation-planner
LLM text
/api/registry/manifest/wanshuiyin-ablation-planner?format=text
Install alias
/api/registry/install/wanshuiyin-ablation-planner
Recommend
/api/registry/recommend?task=Use%20ablation-planner%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Browser automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Good shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
GitHub adoption
PASS16K GitHub stars
Stars/forks activity
PASS16K stars, 1.4K forks; issue activity unavailable in current metadata
Recent maintenance
PASS11d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Use as the primary candidate after human or sandbox review.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: ablation-planner description: "Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission." argument-hint: "[method-description-or-claim]" allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply ---
# Ablation Planner
Systematically design ablation studies that answer the questions reviewers will ask. Codex leads the design (reviewer perspective), CC reviews feasibility and implements.
## Context: $ARGUMENTS
## When to Use
- Main results pass `/result-to-claim` with claim_supported = yes or partial - User explicitly requests ablation planning - `/auto-review-loop` reviewer identifies missing ablations
## Workflow
### Step 1: Prepare Context
CC reads available project files to build the full picture: - Method description and components (from `idea-stage/docs/research_contract.md`, legacy `docs/research_contract.md`, or project CLAUDE.md) - Current experiment results (from EXPERIMENT_LOG.md, EXPERIMENT_TRACKER.md, or W&B) - Confirmed and intended claims (from result-to-claim output or project notes) - Available compute resources (from CLAUDE.md server config, if present)
### Step 2: Codex Designs Ablations
``` mcp__codex__codex: model: gpt-5.6-sol config: {"model_reasoning_effort": "xhigh"} prompt: | You are a rigorous ML reviewer planning ablation studies. Given this method and results, design ablations that:
1. Isolate the contribution of each novel component 2. Answer questions reviewers will definitely ask 3. Test sensitivity to key hyperparameters 4. Compare against natural alternative design choices
Method: [description from project files] Components: [list of removable/replaceable components] Current results: [key metrics from experiments] Claims: [what we claim and current evidence]
For each ablation, specify: - name: what to change (e.g., "remove module X", "replace Y with Z") - what_it_tests: the specific question this answers - expected_if_component_matters: what we predict if the component is important - priority: 1 (must-run) to 5 (nice-to-have)
Also provide: - coverage_assessment: what reviewer questions these ablations answer - unnecessary_ablations: experiments that seem useful but won't add insight - suggested_order: run order optimized for maximum early information - estimated_compute: total GPU-hours estimate ```
### Step 3: Parse Ablation Plan
Normalize Codex response into structured format:
```markdown ## Ablation Plan
### Component Ablations (highest priority) | # | Name | What It Tests | Expected If Matters | Priority | |---|------|---------------|---------------------|----------| | 1 | remove module X | contribution of X | performance drops on metric Y | 1 | | 2 | replace X with simpler Z | value of learned vs fixed | drops, especially on dataset A | 2 |
### Hyperparameter Sensitivity | # | Parameter | Values to Test | What It Tests | Priority | |---|-----------|---------------|---------------|----------| | 3 | lambda | [0.01, 0.1, 1.0] | sensitivity to regularization | 3 |
### Design Choice Comparisons | # | Name | What It Tests | Priority | |---|------|---------------|----------| | 4 | joint vs separate matching | whether joint adds value | 4 |
### Coverage Assessment [What reviewer questions these ablations answer]
### Unnecessary Ablations [Experiments that seem useful but won't add insight — skip these]
### Run Order [Optimized for maximum early information]
### Estimated Compute [Total GPU-hours] ```
### Step 4: CC Reviews Feasibility
Before running anything, CC checks: - Compute budget: can we afford all ablations with available GPUs? - Code changes: which ablations need code modifications vs config-only changes? - Dependencies: which ablations can run in parallel? - Cuts: if budget is tight, propose removing lower-priority ablations and ask Codex to confirm
### Step 5: Implement and Run
1. Create configs/scripts for each ablation (config-only changes first) 2. Smoke test each ablation before full run 3. Run in suggested order, using descriptive names (e.g., `ablation-no-module-X`) 4. Track results in EXPERIMENT_LOG.md 5. After all ablations complete → update findings.md with insights
## Rules
- **Codex leads the design. CC does not pre-filter or bias the ablation list** before Codex sees it. Codex thinks like a reviewer; CC thinks like an engineer. - Every ablation must have a clear `what_it_tests` and `expected_if_component_matters`. No "just try it" experiments. - Config-only ablations take priority over those needing code changes (faster, less error-prone). - If total compute exceeds budget, CC proposes cuts and asks Codex to re-prioritize — don't silently drop ablations. - Component ablations (remove/replace) take priority over hyperparameter sweeps. - Do not generate ablations for components identical to the baseline (no-op ablations). - Record all ablation results in EXPERIMENT_LOG.md, including negative results (component removal had no effect = important finding).
Source provenance
Decision snapshot
15,641 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for ablation-planner, ready for a manual X post.
ablation-planner: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation stud... 15.6K stars https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x
Listing + install path for ablation-planner: https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x Install: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-pla...
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to wanshuiyin but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner/audit)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)wanshuiyin
@wanshuiyin
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Review then install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsReview then install
Install targets
Codex install prompt
Install the "ablation-planner" agent skill from https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/ablation-planner. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wanshuiyin-ablation-planner","task":"Install ablation-planner","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Maintenance
fresh
11d since push
Risk
Safe to try
Quality score needs review
GitHub quality
16K
89/100 Quality · 84/100 Trust
Coverage tags
Review notes
Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Review then installGood shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
Audit
Safe to tryA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Use as the primary candidate after human or sandbox review.
Stars
16K GitHub stars
Repo activity
16K stars, 1.4K forks
Maintenance
11d since push
License
MIT
Install
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
Agent should check
Copy prompt
Task: Use ablation-planner in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install
Install command: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
LLM text format
/api/skills/wanshuiyin-ablation-planner/install?format=text
Find alternatives
/api/skills/search?q=ablation-planner&limit=3
Agent prompt
Use ablation-planner for this task. Review https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install, then install with: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wanshuiyin-ablation-planner
LLM text
/api/registry/manifest/wanshuiyin-ablation-planner?format=text
Install alias
/api/registry/install/wanshuiyin-ablation-planner
Recommend
/api/registry/recommend?task=Use%20ablation-planner%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Browser automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Good shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
GitHub adoption
PASS16K GitHub stars
Stars/forks activity
PASS16K stars, 1.4K forks; issue activity unavailable in current metadata
Recent maintenance
PASS11d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Use as the primary candidate after human or sandbox review.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: ablation-planner description: "Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission." argument-hint: "[method-description-or-claim]" allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply ---
# Ablation Planner
Systematically design ablation studies that answer the questions reviewers will ask. Codex leads the design (reviewer perspective), CC reviews feasibility and implements.
## Context: $ARGUMENTS
## When to Use
- Main results pass `/result-to-claim` with claim_supported = yes or partial - User explicitly requests ablation planning - `/auto-review-loop` reviewer identifies missing ablations
## Workflow
### Step 1: Prepare Context
CC reads available project files to build the full picture: - Method description and components (from `idea-stage/docs/research_contract.md`, legacy `docs/research_contract.md`, or project CLAUDE.md) - Current experiment results (from EXPERIMENT_LOG.md, EXPERIMENT_TRACKER.md, or W&B) - Confirmed and intended claims (from result-to-claim output or project notes) - Available compute resources (from CLAUDE.md server config, if present)
### Step 2: Codex Designs Ablations
``` mcp__codex__codex: model: gpt-5.6-sol config: {"model_reasoning_effort": "xhigh"} prompt: | You are a rigorous ML reviewer planning ablation studies. Given this method and results, design ablations that:
1. Isolate the contribution of each novel component 2. Answer questions reviewers will definitely ask 3. Test sensitivity to key hyperparameters 4. Compare against natural alternative design choices
Method: [description from project files] Components: [list of removable/replaceable components] Current results: [key metrics from experiments] Claims: [what we claim and current evidence]
For each ablation, specify: - name: what to change (e.g., "remove module X", "replace Y with Z") - what_it_tests: the specific question this answers - expected_if_component_matters: what we predict if the component is important - priority: 1 (must-run) to 5 (nice-to-have)
Also provide: - coverage_assessment: what reviewer questions these ablations answer - unnecessary_ablations: experiments that seem useful but won't add insight - suggested_order: run order optimized for maximum early information - estimated_compute: total GPU-hours estimate ```
### Step 3: Parse Ablation Plan
Normalize Codex response into structured format:
```markdown ## Ablation Plan
### Component Ablations (highest priority) | # | Name | What It Tests | Expected If Matters | Priority | |---|------|---------------|---------------------|----------| | 1 | remove module X | contribution of X | performance drops on metric Y | 1 | | 2 | replace X with simpler Z | value of learned vs fixed | drops, especially on dataset A | 2 |
### Hyperparameter Sensitivity | # | Parameter | Values to Test | What It Tests | Priority | |---|-----------|---------------|---------------|----------| | 3 | lambda | [0.01, 0.1, 1.0] | sensitivity to regularization | 3 |
### Design Choice Comparisons | # | Name | What It Tests | Priority | |---|------|---------------|----------| | 4 | joint vs separate matching | whether joint adds value | 4 |
### Coverage Assessment [What reviewer questions these ablations answer]
### Unnecessary Ablations [Experiments that seem useful but won't add insight — skip these]
### Run Order [Optimized for maximum early information]
### Estimated Compute [Total GPU-hours] ```
### Step 4: CC Reviews Feasibility
Before running anything, CC checks: - Compute budget: can we afford all ablations with available GPUs? - Code changes: which ablations need code modifications vs config-only changes? - Dependencies: which ablations can run in parallel? - Cuts: if budget is tight, propose removing lower-priority ablations and ask Codex to confirm
### Step 5: Implement and Run
1. Create configs/scripts for each ablation (config-only changes first) 2. Smoke test each ablation before full run 3. Run in suggested order, using descriptive names (e.g., `ablation-no-module-X`) 4. Track results in EXPERIMENT_LOG.md 5. After all ablations complete → update findings.md with insights
## Rules
- **Codex leads the design. CC does not pre-filter or bias the ablation list** before Codex sees it. Codex thinks like a reviewer; CC thinks like an engineer. - Every ablation must have a clear `what_it_tests` and `expected_if_component_matters`. No "just try it" experiments. - Config-only ablations take priority over those needing code changes (faster, less error-prone). - If total compute exceeds budget, CC proposes cuts and asks Codex to re-prioritize — don't silently drop ablations. - Component ablations (remove/replace) take priority over hyperparameter sweeps. - Do not generate ablations for components identical to the baseline (no-op ablations). - Record all ablation results in EXPERIMENT_LOG.md, including negative results (component removal had no effect = important finding).
Source provenance
Decision snapshot
15,641 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for ablation-planner, ready for a manual X post.
ablation-planner: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation stud... 15.6K stars https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x
Listing + install path for ablation-planner: https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x Install: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-pla...
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to wanshuiyin but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner/audit)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)wanshuiyin
@wanshuiyin
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Review then install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsReview then install
Install targets
Codex install prompt
Install the "ablation-planner" agent skill from https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/ablation-planner. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wanshuiyin-ablation-planner","task":"Install ablation-planner","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Maintenance
fresh
11d since push
Risk
Safe to try
Quality score needs review
GitHub quality
16K
89/100 Quality · 84/100 Trust
Coverage tags
Review notes
Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Review then installGood shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
Audit
Safe to tryA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Use as the primary candidate after human or sandbox review.
Stars
16K GitHub stars
Repo activity
16K stars, 1.4K forks
Maintenance
11d since push
License
MIT
Install
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
Agent should check
Copy prompt
Task: Use ablation-planner in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install
Install command: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
LLM text format
/api/skills/wanshuiyin-ablation-planner/install?format=text
Find alternatives
/api/skills/search?q=ablation-planner&limit=3
Agent prompt
Use ablation-planner for this task. Review https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install, then install with: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wanshuiyin-ablation-planner
LLM text
/api/registry/manifest/wanshuiyin-ablation-planner?format=text
Install alias
/api/registry/install/wanshuiyin-ablation-planner
Recommend
/api/registry/recommend?task=Use%20ablation-planner%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Browser automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Good shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
GitHub adoption
PASS16K GitHub stars
Stars/forks activity
PASS16K stars, 1.4K forks; issue activity unavailable in current metadata
Recent maintenance
PASS11d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Use as the primary candidate after human or sandbox review.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: ablation-planner description: "Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission." argument-hint: "[method-description-or-claim]" allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply ---
# Ablation Planner
Systematically design ablation studies that answer the questions reviewers will ask. Codex leads the design (reviewer perspective), CC reviews feasibility and implements.
## Context: $ARGUMENTS
## When to Use
- Main results pass `/result-to-claim` with claim_supported = yes or partial - User explicitly requests ablation planning - `/auto-review-loop` reviewer identifies missing ablations
## Workflow
### Step 1: Prepare Context
CC reads available project files to build the full picture: - Method description and components (from `idea-stage/docs/research_contract.md`, legacy `docs/research_contract.md`, or project CLAUDE.md) - Current experiment results (from EXPERIMENT_LOG.md, EXPERIMENT_TRACKER.md, or W&B) - Confirmed and intended claims (from result-to-claim output or project notes) - Available compute resources (from CLAUDE.md server config, if present)
### Step 2: Codex Designs Ablations
``` mcp__codex__codex: model: gpt-5.6-sol config: {"model_reasoning_effort": "xhigh"} prompt: | You are a rigorous ML reviewer planning ablation studies. Given this method and results, design ablations that:
1. Isolate the contribution of each novel component 2. Answer questions reviewers will definitely ask 3. Test sensitivity to key hyperparameters 4. Compare against natural alternative design choices
Method: [description from project files] Components: [list of removable/replaceable components] Current results: [key metrics from experiments] Claims: [what we claim and current evidence]
For each ablation, specify: - name: what to change (e.g., "remove module X", "replace Y with Z") - what_it_tests: the specific question this answers - expected_if_component_matters: what we predict if the component is important - priority: 1 (must-run) to 5 (nice-to-have)
Also provide: - coverage_assessment: what reviewer questions these ablations answer - unnecessary_ablations: experiments that seem useful but won't add insight - suggested_order: run order optimized for maximum early information - estimated_compute: total GPU-hours estimate ```
### Step 3: Parse Ablation Plan
Normalize Codex response into structured format:
```markdown ## Ablation Plan
### Component Ablations (highest priority) | # | Name | What It Tests | Expected If Matters | Priority | |---|------|---------------|---------------------|----------| | 1 | remove module X | contribution of X | performance drops on metric Y | 1 | | 2 | replace X with simpler Z | value of learned vs fixed | drops, especially on dataset A | 2 |
### Hyperparameter Sensitivity | # | Parameter | Values to Test | What It Tests | Priority | |---|-----------|---------------|---------------|----------| | 3 | lambda | [0.01, 0.1, 1.0] | sensitivity to regularization | 3 |
### Design Choice Comparisons | # | Name | What It Tests | Priority | |---|------|---------------|----------| | 4 | joint vs separate matching | whether joint adds value | 4 |
### Coverage Assessment [What reviewer questions these ablations answer]
### Unnecessary Ablations [Experiments that seem useful but won't add insight — skip these]
### Run Order [Optimized for maximum early information]
### Estimated Compute [Total GPU-hours] ```
### Step 4: CC Reviews Feasibility
Before running anything, CC checks: - Compute budget: can we afford all ablations with available GPUs? - Code changes: which ablations need code modifications vs config-only changes? - Dependencies: which ablations can run in parallel? - Cuts: if budget is tight, propose removing lower-priority ablations and ask Codex to confirm
### Step 5: Implement and Run
1. Create configs/scripts for each ablation (config-only changes first) 2. Smoke test each ablation before full run 3. Run in suggested order, using descriptive names (e.g., `ablation-no-module-X`) 4. Track results in EXPERIMENT_LOG.md 5. After all ablations complete → update findings.md with insights
## Rules
- **Codex leads the design. CC does not pre-filter or bias the ablation list** before Codex sees it. Codex thinks like a reviewer; CC thinks like an engineer. - Every ablation must have a clear `what_it_tests` and `expected_if_component_matters`. No "just try it" experiments. - Config-only ablations take priority over those needing code changes (faster, less error-prone). - If total compute exceeds budget, CC proposes cuts and asks Codex to re-prioritize — don't silently drop ablations. - Component ablations (remove/replace) take priority over hyperparameter sweeps. - Do not generate ablations for components identical to the baseline (no-op ablations). - Record all ablation results in EXPERIMENT_LOG.md, including negative results (component removal had no effect = important finding).
Source provenance
Decision snapshot
15,641 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for ablation-planner, ready for a manual X post.
ablation-planner: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation stud... 15.6K stars https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x
Listing + install path for ablation-planner: https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x Install: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-pla...
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to wanshuiyin but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner/audit)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)wanshuiyin
@wanshuiyin
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Review then install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsReview then install
Install targets
Codex install prompt
Install the "ablation-planner" agent skill from https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/ablation-planner. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wanshuiyin-ablation-planner","task":"Install ablation-planner","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Maintenance
fresh
11d since push
Risk
Safe to try
Quality score needs review
GitHub quality
16K
89/100 Quality · 84/100 Trust
Coverage tags
Review notes
Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Review then installGood shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
Audit
Safe to tryA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Use as the primary candidate after human or sandbox review.
Stars
16K GitHub stars
Repo activity
16K stars, 1.4K forks
Maintenance
11d since push
License
MIT
Install
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
Agent should check
Copy prompt
Task: Use ablation-planner in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20ablation-planner%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install
Install command: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wanshuiyin-ablation-planner/install
LLM text format
/api/skills/wanshuiyin-ablation-planner/install?format=text
Find alternatives
/api/skills/search?q=ablation-planner&limit=3
Agent prompt
Use ablation-planner for this task. Review https://www.openagentskill.com/api/skills/wanshuiyin-ablation-planner/install, then install with: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-plannerRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wanshuiyin-ablation-planner
LLM text
/api/registry/manifest/wanshuiyin-ablation-planner?format=text
Install alias
/api/registry/install/wanshuiyin-ablation-planner
Recommend
/api/registry/recommend?task=Use%20ablation-planner%20in%20an%20agent%20workflow&limit=3
Agent fit
Browser automation
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Browser automation
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Good shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
GitHub adoption
PASS16K GitHub stars
Stars/forks activity
PASS16K stars, 1.4K forks; issue activity unavailable in current metadata
Recent maintenance
PASS11d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Use as the primary candidate after human or sandbox review.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Operate web apps
I need my agent to control a browser, fill forms, and verify web app workflows.
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Workflow fit
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: ablation-planner description: "Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission." argument-hint: "[method-description-or-claim]" allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply ---
# Ablation Planner
Systematically design ablation studies that answer the questions reviewers will ask. Codex leads the design (reviewer perspective), CC reviews feasibility and implements.
## Context: $ARGUMENTS
## When to Use
- Main results pass `/result-to-claim` with claim_supported = yes or partial - User explicitly requests ablation planning - `/auto-review-loop` reviewer identifies missing ablations
## Workflow
### Step 1: Prepare Context
CC reads available project files to build the full picture: - Method description and components (from `idea-stage/docs/research_contract.md`, legacy `docs/research_contract.md`, or project CLAUDE.md) - Current experiment results (from EXPERIMENT_LOG.md, EXPERIMENT_TRACKER.md, or W&B) - Confirmed and intended claims (from result-to-claim output or project notes) - Available compute resources (from CLAUDE.md server config, if present)
### Step 2: Codex Designs Ablations
``` mcp__codex__codex: model: gpt-5.6-sol config: {"model_reasoning_effort": "xhigh"} prompt: | You are a rigorous ML reviewer planning ablation studies. Given this method and results, design ablations that:
1. Isolate the contribution of each novel component 2. Answer questions reviewers will definitely ask 3. Test sensitivity to key hyperparameters 4. Compare against natural alternative design choices
Method: [description from project files] Components: [list of removable/replaceable components] Current results: [key metrics from experiments] Claims: [what we claim and current evidence]
For each ablation, specify: - name: what to change (e.g., "remove module X", "replace Y with Z") - what_it_tests: the specific question this answers - expected_if_component_matters: what we predict if the component is important - priority: 1 (must-run) to 5 (nice-to-have)
Also provide: - coverage_assessment: what reviewer questions these ablations answer - unnecessary_ablations: experiments that seem useful but won't add insight - suggested_order: run order optimized for maximum early information - estimated_compute: total GPU-hours estimate ```
### Step 3: Parse Ablation Plan
Normalize Codex response into structured format:
```markdown ## Ablation Plan
### Component Ablations (highest priority) | # | Name | What It Tests | Expected If Matters | Priority | |---|------|---------------|---------------------|----------| | 1 | remove module X | contribution of X | performance drops on metric Y | 1 | | 2 | replace X with simpler Z | value of learned vs fixed | drops, especially on dataset A | 2 |
### Hyperparameter Sensitivity | # | Parameter | Values to Test | What It Tests | Priority | |---|-----------|---------------|---------------|----------| | 3 | lambda | [0.01, 0.1, 1.0] | sensitivity to regularization | 3 |
### Design Choice Comparisons | # | Name | What It Tests | Priority | |---|------|---------------|----------| | 4 | joint vs separate matching | whether joint adds value | 4 |
### Coverage Assessment [What reviewer questions these ablations answer]
### Unnecessary Ablations [Experiments that seem useful but won't add insight — skip these]
### Run Order [Optimized for maximum early information]
### Estimated Compute [Total GPU-hours] ```
### Step 4: CC Reviews Feasibility
Before running anything, CC checks: - Compute budget: can we afford all ablations with available GPUs? - Code changes: which ablations need code modifications vs config-only changes? - Dependencies: which ablations can run in parallel? - Cuts: if budget is tight, propose removing lower-priority ablations and ask Codex to confirm
### Step 5: Implement and Run
1. Create configs/scripts for each ablation (config-only changes first) 2. Smoke test each ablation before full run 3. Run in suggested order, using descriptive names (e.g., `ablation-no-module-X`) 4. Track results in EXPERIMENT_LOG.md 5. After all ablations complete → update findings.md with insights
## Rules
- **Codex leads the design. CC does not pre-filter or bias the ablation list** before Codex sees it. Codex thinks like a reviewer; CC thinks like an engineer. - Every ablation must have a clear `what_it_tests` and `expected_if_component_matters`. No "just try it" experiments. - Config-only ablations take priority over those needing code changes (faster, less error-prone). - If total compute exceeds budget, CC proposes cuts and asks Codex to re-prioritize — don't silently drop ablations. - Component ablations (remove/replace) take priority over hyperparameter sweeps. - Do not generate ablations for components identical to the baseline (no-op ablations). - Record all ablation results in EXPERIMENT_LOG.md, including negative results (component removal had no effect = important finding).
Source provenance
Decision snapshot
15,641 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for ablation-planner, ready for a manual X post.
ablation-planner: Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation stud... 15.6K stars https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x
Listing + install path for ablation-planner: https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=x Install: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-pla...
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to wanshuiyin but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner/audit)
[](https://www.openagentskill.com/skills/wanshuiyin-ablation-planner?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)wanshuiyin
@wanshuiyin
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Review then install
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsPermission surface
shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness
Permission surface
shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness
Permission surface
shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness
Permission surface
shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness