Creator · wanshuiyin
Last updated · Sep 2, 2026
Workflow 1.5: Bridge between idea discovery and auto review. Reads EXPERIMENT_PLAN.md, implements experiment code, deploys to GPU, collects initial results. Use when user says \"实现实验\", \"implement experiments\", \"bridge\", \"从计划到跑实验\", \"deploy the plan\", or has an experiment
Sandbox only
Install targets
Codex install prompt
Install the "experiment-bridge" agent skill from https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/experiment-bridge. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Workflow 1.5: Bridge between idea discovery and auto review. Reads EXPERIMENT_PLAN.md, implements experiment code, deploys to GPU, collects initial results. Use when user says \"实现实验\", \"implement experiments\", \"bridge\", \"从计划到跑实验\", \"deploy the plan\", or has an experiment plan ready to execute. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wanshuiyin-experiment-bridge","task":"Install experiment-bridge","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
Coding agents
I need a coding agent that can understand a repository, edit code, and review pull requests.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-bridge
Maintenance
fresh
13d since push
Risk
Needs review
Permission surface may require sandboxing
GitHub quality
16K
89/100 Quality · 82/100 Trust
Coverage tags
Review notes
Permission surface may require sandboxing · Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
16K GitHub stars
Repo activity
16K stars, 1.4K forks
Maintenance
13d since push
License
MIT
Install
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-bridge
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-bridgeDo not use when
Alternative
168.6K Stars
npx skills add mattpocock/skills --skill code-review
Alternative
40.8K Stars
npx skills add appsmithorg/appsmith
Alternative
175.7K Stars
npx skills add mattpocock/skills --skill implement
Alternative
31.0K Stars
npx skills add vercel-labs/agent-skills --skill vercel-react-best-practices
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
medium
Skill may inspect schemas, query databases, or work with persistent stores.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20experiment-bridge%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20experiment-bridge%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wanshuiyin-experiment-bridge/install
Agent should check
Copy prompt
Task: Use experiment-bridge in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20experiment-bridge%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wanshuiyin-experiment-bridge/install
Install command: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-bridge
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wanshuiyin-experiment-bridge/install
LLM text format
/api/skills/wanshuiyin-experiment-bridge/install?format=text
Find alternatives
/api/skills/search?q=experiment-bridge&limit=3
Agent prompt
Use experiment-bridge for this task. Review https://www.openagentskill.com/api/skills/wanshuiyin-experiment-bridge/install, then install with: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-bridgeRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wanshuiyin-experiment-bridge
LLM text
/api/registry/manifest/wanshuiyin-experiment-bridge?format=text
Install alias
/api/registry/install/wanshuiyin-experiment-bridge
Recommend
/api/registry/recommend?task=Use%20experiment-bridge%20in%20an%20agent%20workflow&limit=3
Agent fit
Research agents
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Research agents
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS16K GitHub stars
Stars/forks activity
PASS16K stars, 1.4K forks; issue activity unavailable in current metadata
Recent maintenance
PASS13d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Build and ship code
I need a coding agent that can understand a repository, edit code, and review pull requests.
Manage repositories
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Workflow fit
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Alternative shortlist
Similar skills that may fit this task.
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
Platform to build admin panels, internal tools, and dashboards. Integrates with 25+ databases and any API.
Implement work from an approved spec or ticket set, run focused and full tests, invoke code review, and commit the result to the current branch.
React and Next.js performance guidance for writing, reviewing, and refactoring production UI code.
--- name: experiment-bridge description: "Workflow 1.5: Bridge between idea discovery and auto review. Reads EXPERIMENT_PLAN.md, implements experiment code, deploys to GPU, collects initial results. Use when user says \"实现实验\", \"implement experiments\", \"bridge\", \"从计划到跑实验\", \"deploy the plan\", or has an experiment plan ready to execute." argument-hint: "[experiment-plan-path-or-topic]" allowed-tools: Bash(*), Read, Write, Edit, Grep, Glob, Skill, mcp__codex__codex, mcp__codex__codex-reply ---
# Workflow 1.5: Experiment Bridge
Implement and deploy experiments from plan: **$ARGUMENTS**
## Overview
This skill bridges Workflow 1 (idea discovery + method refinement) and Workflow 2 (auto review loop). It takes the experiment plan and turns it into running experiments with initial results.
``` Workflow 1 output: This skill: Workflow 2 input: refine-logs/EXPERIMENT_PLAN.md → implement → GPT-5.6-Sol review → deploy → collect → initial results ready refine-logs/EXPERIMENT_TRACKER.md code (cross-model) /run-experiment for /auto-review-loop refine-logs/FINAL_PROPOSAL.md ```
## Constants
- **CODE_REVIEW = true** — GPT-5.6-Sol xhigh reviews experiment code before deployment. Catches logic bugs before wasting GPU hours. Set `false` to skip. - **AUTO_DEPLOY = true** — Automatically deploy experiments after implementation + review. Set `false` to manually inspect code before deploying. - **SANITY_FIRST = true** — Run the sanity-stage experiment first (smallest, fastest) before launching the rest. Catches setup bugs early. - **MAX_PARALLEL_RUNS = 4** — Maximum number of experiments to deploy in parallel (limited by available GPUs). - **BASE_REPO = false** — GitHub repo URL to use as base codebase. When set, clone the repo first and implement experiments on top of it. When `false` (default), write code from scratch or reuse existing project files. - **COMPACT = false** — When `true`, (1) read `idea-stage/IDEA_CANDIDATES.md` instead of full `idea-stage/IDEA_REPORT.md` if available, (2) append experiment results to `EXPERIMENT_LOG.md` after collection.
> Override: `/experiment-bridge "EXPERIMENT_PLAN.md" — compact: true, base repo: https://github.com/org/project`
## Inputs
This skill expects one or more of:
1. **`refine-logs/EXPERIMENT_PLAN.md`** (best) — claim-driven experiment roadmap from `/experiment-plan` 2. **`refine-logs/EXPERIMENT_TRACKER.md`** — run-by-run execution table 3. **`refine-logs/FINAL_PROPOSAL.md`** — method description for implementation context 4. **`idea-stage/IDEA_CANDIDATES.md`** — compact idea summary (preferred when `COMPACT: true`) *(fall back to `./IDEA_CANDIDATES.md` if not found)* 5. **`idea-stage/IDEA_REPORT.md`** — full brainstorm output *(fall back to `./IDEA_REPORT.md` if not found)*
If none exist, ask the user what experiments to implement.
## Workflow
### Phase 1: Parse the Experiment Plan
Read `EXPERIMENT_PLAN.md` and extract:
1. **Run order and milestones** — which experiments run first (sanity → baseline → main → ablation → polish) 2. **For each experiment block:** - Dataset / split / task - Compared systems and variants - Metrics to compute - Setup details (backbone, hyperparameters, seeds) - Success criterion - Priority (MUST-RUN vs NICE-TO-HAVE) 3. **Compute budget** — total estimated GPU-hours 4. **Method details** from `FINAL_PROPOSAL.md` — what exactly to implement
Present a brief summary:
``` 📋 Experiment plan loaded: - Milestones: [N] (sanity → baseline → main → ablation) - Must-run experiments: [N] - Nice-to-have: [N] - Estimated GPU-hours: [X]
Proceeding to implementation. ```
**Research-contract fallback**: if `idea-stage/docs/research_contract.md` does not exist yet (idea selected outside `/idea-discovery`, or an older run), create it now from `templates/RESEARCH_CONTRACT_TEMPLATE.md` using the selected idea + claims from the experiment plan. Downstream `/result-to-claim` and `/ablation-planner` read this file as the claims source, and session recovery (`docs/SESSION_RECOVERY_GUIDE.md`) depends on it existing.
### Phase 2: Implement Experiment Code
**If `BASE_REPO` is set** — clone the repo first: ```bash git clone <BASE_REPO> base_repo/ # Read the repo's README, understand its structure, find entry points # Implement experiments by modifying/extending this codebase ```
For each milestone (in order), write the experiment scripts:
1. **Check existing code** — scan the project (or cloned `base_repo/`) for existing experiment scripts, model code, data loaders. Reuse as much as possible.
2. **Implement missing pieces:** - Training scripts with proper argparse (all hyperparameters configurable) - Evaluation scripts computing the specified metrics - Data loading / preprocessing if needed - Baseline implementations if not already present - Fixed random seeds for reproducibility - Results saved to JSON/CSV for later analysis - Proper logging (wandb if configured in CLAUDE.md)
3. **Follow the plan's run order** — implement sanity-stage experiments first, then baselines, then main method, then ablations.
4. **Self-review before deploying:** - Are all hyperparameters from EXPERIMENT_PLAN.md reflected in argparse? - Is the random seed fixed and controllable? - Are results saved in a parseable format (JSON/CSV)? - Does the code match FINAL_PROPOSAL.md's method description?
### Phase 2.5: Cross-Model Code Review (when CODE_REVIEW = true)
**Skip this step if `CODE_REVIEW` is `false`.**
Before deploying, send the experiment code to GPT-5.6-Sol xhigh for review:
``` mcp__codex__codex: model: gpt-5.6-sol config: {"model_reasoning_effort": "xhigh"} prompt: | Review the following experiment implementation for correctness.
## Experiment Plan: [paste key sections from EXPERIMENT_PLAN.md]
## Method Description: [paste from FINAL_PROPOSAL.md]
## Implementation: [paste the experiment scripts]
Check for: 1. Does the code correctly implement the method described in the proposal? 2. Are all hyperparameters from the plan reflected in the code? 3. Are there any logic bugs (wrong loss function, incorrect data split, missing eval)? 4. Is the evaluation metric computed correctly? 5. **CRITICAL: Does evaluation use the dataset's actual ground truth labels — NOT another model's output as ground truth?** This is a common and severe bug. 6. Any potential issues (OOM risk, numerical instability, missing seeds)?
For each issue found, specify: CRITICAL / MAJOR / MINOR and the exact fix.
=== SCOPE LIMITS (these bound what you PROPOSE, never what you look for) === Report anything that is actually wrong here — including a rare-looking case, if this repo actually produces it. Then keep the fix in scope: 1. This is a RESEARCH-WORKFLOW tool, not a security paper. Verification is welcome; over-defense is not. Assume a cooperating operator on their own machine — a malicious local user is NOT in the threat model. 2. Do NOT propose SHA / hash / content-fingerprint / digest-binding schemes. Reporting a real defect in hashing code that already exists is fine. 3. NO speculative machinery: do not add feature flags, migration frameworks, compat layers, wrappers, pins, or similar mechanisms unless evidence shows a current repo defect they fix or an explicit existing invariant they must preserve. "Load-bearing", "compatibility", and "not scaffolding" are labels, not evidence. Point to the failing path/artifact or invariant, and check the proposal's factual premises, such as whether a named package version exists. 4. NO corner-case obsession: exotic encodings, symlink races, RTL text and millisecond races are out of scope unless you can show the case arises here. 5. Where a rubric or checklist is genuinely needed, do not over-mechanize judgement. A clear sentence a human reads beats a scored table nobody maintains. Exception: code that runs remote commands, starts a network service, or installs an MCP server runs on the user's machine with their credentials — trust-boundary findings there are in scope and the default is strict. Say plainly when something is correct. Do not manufacture findings. ```
**On review results:** - **No CRITICAL issues** → proceed to Phase 3 - **CRITICAL issues found** → fix them, then re-submit for review (max 2 rounds) - **Codex MCP unavailable** → skip silently, proceed to Phase 3 (graceful degradation)
### Phase 3: Sanity Check (if SANITY_FIRST = true)
Before deploying the full experiment suite, run the sanity-stage experiment:
``` /run-experiment [sanity experiment command] ```
Wait for completion. Verify: - Training loop runs without errors - Metrics are computed and saved correctly - GPU memory usage is within bounds - Output format matches expectations
If sanity fails → **auto-debug before giving up**. Budget: up to **2 patch attempts** on the same failure, then up to **2 clean reimplements** (4 total):
1. **Read the error** — parse traceback, stderr, and log files. (The same read-the-primary-artifact discipline applies to surprising REVIEWER verdicts: see `shared-references/review-tracing.md` § *Debugging With Traces*.) 2. **Diagnose** — classify the failure: - OOM → reduce batch size or enable gradient checkpointing - ImportError → install missing package - FileNotFoundError → fix path or download data - CUDA error → check GPU availability, reduce model size - NaN/divergence → reduce learning rate, check data preprocessing 3. **Fix and re-run** — apply the fix, re-run sanity 4. **Attempt 2+ still failing? → Call in Codex rescue** (if Codex plugin installed): Before the next retry, invoke `/codex:rescue` to get a second opinion on the root cause. Codex independently reads the code and error logs — it may spot issues Claude missed (wrong tensor shapes, subtle import shadowing, config mismatches, etc.). Apply its suggested fix, then re-run. - If `/codex:rescue` is not available (plugin not installed), continue with Claude's own diagnosis 5. **Both patch attempts failed on the same failure? → Discard and reimplement cleanly** (up to 2 reimplements). Rewriting the failing script from `EXPERIMENT_PLAN.md` / the research contract is a PEER move to another patch, not a last resort — a third patch on top of two wrong ones is usually worse than a clean rebuild. Delete ONLY the attempt's own code/scaffolding (scripts this phase generated); the plan, `EXPERIMENT_TRACKER.md`, user-authored project source, collected data, and results are never deletable (see `shared-references/external-cadence.md` § *Let a broken attempt restart, not just patch*). 6. **Budget exhausted (2 patches + 2 reimplements), or two reimplements failed the SAME way?** → stop, report the failure with all attempted fixes and error logs. Two clean reimplements failing identically usually means the plan or the environment is wrong — say so explicitly in the report, because that (not the broken build itself) is what needs the human. Do not proceed with broken code.
> Never give up on the first failure. Most experiment crashes are fixable without human intervention.
### Phase 4: Deploy Full Experiments
Deploy experiments following the plan's milestone order. **Route by job count**:
**Small batch (≤5 jobs per milestone)** → use `/run-experiment` directly: ``` /run-experiment [experiment commands] ```
**Large batch (≥10 jobs, multi-seed sweeps, or phase dependencies)** → use `/experiment-queue` for proper orchestration: ``` /experiment-queue [grid spec or manifest] ```
Auto-routing rule: if any milestone in `EXPERIMENT_PLAN.md` declares ≥10 jobs (e.g., `seeds: [42, 200, 201, ...]` × `N: [64, 128, 256]` × `n: [50K, 150K, 500K, 652K]` = 36 jobs) or declares teacher→student phase depende
Source provenance
Decision snapshot
15,641 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for experiment-bridge, ready for a manual X post.
A practical pick for source-backed research: experiment-bridge: Workflow 1.5: Bridge between idea discovery and auto review. Reads EXPERIMENT_PLAN.md, implements experiment code, deploys... 15.6K stars https://www.openagentskill.com/skills/wanshuiyin-experiment-bridge?ref=x
Listing + install path for experiment-bridge: https://www.openagentskill.com/skills/wanshuiyin-experiment-bridge?ref=x Install: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-b...
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to wanshuiyin but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-bridge?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-bridge?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-bridge/audit)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-bridge?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)wanshuiyin
@wanshuiyin
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Code Review
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
168.6K StarsAppsmith
Platform to build admin panels, internal tools, and dashboards. Integrates with 25+ databases and any API.
40.8K StarsImplement
Implement work from an approved spec or ticket set, run focused and full tests, invoke code review, and commit the result to the current branch.
175.7K StarsVercel React Best Practices
React and Next.js performance guidance for writing, reviewing, and refactoring production UI code.
31.0K StarsPermission surface
shell or command execution, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness