Creator · wanshuiyin
Last updated · Sep 3, 2026
SSH job queue for multi-seed/multi-config ML experiments with OOM-aware retry, stale-screen cleanup, and wave-transition race prevention. Use when user says "batch experiments", "队列实验", "run grid", "multi-seed sweep", "auto-chain experiments", or when /run-experiment is insuffici
Sandbox only
Install targets
Codex install prompt
Install the "experiment-queue" agent skill from https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/experiment-queue. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: SSH job queue for multi-seed/multi-config ML experiments with OOM-aware retry, stale-screen cleanup, and wave-transition race prevention. Use when user says "batch experiments", "队列实验", "run grid", "multi-seed sweep", "auto-chain experiments", or when /run-experiment is insufficient for 10+ jobs that need orchestration. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wanshuiyin-experiment-queue","task":"Install experiment-queue","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-queue
Maintenance
fresh
5d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
16K
89/100 Quality · 74/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review · Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
16K GitHub stars
Repo activity
16K stars, 1.4K forks
Maintenance
5d since push
License
MIT
Install
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-queue
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-queueDo not use when
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill may drive a browser or interact with web pages.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20experiment-queue%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20experiment-queue%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wanshuiyin-experiment-queue/install
Agent should check
Copy prompt
Task: Use experiment-queue in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20experiment-queue%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wanshuiyin-experiment-queue/install
Install command: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-queue
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wanshuiyin-experiment-queue/install
LLM text format
/api/skills/wanshuiyin-experiment-queue/install?format=text
Find alternatives
/api/skills/search?q=experiment-queue&limit=3
Agent prompt
Use experiment-queue for this task. Review https://www.openagentskill.com/api/skills/wanshuiyin-experiment-queue/install, then install with: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-queueRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wanshuiyin-experiment-queue
LLM text
/api/registry/manifest/wanshuiyin-experiment-queue?format=text
Install alias
/api/registry/install/wanshuiyin-experiment-queue
Recommend
/api/registry/recommend?task=Use%20experiment-queue%20in%20an%20agent%20workflow&limit=3
Agent fit
Research agents
Use-case tags
Platforms
Claude Code
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Research agents
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS16K GitHub stars
Stars/forks activity
PASS16K stars, 1.4K forks; issue activity unavailable in current metadata
Recent maintenance
PASS5d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Workflow fit
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Run multimodal agents that operate desktop interfaces
Connect agents to hundreds of workflow automations
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
--- name: experiment-queue description: SSH job queue for multi-seed/multi-config ML experiments with OOM-aware retry, stale-screen cleanup, and wave-transition race prevention. Use when user says "batch experiments", "队列实验", "run grid", "multi-seed sweep", "auto-chain experiments", or when /run-experiment is insufficient for 10+ jobs that need orchestration. argument-hint: "[manifest-or-grid-spec]" allowed-tools: Bash(*), Read, Grep, Glob, Edit, Write, Skill(run-experiment), Skill(monitor-experiment) ---
# Experiment Queue
> ⏱ **External cadence: visibility only.** This skill already runs its own > detached server-side scheduler (60s poll + `depends_on` + wave transitions). > Use its status output for overnight visibility (N done / N running / N > pending); do **not** wrap it in a second `/loop` / `CronCreate` poll — that > duplicates the scheduler on an uncoordinated clock and races the > wave-transition logic it was built to prevent. See > [`shared-references/external-cadence.md`](../shared-references/external-cadence.md) > ("don't duplicate an existing scheduler").
Orchestrate large batches of ML experiments on SSH remote GPU servers with proper state tracking, OOM retry, stale cleanup, and wave transitions.
## When to Use This Skill
Use when `/run-experiment` is insufficient: - **≥10 jobs** that need batching across GPUs - **Multi-seed sweeps** (e.g., 21 seeds × 12 cells) - **Wave transitions** (run wave 1, wait, run wave 2, wait, run wave 3...) - **Teacher+student chains** (train teacher then distill; auto-trigger student after teacher done) - **OOM-prone configs** where you need to retry with different GPU or wait - **Mixed seed grids** where failed cells need re-running
Do NOT use for: - Single ad-hoc experiment (use `/run-experiment`) - Modal/Vast.ai deployments (those have their own orchestration) - Experiments that need manual inspection between runs
## Why This Exists
Based on session audit (2026-04-16), the major wall-clock sinks in multi-seed grid experiments are:
1. **Stale screens** — python finishes, wandb uploads, screen hangs, next wave blocked 2. **OOM on shared GPU** — previous job's memory not yet released 3. **Wave race** — new wave launches before previous wave fully settles 4. **Missing checkpoints** — student launches before teacher saved 5. **Parser duplication** — rewriting multi-seed analysis python every batch
All of these are pure engineering friction that can be orchestrated.
## Core Concepts
> **Environment contract**: queue jobs assume the target env is already built > and validated per `../shared-references/compute-env-contract.md` (spec-hash > ledger + kernel witness). A wave of jobs dying at import time = the env > contract was skipped, not a queue bug; check the provider's > `.aris/compute/<provider>.md` ledger before re-queueing.
### Job Manifest
A manifest lists jobs with explicit state:
```yaml project: my_grid_experiment cwd: /home/user/your_project conda: my_env # Optional: override conda hook path if conda is not at a standard location. # Can be a bare path (wrapped automatically) or a full `eval "$(... shell.bash hook)"` string. # Falls back to auto-detect of ~/anaconda3, ~/miniconda3, /opt/anaconda3, etc., # or the ARIS_CONDA_HOOK environment variable. # conda_hook: /custom/path/to/conda ssh: gpu-server default_cmd: > python run_distill.py --backbone softmax --lam 0.5 --K 500 --L 96 --W 16 --n_steps 30000 --batch_size 128 --lr 1e-4
preconditions: - type: checkpoint_exists path: checkpoints/transformer/teacher_L96_K500_N{N}.pt
gpus: [0, 1, 2, 3, 4, 5, 6, 7] max_parallel: 8 gpu_free_threshold_mib: 500 # optional, default 500; raise for shared servers, lower for tight packing oom_retry: delay: 120 max_attempts: 3
jobs: - id: s200_N64_n50K args: {seed: 200, n_hidden: 64, n_train_subset: 50000, subset_seed: 2024} - id: s200_N128_n50K args: {seed: 200, n_hidden: 128, n_train_subset: 50000, subset_seed: 2024} # ... 14 more ```
### Job State Machine
``` pending → running → completed ↘ failed_oom → pending (after delay) [retry up to N] ↘ failed_other → stuck (needs manual inspection) stale screen (process gone, screen lingering) → failed_other → stuck ```
> **Operator note on `stuck` (the agent's move, not the queue's):** the queue > deterministically parks `failed_other` jobs as `stuck` — that part is code and > unchanged. Before handing a `stuck` batch to the human, the OPERATING AGENT > should check: if the same failure repeats across jobs, try ONE clean > reimplement of the **agent-generated wrapper/attempt script only** — never > user/project source (`run_*.py` you didn't write), the manifest, queue state, > logs, or results (see `shared-references/external-cadence.md` § *Let a broken > attempt restart, not just patch*). Reserve the human handoff for > contract/environment doubts, not merely broken attempt code.
### Wave Orchestration
A "wave" is a batch of jobs that fit available GPUs. Next wave only starts when: 1. All current-wave python processes have exited 2. No stale screens remain for current-wave tags 3. GPU memory has dropped below threshold (≤500 MiB) 4. Precondition checks pass for next-wave jobs
## Workflow
### Step 1: Parse Manifest / Build from Grid
Input can be: - **YAML manifest** (explicit job list, recommended for complex cases) - **Grid spec** (Cartesian product of param values, e.g., `N=[64,128,256] × n=[50K,150K,500K,652K]`) - **Natural language description** (Claude parses into manifest)
Bind the run identifiers once so every later step (manifest save, scp, launch, monitor, resume) refers to the same paths. Set these as local shell variables before generating the manifest:
```bash # REPLACE the placeholder path before running, or pre-export PROJECT_DIR: PROJECT_DIR="${PROJECT_DIR:?set PROJECT_DIR to the local project root}" RUN_TS=$(date -u +%Y%m%dT%H%M%SZ) # one timestamp per run, reused everywhere LOCAL_RUN_DIR="$PROJECT_DIR/experiment_queue/$RUN_TS" mkdir -p "$LOCAL_RUN_DIR" ```
Save the built manifest to `$LOCAL_RUN_DIR/manifest.json` for reproducibility.
### Step 2: Pre-flight
- Check SSH connection works - Check conda env exists on remote - Check `cwd` exists on remote - Check all preconditions (checkpoints, input files) - Check GPU availability (at least `max_parallel` free GPUs)
If any precondition fails, show user which jobs are blocked and why.
### Step 3: Launch Scheduler
The canonical scheduler implementation lives in `skills/experiment-queue/scripts/queue_manager.py` (Phase 3.3 move, Arch C). `tools/experiment_queue/queue_manager.py` is now a Python `os.execv` shim retained for legacy resolver-chain compatibility. Three preliminaries before launch.
**3a. Resolve the local helper directory.** The two helpers (`queue_manager.py`, `build_manifest.py`) now sit under `skills/experiment-queue/scripts/` in the ARIS repo, with shims at `tools/experiment_queue/` for legacy resolver layers. Use this hybrid chain so the skill works from any project layout:
```bash # Layer 0: self-contained (CC 1.0+ exposes $CLAUDE_SKILL_DIR). QUEUE_TOOLS="" if [ -n "${CLAUDE_SKILL_DIR:-}" ] && [ -f "$CLAUDE_SKILL_DIR/scripts/queue_manager.py" ]; then QUEUE_TOOLS="$CLAUDE_SKILL_DIR/scripts" fi # Layers 1-4: legacy chain via tools/experiment_queue/ shims. if [ -z "$QUEUE_TOOLS" ]; then cd "$(git rev-parse --show-toplevel 2>/dev/null || pwd)" || exit 1 if [ -z "${ARIS_REPO:-}" ] && [ -f .aris/installed-skills.txt ]; then ARIS_REPO=$(awk -F'\t' '$1=="repo_root"{print $2; exit}' .aris/installed-skills.txt 2>/dev/null) || true fi if [ -z "${ARIS_REPO:-}" ] && [ -f "$HOME/.aris/repo" ]; then ARIS_REPO=$(cat "$HOME/.aris/repo" 2>/dev/null) || true fi QUEUE_TOOLS=".aris/tools/experiment_queue" [ -f "$QUEUE_TOOLS/queue_manager.py" ] || QUEUE_TOOLS="tools/experiment_queue" [ -f "$QUEUE_TOOLS/queue_manager.py" ] || { [ -n "${ARIS_REPO:-}" ] && QUEUE_TOOLS="$ARIS_REPO/tools/experiment_queue"; } [ -f "$QUEUE_TOOLS/queue_manager.py" ] || QUEUE_TOOLS="" fi [ -z "$QUEUE_TOOLS" ] && { echo "ERROR: experiment_queue helpers not found (layer 0: \$CLAUDE_SKILL_DIR/scripts/; layers 1-4: .aris/tools/, tools/, \$ARIS_REPO/tools/, \$ARIS_REPO/tools/ via ~/.aris/repo). Rerun install_aris.sh or smart_update.sh (refreshes ~/.aris/repo), set ARIS_REPO, or copy the canonical scripts from \$ARIS_REPO/skills/experiment-queue/scripts/." >&2; exit 1; } ```
The `.aris/tools` symlink is set up by `install_aris.sh` (#174). Older installs without that symlink fall through to `tools/experiment_queue` (works if invoked from inside the ARIS repo), `$ARIS_REPO/tools/experiment_queue`, or the same path resolved via the global pointer file `~/.aris/repo` (#366, for installs with no project-local manifest). After Phase 3.3, each of those legacy paths contains a Python `os.execv` shim that forwards to the canonical `skills/experiment-queue/scripts/` location, so existing users do not need to re-run anything.
**3b. Compute remote paths.** Use both a remote-relative form (for `scp` destinations — modern `scp` runs in SFTP mode and does NOT reliably expand `$HOME` in destination paths) and a `$HOME`-prefixed form (for `ssh ... command` strings, where remote bash WILL expand `$HOME`):
```bash REMOTE_RUN_REL=".aris_queue/runs/$RUN_TS" # for scp destinations (relative to remote home) REMOTE_RUN_DIR="\$HOME/$REMOTE_RUN_REL" # for ssh command strings (literal $HOME, expanded on remote) ```
**3c. Bootstrap the remote run directory and copy helpers + manifest.** Per-invocation and idempotent. Use a unique run directory rather than `/tmp` so concurrent queues do not collide and so resume-after-crash is reproducible.
```bash ssh <server> "mkdir -p \"$REMOTE_RUN_DIR/logs\" \"\$HOME/.aris_queue\"" scp "$QUEUE_TOOLS/queue_manager.py" "$QUEUE_TOOLS/build_manifest.py" <server>:.aris_queue/ scp "$LOCAL_RUN_DIR/manifest.json" <server>:"$REMOTE_RUN_REL/manifest.json" ```
**3d. Launch the scheduler as a detached `nohup` process on the SSH host:**
```bash ssh <server> "nohup python3 \"\$HOME/.aris_queue/queue_manager.py\" \\ --manifest \"$REMOTE_RUN_DIR/manifest.json\" \\ --state \"$REMOTE_RUN_DIR/queue_state.json\" \\ --log-dir \"$REMOTE_RUN_DIR/logs\" \\ > \"$REMOTE_RUN_DIR/queue_mgr.log\" 2>&1 &" ```
Notes for callers: - `--log-dir` is what `queue_manager.py` actually consumes (per-job log files for OOM detection). Do NOT pass `--log <path>` — that flag is declared but unused, and a single combined log breaks the per-job stale-screen / OOM heuristics. - Persist `RUN_TS` / `REMOTE_RUN_REL` / `REMOTE_RUN_DIR` to disk so monitoring and resume can reload them without regenerating:
```bash { printf 'PROJECT_DIR=%q\n' "$PROJECT_DIR" printf 'RUN_TS=%q\n' "$RUN_TS" printf 'LOCAL_RUN_DIR=%q\n' "$LOCAL_RUN_DIR" printf 'REMOTE_RUN_REL=%q\n' "$REMOTE_RUN_REL" printf 'REMOTE_RUN_DIR=%q\n' "$REMOTE_RUN_DIR" } > "$LOCAL_RUN_DIR/run_meta.txt" ```
`%q` shell-escapes the values so the file is safely sourceable later. Note that `REMOTE_RUN_DIR` keeps a literal `$HOME` (do not expand it locally), which is the right form for re-use inside `ssh "..."` strings later.
**3e. Resume an existing queue (only when the user asks).** A fresh `RUN_TS` per invocation is correct for *new* queues. To resume a crashed queue, do NOT regenerate `RUN_TS` — reload the recorded values and re-run only the launch command (Step 3d), not the bootstrap (Step 3c):
```bash LOCAL_RUN_DIR="/abs/path/to/project/experiment_queue/<existing-run-ts>" # the run dir to resume . "$LOCAL_RUN_DIR/run_meta.txt" # reloads PROJECT_DIR / RUN_TS / REMOTE_RUN_REL / REMOTE_RUN_DIR # Then re-run Step 3d verbatim. Do NOT re-run Step 3c (would overwrite manifest.json + state.json). ```
A `queue_state.json` written before the 2026-08 scheduler fix records jobs the old code mis-judged a
Source provenance
Decision snapshot
15,660 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for experiment-queue, ready for a manual X post.
A practical pick for source-backed research: experiment-queue: SSH job queue for multi-seed/multi-config ML experiments with OOM-aware retry, stale-screen cleanup, and wave-transition ra... 15.7K stars https://www.openagentskill.com/skills/wanshuiyin-experiment-queue?ref=x
Listing + install path for experiment-queue: https://www.openagentskill.com/skills/wanshuiyin-experiment-queue?ref=x Install: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-q...
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to wanshuiyin but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-queue?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-queue?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-queue/audit)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-queue?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)wanshuiyin
@wanshuiyin
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
UI-TARS Desktop
Run multimodal agents that operate desktop interfaces
37.0K Starsn8n
Connect agents to hundreds of workflow automations
194.1K StarsMoneyPrinterTurbo
利用AI大模型,一键生成高清短视频 Generate short videos with one click using AI LLM.
88.5K StarsTasmota
Alternative firmware for ESP8266 and ESP32 based devices with easy configuration using webUI, OTA updates, automation using timers or rules, expandability and entirely local control over MQTT, HTTP, Serial or KNX. Full documentation at
24.7K StarsPermission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness