arbor
Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many exper
Supply asset profile
Research and knowledge work
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
Maintenance
fresh
2d since push
Risk
Safe to try
No major risk signals from available metadata
GitHub quality
34K
92/100 Quality · 84/100 Trust
Coverage tags
Review notes
No major risk signals from available metadata
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Review then installGood shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
Audit
Safe to tryA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Use as the primary candidate after human or sandbox review.
Stars
34K GitHub stars
Repo activity
34K stars, 3.3K forks
Maintenance
2d since push
License
MIT license
Install
npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
Install safety
standard package or runtime install path
Permission surface
shell or command execution, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Low metadata risk
- No major trust warnings detected from available metadata
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- Research agents workflows
- Claude Code teams
- teams that value GitHub adoption signals
- Search sources
Suited agents
Install decision
- Command
- npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
- Policy
- review
- Human review
- yes
Trust and risk
- Trust
- 79/100
- Audit
- 89/100
- Risk level
- Safe to try
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add K-Dense-AI/scientific-agent-skills --skill arborDo not use when
- teams that need a vendor-supported SLA
- high-compliance environments without internal security review
- No major risk signals from current metadata
- High-risk permission hints: Shell or command execution
- No major trust warnings detected from available metadata
Alternative
Last30days Skill
53.5K Stars
npx skills add mvanhorn/last30days-skill -g
Alternative
Academic Research Skills
38.4K Stars
npx skills add Imbad0202/academic-research-skills
Alternative
GPT Researcher
28.0K Stars
npx skills add assafelovic/gpt-researcher
Alternative
DeepResearch
19.8K Stars
npx skills add Alibaba-NLP/DeepResearch
Agent safety v2
61/100 · Review before install
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
high
Shell or command execution
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Filesystem access
Skill may read or write project files, documents, generated artifacts, or local workspace state.
- High-risk permission hints: Shell or command execution
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install k-dense-ai-arborAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20arbor%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20arbor%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/k-dense-ai-arbor/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use arbor in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20arbor%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/k-dense-ai-arbor/install
Install command: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/k-dense-ai-arbor/install
LLM text format
/api/skills/k-dense-ai-arbor/install?format=text
Find alternatives
/api/skills/search?q=arbor&limit=3
Agent prompt
Use arbor for this task. Review https://www.openagentskill.com/api/skills/k-dense-ai-arbor/install, then install with: npx skills add K-Dense-AI/scientific-agent-skills --skill arborRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/k-dense-ai-arbor
LLM text
/api/registry/manifest/k-dense-ai-arbor?format=text
Install alias
/api/registry/install/k-dense-ai-arbor
Recommend
/api/registry/recommend?task=Use%20arbor%20in%20an%20agent%20workflow&limit=3
Agent fit
Research agents
Use-case tags
Platforms
Claude Code
Audit report
Safe to try · 89/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Primary pick for Research agents
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Research agents
Trust label
Production-ready
Install path
Command ready
Use when
- Research agents workflows
- Claude Code teams
- teams that value GitHub adoption signals
Evidence
- 33,974 GitHub stars
- recent repository activity
- install command or GitHub repo available
- 92/100 quality profile
- 19 OpenAgentSkill engagement events
review first
- No major risk signals from current metadata
Implementation path
- 1Install it in a sandbox agent and run one Research agents task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Review then install
Good shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
GitHub adoption
PASS34K GitHub stars
Stars/forks activity
PASS34K stars, 3.3K forks; issue activity unavailable in current metadata
Recent maintenance
PASS2d since push
License clarity
PASSMIT license
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Large GitHub adoption signal
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Use as the primary candidate after human or sandbox review.
Quality profile
Excellent candidate for agent workflows
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Use this skill in these scenarios
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Search private knowledge
RAG and knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Operate web apps
Browser automation
I need my agent to control a browser, fill forms, and verify web app workflows.
Workflow fit
Add it to a complete workflow
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Ingest, retrieve, and cite
RAG knowledge base
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
Academic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
GPT Researcher
Run autonomous deep research over web and local sources
DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Overview
--- name: arbor description: Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md. allowed-tools: Read Write Edit Bash Agent license: MIT license metadata: version: "1.1" skill-author: K-Dense Inc. ---
# Arbor — Autonomous Optimization via Hypothesis Tree Refinement
## Overview
This skill runs an **Autonomous Optimization (AO)** loop: starting from an existing artifact and a measurable objective, improve it through many rounds of experiment and evaluation — without step-by-step human supervision and without overfitting to the feedback signal. It's the right tool when the bottleneck isn't writing one good change, but *organizing dozens of trials* so that lessons accumulate instead of evaporating.
It implements **Hypothesis Tree Refinement (HTR)** from *Arbor* (Jin et al., 2026). The key idea: keep the research state in a persistent **hypothesis tree** rather than in conversation history. Each node binds a hypothesis, the distilled insight it produced, and a pointer to the artifact version that realizes it. You play the long-lived **coordinator** that owns this tree and decides where to search; short-lived **executor** subagents test one hypothesis each in isolated git worktrees and report back. A **held-out merge gate** admits a change only when it improves on a *test* evaluator the search never optimized against. This is what turns trial-and-error into cumulative, auditable research.
Use the `scripts/tree.py` state manager for all the bookkeeping (creating nodes, writing evidence, propagating insights, pruning, the merge gate, the Observe projection). It keeps the state consistent and frees you to spend judgment on what the evidence *means*.
## When to use this skill
Reach for Arbor when the task is **iterative improvement of a concrete artifact under an evaluator**: - Model training: optimizer/architecture/recipe changes to lower loss or hit a target in fewer steps. - Harness/agent engineering: raising pass rate or accuracy of an agent loop, search harness, or tool-use scaffold. - Data synthesis: improving a generation/filtering pipeline judged by downstream model behavior. - Benchmark optimization: MLE-bench / Kaggle-style "improve the submission" tasks. - Prompt/system optimization where you can score outputs automatically.
The distinguishing signals: there's an **artifact you can modify**, an **objective**, a way to **score** candidates, and you expect to run **many experiments**. If the user only wants a single fix or a one-shot answer, this is overkill — just do the work directly. If they want open-ended ideation with no evaluator, use `hypothesis-generation` or `scientific-brainstorming` instead.
## The AO setup — pin this down first
Before any experiments, establish the task tuple `(M_0, O, E_dev, E_test)`. Getting this right matters more than any later decision, so confirm it explicitly:
- **M_0 — initial material**: the artifact to improve (a repo, a script, a config, a prompt). Make sure it's under git and currently runs. - **O — objective**: the natural-language goal and the metric *direction* (maximize accuracy? minimize loss/steps?). - **E_dev — development evaluator**: a command you can run freely during search to score a candidate. Fast, repeatable. - **E_test — held-out test evaluator**: a *separate* evaluator (different seeds, different split, or a larger run) used only at the merge gate. It must not be used as a search oracle — that's the whole point.
If the user hasn't given you a clean dev/test split, **construct one and say so**. The dev/test separation is the mechanism that catches overfitting: a candidate that wins on dev but not on test isn't a success, it's a warning that you're exploiting the feedback signal. Without it, autonomous search reliably overfits.
Initialize the run:
```bash python scripts/tree.py init \ --objective "Improve BrowseComp answer accuracy on the search harness" \ --dev-eval "python eval.py --split dev --n 50" \ --test-eval "python eval.py --split test --n 300" \ --material "." --metric-direction max --branching 3 --max-depth 2 --budget 12 ```
`--branching` is how many sibling hypotheses you propose per parent; `--max-depth 2` keeps directions at depth 1 and concrete interventions at depth 2 (the paper's default); `--budget` is the number of coordinator cycles. Start small (10–20 cycles) — structured search beats brute force, and you can extend if progress is still being made.
## The coordinator loop
You run repeated cycles of six steps. This is the heart of HTR; do not collapse it into ad-hoc editing. Run `python scripts/tree.py cycle` once per cycle to track the budget.
### 1. Observe Begin every cycle by re-grounding in the tree, not in your memory of the conversation:
```bash python scripts/tree.py observe ```
This prints the objective, global insights, the active frontier (selectable hypotheses), executed nodes with their evidence, pruned lessons (negative constraints), and the current best artifact. Treating the tree as the source of truth is what keeps you coherent over a long run, after context compression has thrown away the details.
### 2. Ideate Pick a promising parent and propose a few child hypotheses under it. **Condition on the tree's evidence** — this is the difference between Arbor and random search: - Validated insights are assumptions you can build on. - Pruned nodes are dead ends to avoid. - A "half-right" result is a *starting point for a sharper hypothesis*, not a reason to abandon the direction.
Each hypothesis should be a **falsifiable claim about how changing the artifact will move the metric**, not a vague intention. Depth-1 nodes are broad directions ("the search harness loses correct answers it already retrieved"); depth-2 nodes are concrete, executable interventions ("run K=5 independent rollouts and aggregate by evidence dossier instead of majority vote").
```bash python scripts/tree.py add-node --parent n0 --hypothesis "Verification, not retrieval, is the bottleneck: candidates are found but discarded" python scripts/tree.py add-node --parent n4 --hypothesis "Decompose the question into atomic constraints and verify each independently" ```
### 3. Select Choose which pending leaves to run next. **Selection is not pure score-maximization** — pick a hypothesis because it has strong prior evidence, because it would resolve an ambiguity its siblings exposed, or because its failure would clarify an important assumption. Frontier control under delayed feedback rewards informative experiments, not just promising ones.
### 4. Dispatch Run each selected hypothesis as an **executor subagent in an isolated worktree** (use the Agent tool with `isolation: "worktree"`, or have the executor create one with `git worktree add`). Isolation matters: parallel experiments must not clobber each other or the current best, and exploratory changes stay quarantined until they pass the merge gate.
Dispatch siblings **in parallel** (multiple Agent calls in one message) when they're independent — comparative evidence within one direction is exactly what makes later pruning and abstraction possible.
Give each executor a tight, **hypothesis-bound** brief. See `references/executor-brief.md` for the full template. The contract that makes HTR work: **the executor may not change the hypothesis when the metric stalls.** It repairs its own code and reruns, but `h_n` is fixed — otherwise the returned score is no longer evidence about the assigned node and the tree's semantics break. The executor returns exactly four things: - **dev_score** — the dev evaluator result (for selection); - **result** — a factual summary of what happened; - **insight** — the distilled, reusable lesson (*why* the result supports, weakens, or bounds the hypothesis); - **branch_ref** — the git branch/commit/worktree path holding the artifact.
Mark a node `running` before dispatch (`tree.py set-status --node n5 --status running`) so the Observe projection stays accurate.
### 5. Backpropagate When an executor returns, write its report into the node, then **abstract the lesson upward**:
```bash python scripts/tree.py set-evidence --node n5 --dev-score 70.0 \ --result "K=5 dossier aggregation recovers answers in minority rollouts" \ --insight "Correct answers often appear in a minority of rollouts; aggregation beats majority vote" \ --branch-ref "wt/n5"
python scripts/tree.py propagate --node n5 \ --insight "Candidate coverage, not verification, limits this direction" --to-root ```
This is the step that makes the tree more than a log. A leaf-level observation ("data-interface mismatch") should become a direction-level constraint and, if it generalizes, a global prior that shapes future ideation. **Insight propagation is the component that drives most of HTR's gains** — in the paper's MLE-Bench Lite ablation, a tree *without* insight feedback scored even lower than a flat experiment queue with no tree at all (54.5% vs. 63.6% any-medal, against 81.8% for the full system). Hierarchy alone isn't enough: the semantic memory is what matters. So spend real thought on the abstraction; don't just copy the leaf insight upward verbatim.
### 6. Decide Decide what to do with the new evidence: keep expanding a direction, prune a falsified subtree, or attempt to merge a candidate.
- **Prune** dead ends, recording *why* — the reason becomes a negative constraint: ```bash python scripts/tree.py prune --node n7 --reason "search-augmented judge overfits dev questions; no test transfer" ``` - **Merge gate** — promote a candidate to the new best **only if it improves on `E_test`**. Run the test evaluator in a *fresh* worktree (not the dev worktree, to avoid leakage), then: ```bash python scripts/tree.py merge --node n5 --test-score 67.67 --branch-ref "wt/n5" ``` If the gate rejects it, that's informative: a high-dev / low-test candidate is evidence the direction may be exploiting the dev signal rather than producing a transferable improvement. Record that lesson; don't quietly promote it anyway.
Repeat until the budget is spent, the frontier is exhausted, or progress has clearly stalled.
## Finishing the run
When you stop, produce a short report (see `references/report-template.md`) covering: - the final best artifact, its test score, and its delta over `M_0`; - the tree (`python scripts/tree.py status`) as the audit trail of what was tried; - the main hypothesis shifts — how task understanding deepened across the run (early nodes test broad mechanisms; later nodes find their limits; ancestor insights compress these into the constraints behind the final design); - merged vs. explored: many nodes improve dev, far fewer pass the test gate — report that gap honestly rather than overstating dev wins.
Always leave `M_best` as a real, runnable artifact on a named branch, and tell the user how to check it out.
## Principles that make this work (not rote rules)
These come from the paper's analysis; understanding *why* matters more than following them mechanically.
- **The tree is the memory; conversatio
Technical details
- Version
- 1.0.0
- License
- MIT license
- Last updated
- Aug 20, 2026
- Published
- Aug 20, 2026
Decision snapshot
Primary pick
33,974 GitHub stars
Audit
Install review
Install and adoption review
- Security
- 83/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for arbor, ready for a manual X post.
A practical pick for source-backed research: arbor: Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and... 34.0K stars https://www.openagentskill.com/skills/k-dense-ai-arbor?ref=x
Optional reply with install command
Listing + install path for arbor: https://www.openagentskill.com/skills/k-dense-ai-arbor?ref=x Install: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
Listing source
Registry indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- K-Dense-AI
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Registry indexed listing is attributed to K-Dense-AI but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/k-dense-ai-arbor)
[](https://www.openagentskill.com/skills/k-dense-ai-arbor)
[](https://www.openagentskill.com/skills/k-dense-ai-arbor/audit)
[](https://www.openagentskill.com/skills/k-dense-ai-arbor)Author
K-Dense-AI
@k-dense-ai
Tags
Platform fit
Health signals
- GitHub stars
- 34.0K
- Quality score
- 55/100
- Last GitHub push
- Aug 20, 2026
- Framework hints
- Unknown
- OpenAgentSkill views
- 19
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Review then install
- GitHub adoption34K GitHub starsPASS
- Stars/forks activity34K stars, 3.3K forks; issue activity unavailable in current metadataPASS
- Recent maintenance2d since pushPASS
- License clarityMIT licensePASS
- README/SKILL.md completenessPublic metadata needs stronger README/SKILL.md contextINFO
- Dependency/runtime riskcommand execution surfaceINFO
Related skills
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
53.5K StarsAcademic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
38.4K StarsGPT Researcher
Run autonomous deep research over web and local sources
28.0K StarsDeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
19.8K Stars