arbor

強い · 79
Registry に収録

Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many exper

Verified installs0
スター34.0K
バージョン1.0.0
品質92/100 · 優秀
信頼79/100 · レビュー後にインストール
監査89/100 · 試用可

供給アセットの概要

リサーチとナレッジ作業

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

カテゴリを見る

シナリオ

リサーチ Agent

I need my agent to research a topic, compare sources, and produce a concise report.

Agent 適合

Claude Code + CLI + Codex

Codex、Claude Code、Cursor、CLI、またはカスタム Agent に対応します。

インストール

準備完了

npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

メンテナンス

新しい

最終プッシュから 2 日

リスク

試用可

利用可能なメタデータに重大なリスクシグナルはありません

GitHub 品質

34K

92/100 品質 · 84/100 信頼

対象タグ

リサーチリサーチ Agentagent-skill

レビュー注記

利用可能なメタデータに重大なリスクシグナルはありません

Agent 導入スコアカード

信頼、監査、インストール準備状況を一目で確認

公開リポジトリのメタデータ、OpenAgentSkill のレビューシグナル、保守の鮮度、インストール準備状況を組み合わせたスコアです。候補選定の目安であり、人によるレビューの代替ではありません。

品質

優秀
92

採用度と保守性のシグナルが強い高信頼候補です。

信頼

レビュー後にインストール
79

有望な候補ですが、Agent は実行前に監査メモ、インストールポリシー、成果エビデンスを確認する必要があります。

監査

試用可
89

インストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。

OpenAgentSkill Trust Score v5

インストール前に人のレビュー

人による確認またはサンドボックス検証後に第一候補として使用します。

CodexClaude CodeCursorOpenAgentSkill CLI

スター

GitHub スター 34K

リポジトリ活動

スター 34K、フォーク 3.3K

メンテナンス

最終プッシュから 2 日

ライセンス

MIT license

インストール

npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

インストール安全性

標準パッケージまたはランタイムのインストールパス

権限範囲

shell or command execution, filesystem or document access

Agent の成果

Agent の成果データはまだありません

ドキュメント

Usable metadata, review docs

リスク概要

低いメタデータリスク

  • 利用可能なメタデータに重大な信頼警告はありません

インストール準備状況

インストールパスを利用可能

  • インストールパスを利用できます
  • リポジトリの根拠を利用できます
  • ライセンスが明示されています
  • Agent-Proven の成果エビデンスはまだありません

Agent 可読メタデータ

このスキルの機械可読な判断データ。

このブロックまたは埋め込み JSON を使い、Agent がこのスキルをインストールすべきか、代替を選ぶべきか、先に人のレビューを求めるべきかを判断できます。

JSON を開く

適したタスク

  • リサーチ Agent ワークフロー
  • Claude Code チーム
  • GitHub 採用シグナルを重視するチーム
  • 検索ソース

適した Agent

CodexClaude CodeCursorOpenAgentSkill CLICLI

インストール判断

コマンド
npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
ポリシー
レビュー
人によるレビュー
はい

信頼とリスク

信頼
79/100
監査
89/100
リスクレベル
試用可

成果ループ

エンドポイント
/api/agent/outcome
イベント ID
resolve
成果
5

インストールコマンド

npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

使わない場合

  • ベンダー提供の SLA が必要なチーム
  • 内部セキュリティレビューのない高コンプライアンス環境
  • 現在のメタデータに重大なリスクシグナルはありません
  • 高リスク権限のヒント: Shell またはコマンド実行
  • 利用可能なメタデータに重大な信頼警告はありません

Agent セーフティ v2

61/100 · インストール前にレビュー

権限メモ付きでレビュー済みレビュー

利用可能な候補ですが、Agent はインストール前に権限と監査メモを提示する必要があります。

実際のワークスペースへインストールする前に人の承認が必要です。

API で解決

Shell またはコマンド実行

Skill メタデータに端末、CLI、Shell、サブプロセス、またはコマンド実行のワークフローが含まれます。

ネットワークアクセス

Skill はリモートページ、API、リポジトリ、外部サービスにアクセスする可能性があります。

ファイルシステムアクセス

Skill はプロジェクトファイル、ドキュメント、生成物、ローカルワークスペース状態を読み書きする可能性があります。

  • 高リスク権限のヒント: Shell またはコマンド実行

インストール先

Agent ワークフローにこのスキルをインストール

公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install k-dense-ai-arbor

Agent 解決プラン

インストール前に Agent に適合性を検証させます。

Resolve API は第一候補、代替、安全ポリシー、監査メモ、インストール先、Agent がそのまま使えるプロンプトを返します。

テキストプランを開く

Agent が確認すべきこと

  • Resolve API でタスク適合と代替を確認。
  • 監査・信頼スコアと安全ポリシーの警告を確認。
  • Codex、Claude Code、Cursor、CLI のインストール先互換性を確認。

プロンプトをコピー

Task: Use arbor in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20arbor%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/k-dense-ai-arbor/install
Install command: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 引き継ぎ

別のディレクトリではなく、インストール経路を Agent に渡します。

公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。

Install API を開く

Agent プロンプト

Use arbor for this task. Review https://www.openagentskill.com/api/skills/k-dense-ai-arbor/install, then install with: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

Registry メタデータ

自動スキル選択用の Agent 可読プロファイル。

Registry API 経由で判断、信頼、監査、ユースケース、インストールのシグナルを提供し、UI をスクレイピングせずに Agent が順位付けできます。

Manifest を開く

Agent 適合

100/100

リサーチ Agent

プラットフォーム

Claude Code

監査レポート

試用可 · 89/100

インストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。

監査レポートを見る評価レポートを見る

Agent 判断パネル

リサーチ Agent 向けの第一候補

有力候補として扱い、自分の Agent スタックで README とインストール経路を検証してください。

100
準備状況
採用
段階

スタック内の役割

第一候補

主な適合

リサーチ Agent

信頼ラベル

本番対応

インストールパス

コマンド準備済み

使う場面

  • リサーチ Agent ワークフロー
  • Claude Code チーム
  • GitHub 採用シグナルを重視するチーム

根拠

  • GitHub スター 33,974
  • 最近のリポジトリ活動
  • インストールコマンドまたは GitHub リポジトリが利用可能
  • 品質プロファイル 92/100
  • OpenAgentSkill エンゲージメント 19 件

先にレビュー

  • 現在のメタデータに重大なリスクシグナルはありません

実装パス

  1. 1サンドボックスの Agent にインストールし、リサーチ Agent タスクを一度最初から最後まで実行します。
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

信頼プロファイル

レビュー後にインストール

有望な候補ですが、Agent は実行前に監査メモ、インストールポリシー、成果エビデンスを確認する必要があります。

79
OpenAgentSkill Trust Score

GitHub 採用度

合格

GitHub スター 34K

スター/フォーク活動

合格

スター 34K、フォーク 3.3K; 現在のメタデータでは Issue 活動を利用できません

最近のメンテナンス

合格

最終プッシュから 2 日

ライセンスの明確さ

合格

MIT license

良いシグナル

  • AI レビュー承認済み
  • インストールパスを利用できます
  • リポジトリの根拠を利用できます
  • 最近保守されたリポジトリ
  • Large GitHub adoption signal
  • インストールコマンドに明確な高リスクパターンはありません
  • 成果ループは準備済みですが、最初の実行が必要です

インストール前にレビュー

  • 実際の Agent 成果レポートはまだありません
  • 無人インストールの前に人によるレビューが必要です

推奨アクション

人による確認またはサンドボックス検証後に第一候補として使用します。

品質プロファイル

優秀 Agent ワークフロー向けの候補

採用度と保守性のシグナルが強い高信頼候補です。

92
GitHub スター
34K
鮮度
2 日前
インストール準備完了
はい
ライセンス
MIT license

ワークフロー適合

このスキルを使うシナリオ

ワークフロー適合

完全なワークフローに追加

代替候補

インストール前に比較

このタスクに適する可能性のある類似スキル。

すべて比較

概要

--- name: arbor description: Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md. allowed-tools: Read Write Edit Bash Agent license: MIT license metadata: version: "1.1" skill-author: K-Dense Inc. ---

# Arbor — Autonomous Optimization via Hypothesis Tree Refinement

## Overview

This skill runs an **Autonomous Optimization (AO)** loop: starting from an existing artifact and a measurable objective, improve it through many rounds of experiment and evaluation — without step-by-step human supervision and without overfitting to the feedback signal. It's the right tool when the bottleneck isn't writing one good change, but *organizing dozens of trials* so that lessons accumulate instead of evaporating.

It implements **Hypothesis Tree Refinement (HTR)** from *Arbor* (Jin et al., 2026). The key idea: keep the research state in a persistent **hypothesis tree** rather than in conversation history. Each node binds a hypothesis, the distilled insight it produced, and a pointer to the artifact version that realizes it. You play the long-lived **coordinator** that owns this tree and decides where to search; short-lived **executor** subagents test one hypothesis each in isolated git worktrees and report back. A **held-out merge gate** admits a change only when it improves on a *test* evaluator the search never optimized against. This is what turns trial-and-error into cumulative, auditable research.

Use the `scripts/tree.py` state manager for all the bookkeeping (creating nodes, writing evidence, propagating insights, pruning, the merge gate, the Observe projection). It keeps the state consistent and frees you to spend judgment on what the evidence *means*.

## When to use this skill

Reach for Arbor when the task is **iterative improvement of a concrete artifact under an evaluator**: - Model training: optimizer/architecture/recipe changes to lower loss or hit a target in fewer steps. - Harness/agent engineering: raising pass rate or accuracy of an agent loop, search harness, or tool-use scaffold. - Data synthesis: improving a generation/filtering pipeline judged by downstream model behavior. - Benchmark optimization: MLE-bench / Kaggle-style "improve the submission" tasks. - Prompt/system optimization where you can score outputs automatically.

The distinguishing signals: there's an **artifact you can modify**, an **objective**, a way to **score** candidates, and you expect to run **many experiments**. If the user only wants a single fix or a one-shot answer, this is overkill — just do the work directly. If they want open-ended ideation with no evaluator, use `hypothesis-generation` or `scientific-brainstorming` instead.

## The AO setup — pin this down first

Before any experiments, establish the task tuple `(M_0, O, E_dev, E_test)`. Getting this right matters more than any later decision, so confirm it explicitly:

- **M_0 — initial material**: the artifact to improve (a repo, a script, a config, a prompt). Make sure it's under git and currently runs. - **O — objective**: the natural-language goal and the metric *direction* (maximize accuracy? minimize loss/steps?). - **E_dev — development evaluator**: a command you can run freely during search to score a candidate. Fast, repeatable. - **E_test — held-out test evaluator**: a *separate* evaluator (different seeds, different split, or a larger run) used only at the merge gate. It must not be used as a search oracle — that's the whole point.

If the user hasn't given you a clean dev/test split, **construct one and say so**. The dev/test separation is the mechanism that catches overfitting: a candidate that wins on dev but not on test isn't a success, it's a warning that you're exploiting the feedback signal. Without it, autonomous search reliably overfits.

Initialize the run:

```bash python scripts/tree.py init \ --objective "Improve BrowseComp answer accuracy on the search harness" \ --dev-eval "python eval.py --split dev --n 50" \ --test-eval "python eval.py --split test --n 300" \ --material "." --metric-direction max --branching 3 --max-depth 2 --budget 12 ```

`--branching` is how many sibling hypotheses you propose per parent; `--max-depth 2` keeps directions at depth 1 and concrete interventions at depth 2 (the paper's default); `--budget` is the number of coordinator cycles. Start small (10–20 cycles) — structured search beats brute force, and you can extend if progress is still being made.

## The coordinator loop

You run repeated cycles of six steps. This is the heart of HTR; do not collapse it into ad-hoc editing. Run `python scripts/tree.py cycle` once per cycle to track the budget.

### 1. Observe Begin every cycle by re-grounding in the tree, not in your memory of the conversation:

```bash python scripts/tree.py observe ```

This prints the objective, global insights, the active frontier (selectable hypotheses), executed nodes with their evidence, pruned lessons (negative constraints), and the current best artifact. Treating the tree as the source of truth is what keeps you coherent over a long run, after context compression has thrown away the details.

### 2. Ideate Pick a promising parent and propose a few child hypotheses under it. **Condition on the tree's evidence** — this is the difference between Arbor and random search: - Validated insights are assumptions you can build on. - Pruned nodes are dead ends to avoid. - A "half-right" result is a *starting point for a sharper hypothesis*, not a reason to abandon the direction.

Each hypothesis should be a **falsifiable claim about how changing the artifact will move the metric**, not a vague intention. Depth-1 nodes are broad directions ("the search harness loses correct answers it already retrieved"); depth-2 nodes are concrete, executable interventions ("run K=5 independent rollouts and aggregate by evidence dossier instead of majority vote").

```bash python scripts/tree.py add-node --parent n0 --hypothesis "Verification, not retrieval, is the bottleneck: candidates are found but discarded" python scripts/tree.py add-node --parent n4 --hypothesis "Decompose the question into atomic constraints and verify each independently" ```

### 3. Select Choose which pending leaves to run next. **Selection is not pure score-maximization** — pick a hypothesis because it has strong prior evidence, because it would resolve an ambiguity its siblings exposed, or because its failure would clarify an important assumption. Frontier control under delayed feedback rewards informative experiments, not just promising ones.

### 4. Dispatch Run each selected hypothesis as an **executor subagent in an isolated worktree** (use the Agent tool with `isolation: "worktree"`, or have the executor create one with `git worktree add`). Isolation matters: parallel experiments must not clobber each other or the current best, and exploratory changes stay quarantined until they pass the merge gate.

Dispatch siblings **in parallel** (multiple Agent calls in one message) when they're independent — comparative evidence within one direction is exactly what makes later pruning and abstraction possible.

Give each executor a tight, **hypothesis-bound** brief. See `references/executor-brief.md` for the full template. The contract that makes HTR work: **the executor may not change the hypothesis when the metric stalls.** It repairs its own code and reruns, but `h_n` is fixed — otherwise the returned score is no longer evidence about the assigned node and the tree's semantics break. The executor returns exactly four things: - **dev_score** — the dev evaluator result (for selection); - **result** — a factual summary of what happened; - **insight** — the distilled, reusable lesson (*why* the result supports, weakens, or bounds the hypothesis); - **branch_ref** — the git branch/commit/worktree path holding the artifact.

Mark a node `running` before dispatch (`tree.py set-status --node n5 --status running`) so the Observe projection stays accurate.

### 5. Backpropagate When an executor returns, write its report into the node, then **abstract the lesson upward**:

```bash python scripts/tree.py set-evidence --node n5 --dev-score 70.0 \ --result "K=5 dossier aggregation recovers answers in minority rollouts" \ --insight "Correct answers often appear in a minority of rollouts; aggregation beats majority vote" \ --branch-ref "wt/n5"

python scripts/tree.py propagate --node n5 \ --insight "Candidate coverage, not verification, limits this direction" --to-root ```

This is the step that makes the tree more than a log. A leaf-level observation ("data-interface mismatch") should become a direction-level constraint and, if it generalizes, a global prior that shapes future ideation. **Insight propagation is the component that drives most of HTR's gains** — in the paper's MLE-Bench Lite ablation, a tree *without* insight feedback scored even lower than a flat experiment queue with no tree at all (54.5% vs. 63.6% any-medal, against 81.8% for the full system). Hierarchy alone isn't enough: the semantic memory is what matters. So spend real thought on the abstraction; don't just copy the leaf insight upward verbatim.

### 6. Decide Decide what to do with the new evidence: keep expanding a direction, prune a falsified subtree, or attempt to merge a candidate.

- **Prune** dead ends, recording *why* — the reason becomes a negative constraint: ```bash python scripts/tree.py prune --node n7 --reason "search-augmented judge overfits dev questions; no test transfer" ``` - **Merge gate** — promote a candidate to the new best **only if it improves on `E_test`**. Run the test evaluator in a *fresh* worktree (not the dev worktree, to avoid leakage), then: ```bash python scripts/tree.py merge --node n5 --test-score 67.67 --branch-ref "wt/n5" ``` If the gate rejects it, that's informative: a high-dev / low-test candidate is evidence the direction may be exploiting the dev signal rather than producing a transferable improvement. Record that lesson; don't quietly promote it anyway.

Repeat until the budget is spent, the frontier is exhausted, or progress has clearly stalled.

## Finishing the run

When you stop, produce a short report (see `references/report-template.md`) covering: - the final best artifact, its test score, and its delta over `M_0`; - the tree (`python scripts/tree.py status`) as the audit trail of what was tried; - the main hypothesis shifts — how task understanding deepened across the run (early nodes test broad mechanisms; later nodes find their limits; ancestor insights compress these into the constraints behind the final design); - merged vs. explored: many nodes improve dev, far fewer pass the test gate — report that gap honestly rather than overstating dev wins.

Always leave `M_best` as a real, runnable artifact on a named branch, and tell the user how to check it out.

## Principles that make this work (not rote rules)

These come from the paper's analysis; understanding *why* matters more than following them mechanically.

- **The tree is the memory; conversatio

技術詳細

バージョン
1.0.0
ライセンス
MIT license
最終更新
2026年8月20日
公開日
2026年8月20日

判断の要約

第一候補

100
準備完了
採用
段階

GitHub スター 33,974

監査

インストールレビュー

インストールと採用のレビュー

89
試用可
セキュリティ
83/100
メンテナンス
100/100
インストール
92/100
完全な監査を開く評価レポートを見る

Agent 実証エビデンス

Agent 実証エビデンス

Resolve、レビュー、インストール、限定実行後の成果レポート。

0
実証済み
Needs first agent run自動インストール: 先にレビュー最新: 不明
成功率
直近の失敗
成果
0
出力品質
失敗
0
非該当
0
インストール数
0
リスクによりブロック
0
設定が必要
0
本番
0

Agent の実行結果はまだありません。最初の実行では /api/agent/outcome を通じて成功、設定要件、リスクによるブロック、失敗、非該当を報告できます。

インストール

Agent ワークフローに追加

無料・オープンソース. 本番 Agent にインストールする前にレポートを確認してください。

成長ループ

共有キット

X

arbor 用のシナリオベース草案です。X へ手動投稿できます。

キュレーターノート
A practical pick for source-backed research:

arbor: Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and...

34.0K stars

https://www.openagentskill.com/skills/k-dense-ai-arbor?ref=x
X 下書きを開く
任意:インストールコマンド付きの返信
Listing + install path for arbor:
https://www.openagentskill.com/skills/k-dense-ai-arbor?ref=x

Install: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
返信の下書きを開く

掲載元

Registry により登録

申請可能

この掲載は公開ソースから登録されており、メンテナー申請が承認されるまで公式として表示されません。

作成者
K-Dense-AI
インデックス作成者
OpenAgentSkill コミュニティインデックス

帰属は公開リポジトリまたは作成者プロフィールにリンクされています。作成者は掲載を申請して所有権シグナルを更新できます。

このスキルを申請

所有者の申請

このスキル掲載を申請

この Registry により登録 掲載は K-Dense-AI に帰属していますが、まだ公式として表示されていません。申請すると、確認済み所有者シグナルが追加され、今後の公開、インストール、監査更新の信頼性が高まります。

クリエイター被リンクキット

README にエビデンスバッジを追加

開発者がリポジトリを評価する場所で、正規掲載、現在の信頼・監査シグナル、実際の Agent-Proven エビデンスを表示します。

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=listed&label=Listed)](https://www.openagentskill.com/skills/k-dense-ai-arbor)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=trust&label=Trust)](https://www.openagentskill.com/skills/k-dense-ai-arbor)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=audit&label=Audit)](https://www.openagentskill.com/skills/k-dense-ai-arbor/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/k-dense-ai-arbor)

作者

K

K-Dense-AI

@k-dense-ai

プラットフォーム適合

健全性シグナル

GitHub スター
34.0K
品質スコア
55/100
最終 GitHub プッシュ
2026年8月20日
フレームワークのヒント
不明
OpenAgentSkill 閲覧数
19
インストールコピー数
0
外部クリック
0

コミュニティシグナル

このスキルが Agent ワークフローに役立つかを共有してください。集約されたフィードバックがランキングを改善します。

信頼と安全性

レビュー後にインストール

79
  • GitHub 採用度GitHub スター 34K合格
  • スター/フォーク活動スター 34K、フォーク 3.3K; 現在のメタデータでは Issue 活動を利用できません合格
  • 最近のメンテナンス最終プッシュから 2 日合格
  • ライセンスの明確さMIT license合格
  • README/SKILL.md の完全性公開メタデータにはより十分な README/SKILL.md の文脈が必要です情報
  • 依存関係/ランタイムのリスクコマンド実行範囲情報