arbor

Sólido · 79
Indexado en Registry

Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many exper

Verified installs0
Estrellas34.0K
Versión1.0.0
Calidad92/100 · Excelente
Confianza79/100 · Revisar antes de instalar
Auditoría89/100 · Seguro para probar

Perfil del activo

Investigación y trabajo de conocimiento

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

Ver categoría

Escenario

Agents de investigación

I need my agent to research a topic, compare sources, and produce a concise report.

Afinidad con Agent

Claude Code + CLI + Codex

Funciona con Codex, Claude Code, Cursor, CLI o Agents personalizados.

Instalar

Listo

npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

Mantenimiento

Actual

3 días desde el último push

Riesgo

Seguro para probar

No major risk signals from available metadata

Calidad de GitHub

34K

92/100 Calidad · 84/100 Confianza

Etiquetas de cobertura

InvestigaciónAgents de investigaciónagent-skill

Notas de revisión

No major risk signals from available metadata

Tarjeta de adopción del Agent

Confianza, auditoría y preparación de instalación de un vistazo

Estas puntuaciones combinan metadatos públicos del repositorio, señales de revisión de OpenAgentSkill, actualidad de mantenimiento y preparación de instalación. Sirven para preseleccionar; no sustituyen la revisión humana.

Calidad

Excelente
92

High-confidence pick with strong adoption and healthy maintenance signals.

Confianza

Revisar antes de instalar
79

Buena señal para la preselección, pero el Agent debe revisar las notas de auditoría, la política de instalación y la evidencia de resultados antes de ejecutarlo.

Auditoría

Seguro para probar
89

Revisión legible por máquina de la preparación de instalación, los metadatos de seguridad, el mantenimiento y el riesgo de adopción.

Trust Score de OpenAgentSkill v5

Revisión humana antes de instalar

Úsalo como candidato principal tras revisión humana o en sandbox.

CodexClaude CodeCursorOpenAgentSkill CLI

Estrellas

34K estrellas de GitHub

Actividad del repositorio

34K estrellas y 3.3K forks

Mantenimiento

3 días desde el último push

Licencia

MIT license

Instalar

npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

Seguridad de instalación

Ruta estándar de paquete o instalación en tiempo de ejecución

Superficie de permisos

shell or command execution, filesystem or document access

Resultados del Agent

Aún no hay datos de resultados del Agent

Documentación

Usable metadata, review docs

Resumen de riesgo

Riesgo de metadatos bajo

  • No major trust warnings detected from available metadata

Preparación de instalación

Ruta de instalación disponible

  • La ruta de instalación está disponible
  • La evidencia del repositorio está disponible
  • La licencia está declarada
  • Aún no hay evidencia de resultados Agent-Proven

Metadatos legibles por Agent

Datos de decisión legibles por máquina para este skill.

Usa este bloque o el JSON integrado para decidir si un Agent debe instalar este skill, elegir una alternativa o pedir revisión humana primero.

Abrir JSON

Tareas adecuadas

  • Flujos de Agents de investigación
  • Equipos de Claude Code
  • Equipos que valoran señales de adopción de GitHub
  • Fuentes de búsqueda

Agents adecuados

CodexClaude CodeCursorOpenAgentSkill CLICLI

Decisión de instalación

Comando
npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
Política
Revisar
Revisión humana

Confianza y riesgo

Confianza
79/100
Auditoría
89/100
Nivel de riesgo
Seguro para probar

Ciclo de resultados

Endpoint
/api/agent/outcome
ID del evento
resolve
Resultados
5

Comando de instalación

npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

No usar cuando

  • Equipos que necesitan un SLA con soporte del proveedor
  • Entornos de alta conformidad sin revisión interna de seguridad
  • No major risk signals from current metadata
  • Indicios de permisos de alto riesgo: ejecución de shell o comandos
  • No major trust warnings detected from available metadata

Seguridad de Agent v2

61/100 · Revisar antes de instalar

Revisado con notas de permisosRevisar

Candidato utilizable, pero el Agent debe mostrar las notas de permisos y auditoría antes de instalar.

Requiere aprobación humana antes de instalar en un espacio de trabajo real.

Resolver con API

Alto

Ejecución de shell o comandos

Los metadatos del skill hacen referencia a terminal, CLI, shell, subprocesos o flujos de ejecución de comandos.

Medio

Acceso a red

El skill probablemente consulta páginas remotas, API, repositorios o servicios externos.

Medio

Acceso al sistema de archivos

El skill puede leer o escribir archivos de proyecto, documentos, artefactos generados o estado local.

  • Indicios de permisos de alto riesgo: ejecución de shell o comandos

Destinos de instalación

Instala este skill en tu flujo de Agent

Usa el endpoint público para obtener el comando, la lista de seguridad, prompts y enlaces canónicos.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install k-dense-ai-arbor

Plan de resolución de Agent

Deja que un Agent valide el ajuste antes de instalar.

La API Resolve devuelve la skill elegida, alternativas, política de seguridad, notas de auditoría, destino de instalación y un prompt listo para usar.

Abrir plan de texto

Agent debe revisar

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Copiar prompt

Task: Use arbor in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20arbor%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/k-dense-ai-arbor/install
Install command: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Traspaso de Agent

Da al Agent la ruta de instalación, no otro directorio.

Usa el endpoint público para obtener el comando, la lista de seguridad, prompts y enlaces canónicos.

Abrir API de instalación

Prompt de Agent

Use arbor for this task. Review https://www.openagentskill.com/api/skills/k-dense-ai-arbor/install, then install with: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

Metadatos del Registry

Perfil legible por Agent para seleccionar skills automáticamente.

La API Registry expone señales de decisión, confianza, auditoría, casos de uso e instalación sin raspar la interfaz.

Abrir Manifest

Afinidad con Agent

100/100

Agents de investigación

Plataformas

Claude Code

Informe de auditoría

Seguro para probar · 89/100

Revisión legible por máquina de la preparación de instalación, los metadatos de seguridad, el mantenimiento y el riesgo de adopción.

Ver informe de auditoríaVer informe de evaluación

Panel de decisión de Agent

Elección principal para Agents de investigación

Use this as a leading candidate, then validate the README and install path in your own agent stack.

100
Preparación
Adoptar
Etapa

Rol en la pila

Elección principal

Ajuste principal

Agents de investigación

Etiqueta de confianza

Listo para producción

Ruta de instalación

Comando listo

Úsalo cuando

  • Flujos de Agents de investigación
  • Equipos de Claude Code
  • Equipos que valoran señales de adopción de GitHub

Evidencia

  • 33,974 estrellas de GitHub
  • recent repository activity
  • install command or GitHub repo available
  • perfil de calidad 92/100
  • 19 eventos de interacción de OpenAgentSkill

revisar primero

  • No major risk signals from current metadata

Ruta de implementación

  1. 1Instálalo en un Agent de sandbox y ejecuta una tarea de Agents de investigación de principio a fin.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Perfil de confianza

Revisar antes de instalar

Buena señal para la preselección, pero el Agent debe revisar las notas de auditoría, la política de instalación y la evidencia de resultados antes de ejecutarlo.

79
Trust Score de OpenAgentSkill

Adopción en GitHub

Aprobado

34K estrellas de GitHub

Actividad de stars/forks

Aprobado

34K estrellas y 3.3K forks; la actividad de issues no está disponible en los metadatos actuales

Mantenimiento reciente

Aprobado

3 días desde el último push

Claridad de licencia

Aprobado

MIT license

Señales positivas

  • Revisión de IA aprobada
  • La ruta de instalación está disponible
  • La evidencia del repositorio está disponible
  • Repositorio mantenido recientemente
  • Large GitHub adoption signal
  • El comando de instalación no muestra un patrón de alto riesgo evidente
  • El ciclo de resultados está listo, pero necesita la primera ejecución real de Agent

Revisar antes de instalar

  • Aún no hay informes reales de resultados del Agent
  • Se requiere revisión humana antes de una instalación desatendida

Acción recomendada

Úsalo como candidato principal tras revisión humana o en sandbox.

Perfil de calidad

Excelente candidato para flujos de Agent

High-confidence pick with strong adoption and healthy maintenance signals.

92
Estrellas de GitHub
34K
Actualidad
hace 3 días
Listo para instalar
Licencia
MIT license

Ajuste de flujo

Usa esta skill en estos escenarios

Ajuste de flujo

Añadir a un flujo completo

Lista de alternativas

Compara antes de instalar

Similar skills that may fit this task.

Comparar todo

Resumen

--- name: arbor description: Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md. allowed-tools: Read Write Edit Bash Agent license: MIT license metadata: version: "1.1" skill-author: K-Dense Inc. ---

# Arbor — Autonomous Optimization via Hypothesis Tree Refinement

## Overview

This skill runs an **Autonomous Optimization (AO)** loop: starting from an existing artifact and a measurable objective, improve it through many rounds of experiment and evaluation — without step-by-step human supervision and without overfitting to the feedback signal. It's the right tool when the bottleneck isn't writing one good change, but *organizing dozens of trials* so that lessons accumulate instead of evaporating.

It implements **Hypothesis Tree Refinement (HTR)** from *Arbor* (Jin et al., 2026). The key idea: keep the research state in a persistent **hypothesis tree** rather than in conversation history. Each node binds a hypothesis, the distilled insight it produced, and a pointer to the artifact version that realizes it. You play the long-lived **coordinator** that owns this tree and decides where to search; short-lived **executor** subagents test one hypothesis each in isolated git worktrees and report back. A **held-out merge gate** admits a change only when it improves on a *test* evaluator the search never optimized against. This is what turns trial-and-error into cumulative, auditable research.

Use the `scripts/tree.py` state manager for all the bookkeeping (creating nodes, writing evidence, propagating insights, pruning, the merge gate, the Observe projection). It keeps the state consistent and frees you to spend judgment on what the evidence *means*.

## When to use this skill

Reach for Arbor when the task is **iterative improvement of a concrete artifact under an evaluator**: - Model training: optimizer/architecture/recipe changes to lower loss or hit a target in fewer steps. - Harness/agent engineering: raising pass rate or accuracy of an agent loop, search harness, or tool-use scaffold. - Data synthesis: improving a generation/filtering pipeline judged by downstream model behavior. - Benchmark optimization: MLE-bench / Kaggle-style "improve the submission" tasks. - Prompt/system optimization where you can score outputs automatically.

The distinguishing signals: there's an **artifact you can modify**, an **objective**, a way to **score** candidates, and you expect to run **many experiments**. If the user only wants a single fix or a one-shot answer, this is overkill — just do the work directly. If they want open-ended ideation with no evaluator, use `hypothesis-generation` or `scientific-brainstorming` instead.

## The AO setup — pin this down first

Before any experiments, establish the task tuple `(M_0, O, E_dev, E_test)`. Getting this right matters more than any later decision, so confirm it explicitly:

- **M_0 — initial material**: the artifact to improve (a repo, a script, a config, a prompt). Make sure it's under git and currently runs. - **O — objective**: the natural-language goal and the metric *direction* (maximize accuracy? minimize loss/steps?). - **E_dev — development evaluator**: a command you can run freely during search to score a candidate. Fast, repeatable. - **E_test — held-out test evaluator**: a *separate* evaluator (different seeds, different split, or a larger run) used only at the merge gate. It must not be used as a search oracle — that's the whole point.

If the user hasn't given you a clean dev/test split, **construct one and say so**. The dev/test separation is the mechanism that catches overfitting: a candidate that wins on dev but not on test isn't a success, it's a warning that you're exploiting the feedback signal. Without it, autonomous search reliably overfits.

Initialize the run:

```bash python scripts/tree.py init \ --objective "Improve BrowseComp answer accuracy on the search harness" \ --dev-eval "python eval.py --split dev --n 50" \ --test-eval "python eval.py --split test --n 300" \ --material "." --metric-direction max --branching 3 --max-depth 2 --budget 12 ```

`--branching` is how many sibling hypotheses you propose per parent; `--max-depth 2` keeps directions at depth 1 and concrete interventions at depth 2 (the paper's default); `--budget` is the number of coordinator cycles. Start small (10–20 cycles) — structured search beats brute force, and you can extend if progress is still being made.

## The coordinator loop

You run repeated cycles of six steps. This is the heart of HTR; do not collapse it into ad-hoc editing. Run `python scripts/tree.py cycle` once per cycle to track the budget.

### 1. Observe Begin every cycle by re-grounding in the tree, not in your memory of the conversation:

```bash python scripts/tree.py observe ```

This prints the objective, global insights, the active frontier (selectable hypotheses), executed nodes with their evidence, pruned lessons (negative constraints), and the current best artifact. Treating the tree as the source of truth is what keeps you coherent over a long run, after context compression has thrown away the details.

### 2. Ideate Pick a promising parent and propose a few child hypotheses under it. **Condition on the tree's evidence** — this is the difference between Arbor and random search: - Validated insights are assumptions you can build on. - Pruned nodes are dead ends to avoid. - A "half-right" result is a *starting point for a sharper hypothesis*, not a reason to abandon the direction.

Each hypothesis should be a **falsifiable claim about how changing the artifact will move the metric**, not a vague intention. Depth-1 nodes are broad directions ("the search harness loses correct answers it already retrieved"); depth-2 nodes are concrete, executable interventions ("run K=5 independent rollouts and aggregate by evidence dossier instead of majority vote").

```bash python scripts/tree.py add-node --parent n0 --hypothesis "Verification, not retrieval, is the bottleneck: candidates are found but discarded" python scripts/tree.py add-node --parent n4 --hypothesis "Decompose the question into atomic constraints and verify each independently" ```

### 3. Select Choose which pending leaves to run next. **Selection is not pure score-maximization** — pick a hypothesis because it has strong prior evidence, because it would resolve an ambiguity its siblings exposed, or because its failure would clarify an important assumption. Frontier control under delayed feedback rewards informative experiments, not just promising ones.

### 4. Dispatch Run each selected hypothesis as an **executor subagent in an isolated worktree** (use the Agent tool with `isolation: "worktree"`, or have the executor create one with `git worktree add`). Isolation matters: parallel experiments must not clobber each other or the current best, and exploratory changes stay quarantined until they pass the merge gate.

Dispatch siblings **in parallel** (multiple Agent calls in one message) when they're independent — comparative evidence within one direction is exactly what makes later pruning and abstraction possible.

Give each executor a tight, **hypothesis-bound** brief. See `references/executor-brief.md` for the full template. The contract that makes HTR work: **the executor may not change the hypothesis when the metric stalls.** It repairs its own code and reruns, but `h_n` is fixed — otherwise the returned score is no longer evidence about the assigned node and the tree's semantics break. The executor returns exactly four things: - **dev_score** — the dev evaluator result (for selection); - **result** — a factual summary of what happened; - **insight** — the distilled, reusable lesson (*why* the result supports, weakens, or bounds the hypothesis); - **branch_ref** — the git branch/commit/worktree path holding the artifact.

Mark a node `running` before dispatch (`tree.py set-status --node n5 --status running`) so the Observe projection stays accurate.

### 5. Backpropagate When an executor returns, write its report into the node, then **abstract the lesson upward**:

```bash python scripts/tree.py set-evidence --node n5 --dev-score 70.0 \ --result "K=5 dossier aggregation recovers answers in minority rollouts" \ --insight "Correct answers often appear in a minority of rollouts; aggregation beats majority vote" \ --branch-ref "wt/n5"

python scripts/tree.py propagate --node n5 \ --insight "Candidate coverage, not verification, limits this direction" --to-root ```

This is the step that makes the tree more than a log. A leaf-level observation ("data-interface mismatch") should become a direction-level constraint and, if it generalizes, a global prior that shapes future ideation. **Insight propagation is the component that drives most of HTR's gains** — in the paper's MLE-Bench Lite ablation, a tree *without* insight feedback scored even lower than a flat experiment queue with no tree at all (54.5% vs. 63.6% any-medal, against 81.8% for the full system). Hierarchy alone isn't enough: the semantic memory is what matters. So spend real thought on the abstraction; don't just copy the leaf insight upward verbatim.

### 6. Decide Decide what to do with the new evidence: keep expanding a direction, prune a falsified subtree, or attempt to merge a candidate.

- **Prune** dead ends, recording *why* — the reason becomes a negative constraint: ```bash python scripts/tree.py prune --node n7 --reason "search-augmented judge overfits dev questions; no test transfer" ``` - **Merge gate** — promote a candidate to the new best **only if it improves on `E_test`**. Run the test evaluator in a *fresh* worktree (not the dev worktree, to avoid leakage), then: ```bash python scripts/tree.py merge --node n5 --test-score 67.67 --branch-ref "wt/n5" ``` If the gate rejects it, that's informative: a high-dev / low-test candidate is evidence the direction may be exploiting the dev signal rather than producing a transferable improvement. Record that lesson; don't quietly promote it anyway.

Repeat until the budget is spent, the frontier is exhausted, or progress has clearly stalled.

## Finishing the run

When you stop, produce a short report (see `references/report-template.md`) covering: - the final best artifact, its test score, and its delta over `M_0`; - the tree (`python scripts/tree.py status`) as the audit trail of what was tried; - the main hypothesis shifts — how task understanding deepened across the run (early nodes test broad mechanisms; later nodes find their limits; ancestor insights compress these into the constraints behind the final design); - merged vs. explored: many nodes improve dev, far fewer pass the test gate — report that gap honestly rather than overstating dev wins.

Always leave `M_best` as a real, runnable artifact on a named branch, and tell the user how to check it out.

## Principles that make this work (not rote rules)

These come from the paper's analysis; understanding *why* matters more than following them mechanically.

- **The tree is the memory; conversatio

Detalles técnicos

Versión
1.0.0
Licencia
MIT license
Última actualización
20 ago 2026
Publicado
20 ago 2026

Resumen de decisión

Elección principal

100
Listo
Adoptar
Etapa

33,974 estrellas de GitHub

Auditoría

Revisión de instalación

Revisión de instalación y adopción

89
Seguro para probar
Seguridad
83/100
Mantenimiento
100/100
Instalar
92/100
Abrir auditoría completaVer informe de evaluación

Evidencia probada por Agent

Evidencia probada por Agent

Informes de resultados tras resolver, revisar, instalar y una ejecución limitada.

0
Probado
Needs first agent runAuto-instalación: revisar primeroÚltimo: Desconocido
Tasa de éxito
Fallo reciente
Resultados
0
Calidad de salida
Fallidos
0
No relevante
0
Instalaciones
0
Bloqueado por riesgo
0
Configuración necesaria
0
Producción
0

Aún no hay datos de resultados de Agent. La primera ejecución puede informar éxito, configuración necesaria, bloqueos de riesgo, fallo o irrelevancia mediante /api/agent/outcome.

Instalar

Añadir al flujo de Agent

Gratis y de código abierto. Revisa el informe antes de instalar en Agents de producción.

Bucle de crecimiento

Kit para compartir

X

Borrador basado en un caso para arbor, listo para publicar manualmente en X.

Nota del curador
A practical pick for source-backed research:

arbor: Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and...

34.0K stars

https://www.openagentskill.com/skills/k-dense-ai-arbor?ref=x
Abrir borrador de X
Respuesta opcional con comando de instalación
Listing + install path for arbor:
https://www.openagentskill.com/skills/k-dense-ai-arbor?ref=x

Install: npx skills add K-Dense-AI/scientific-agent-skills --skill arbor

Fuente de la ficha

Indexado por Registry

Reclamable

Esta ficha se indexó desde fuentes públicas y no está marcada como oficial hasta que se apruebe una reclamación de mantenedor.

Creador
K-Dense-AI
Indexado por
Índice comunitario de OpenAgentSkill

La atribución enlaza al repositorio público o al perfil del creador. Los creadores pueden reclamar la ficha para actualizar las señales de propiedad.

Reclamar este skill

Reclamación del propietario

Reclamar esta ficha de skill

Esta ficha Indexado por Registry se atribuye a K-Dense-AI, pero aún no está marcada como oficial. Reclámala para añadir una señal de propietario verificado y hacer más fiables futuras actualizaciones de lanzamiento, instalación y auditoría.

Kit de enlaces para creadores

Añade las insignias de evidencia a tu README

Muestra la ficha canónica, las señales actuales de confianza y auditoría, y evidencia real de Agent-Proven donde los desarrolladores evalúan el repositorio.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=listed&label=Listed)](https://www.openagentskill.com/skills/k-dense-ai-arbor)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=trust&label=Trust)](https://www.openagentskill.com/skills/k-dense-ai-arbor)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=audit&label=Audit)](https://www.openagentskill.com/skills/k-dense-ai-arbor/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/k-dense-ai-arbor?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/k-dense-ai-arbor)

Autor

K

K-Dense-AI

@k-dense-ai

Etiquetas

Afinidad con plataforma

Señales de salud

Estrellas de GitHub
34.0K
Puntuación de calidad
55/100
Último push de GitHub
20 ago 2026
Pistas del framework
Desconocido
Vistas de OpenAgentSkill
19
Copias de instalación
0
Clics externos
0

Señal de comunidad

Comparte si este skill resulta útil para tu flujo de Agent. Los comentarios agregados mejoran la clasificación con el tiempo.

Confianza y seguridad

Revisar antes de instalar

79
  • Adopción en GitHub34K estrellas de GitHubAprobado
  • Actividad de stars/forks34K estrellas y 3.3K forks; la actividad de issues no está disponible en los metadatos actualesAprobado
  • Mantenimiento reciente3 días desde el último pushAprobado
  • Claridad de licenciaMIT licenseAprobado
  • Completitud de README/SKILL.mdLos metadatos públicos necesitan más contexto de README/SKILL.mdInfo
  • Riesgo de dependencias/runtimeSuperficie de ejecución de comandosInfo