arboreto

REVIEW · 75
Registry indexed

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Suppo

Verified installs0
Stars34.0K
Version1.0.0
Quality92/100 · Excellent
Trust75/100 · Sandbox only
Audit88/100 · Needs review

Supply asset profile

Coding and developer agents

Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.

Browse track

Scenario

GitHub automation

I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.

Agent fit

Claude Code + CLI + Codex

Codex, Claude Code, Cursor, CLI, or custom agents.

Install

Ready

npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto

Maintenance

fresh

2d since push

Risk

Needs review

Permission surface may require sandboxing

GitHub quality

34K

92/100 Quality · 83/100 Trust

Coverage tags

CodingGitHub automationdata-analysisagent-skill

Review notes

Permission surface may require sandboxing · Quality score needs review

Agent adoption scorecard

Trust, audit, and install readiness at a glance

These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.

Quality

Excellent
92

High-confidence pick with strong adoption and healthy maintenance signals.

Trust

Sandbox only
75

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

Audit

Needs review
88

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

OpenAgentSkill Trust Score v5

Human review before install

Run only in a sandbox and compare close alternatives before using it for real work.

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

34K GitHub stars

Repo activity

34K stars, 3.3K forks

Maintenance

2d since push

License

BSD-3-Clause license

Install

npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto

Install safety

standard package or runtime install path

Permission surface

shell or command execution, filesystem or document access

Agent outcomes

No agent outcome data yet

Docs

Strong README/SKILL.md context

Risk summary

Review before production

  • Quality score needs review
  • Permission surface needs review: shell or command execution, filesystem or document access
  • Permission surface: shell or command execution, filesystem or document access

Install readiness

Install path available

  • Install path is available
  • Repository evidence is available
  • License is declared
  • No Agent Proven outcome evidence yet

Agent-readable metadata

Machine-readable decision data for this skill.

Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.

Open JSON

Suited tasks

  • Workflow automation workflows
  • Claude Code teams
  • teams that value GitHub adoption signals
  • Move data between tools

Suited agents

CodexClaude CodeCursorOpenAgentSkill CLICLI

Install decision

Command
npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto
Policy
review
Human review
yes

Trust and risk

Trust
75/100
Audit
88/100
Risk level
Needs review

Outcome loop

Endpoint
/api/agent/outcome
Event ID
resolve
Outcomes
5

Install command

npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto

Do not use when

  • teams that need a vendor-supported SLA
  • high-compliance environments without internal security review
  • No major risk signals from current metadata
  • High-risk permission hints: Shell or command execution
  • Permission surface may require sandboxing

Agent safety v2

56/100 · Review before install

Experimentalreview

Sparse or mixed signals. Useful for discovery, but not for autonomous installation.

Test manually in an isolated workspace and compare against safer alternatives.

Resolve via API

high

Shell or command execution

Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.

medium

Network access

Skill likely fetches remote pages, APIs, repositories, or external services.

medium

Filesystem access

Skill may read or write project files, documents, generated artifacts, or local workspace state.

medium

Database access

Skill may inspect schemas, query databases, or work with persistent stores.

  • High-risk permission hints: Shell or command execution
  • Permission surface may require sandboxing

Install targets

Install this skill in your agent workflow

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install k-dense-ai-arboreto

Agent resolve plan

Let an agent verify fit before installing.

The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.

Open text plan

Agent should check

  • Task fit and alternatives from Resolve API.
  • Audit score, trust score, and safety policy warnings.
  • Install target compatibility for Codex, Claude Code, Cursor, or CLI.

Copy prompt

Task: Use arboreto in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20arboreto%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/k-dense-ai-arboreto/install
Install command: npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent handoff

Give an agent the install path, not another directory page.

Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.

Open install API

Agent prompt

Use arboreto for this task. Review https://www.openagentskill.com/api/skills/k-dense-ai-arboreto/install, then install with: npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto

Registry metadata

Agent-readable profile for automatic skill selection.

This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.

Open manifest

Agent fit

100/100

Workflow automation

Platforms

Claude Code

Audit report

Needs review · 88/100

A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.

View audit reportView eval report

Agent decision cockpit

Primary pick for Workflow automation

Use this as a leading candidate, then validate the README and install path in your own agent stack.

100
Readiness
Adopt
Stage

Role in stack

Primary pick

Primary fit

Workflow automation

Trust label

Production-ready

Install path

Command ready

Use when

  • Workflow automation workflows
  • Claude Code teams
  • teams that value GitHub adoption signals

Evidence

  • 33,974 GitHub stars
  • recent repository activity
  • install command or GitHub repo available
  • 92/100 quality profile
  • 8 OpenAgentSkill engagement events

review first

  • No major risk signals from current metadata

Implementation path

  1. 1Install it in a sandbox agent and run one Workflow automation task end to end.
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

Trust profile

Sandbox only

Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.

75
OpenAgentSkill Trust Score

GitHub adoption

PASS

34K GitHub stars

Stars/forks activity

PASS

34K stars, 3.3K forks; issue activity unavailable in current metadata

Recent maintenance

PASS

2d since push

License clarity

PASS

BSD-3-Clause license

Good signals

  • AI review approved
  • Install path is available
  • Repository evidence is available
  • Recently maintained repository
  • Large GitHub adoption signal
  • Install command has no obvious high-risk pattern
  • Outcome loop is ready but needs first real agent run

Review before install

  • Quality score needs review
  • Permission surface needs review: shell or command execution, filesystem or document access
  • Permission surface: shell or command execution, filesystem or document access
  • No real agent outcome reports yet
  • Human review required before unattended installation

Recommended action

Run only in a sandbox and compare close alternatives before using it for real work.

Quality profile

Excellent candidate for agent workflows

High-confidence pick with strong adoption and healthy maintenance signals.

92
GitHub stars
34K
Freshness
2d ago
Install ready
Yes
License
BSD-3-Clause license

Workflow fit

Use this skill in these scenarios

Workflow fit

Add it to a complete workflow

Alternative shortlist

Compare before you install

Similar skills that may fit this task.

Compare all

Overview

--- name: arboreto description: Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets. license: BSD-3-Clause license metadata: version: "1.0" skill-author: K-Dense Inc. ---

# Arboreto

## Overview

Arboreto is a Python library from [Aerts Lab](https://github.com/aertslab/arboreto) for inferring gene regulatory networks (GRNs) from gene expression data. It parallelizes tree-based ensemble regression (GRNBoost2, GENIE3) with [Dask](https://distributed.dask.org/) across local cores or remote clusters.

**Core capability**: Identify which transcription factors (TFs) regulate which target genes based on expression patterns across observations (cells, samples, conditions).

**Upstream**: PyPI **0.1.6** (2021-02-09, latest). Docs: [arboreto.readthedocs.io](https://arboreto.readthedocs.io/en/latest/). Primary downstream consumer: [pySCENIC](https://github.com/aertslab/pySCENIC).

## Quick Start

Install arboreto: ```bash uv pip install arboreto ```

Basic GRN inference: ```python import pandas as pd from arboreto.algo import grnboost2

if __name__ == '__main__': # Load expression data (genes as columns) expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')

# Infer regulatory network network = grnboost2(expression_data=expression_matrix)

# Save results (TF, target, importance) network.to_csv('network.tsv', sep='\t', index=False, header=False) ```

**Critical**: Always use `if __name__ == '__main__':` guard because Dask spawns new processes.

## Core Capabilities

### 1. Basic GRN Inference

For standard GRN inference workflows including: - Input data preparation (Pandas DataFrame or NumPy array) - Running inference with GRNBoost2 or GENIE3 - Filtering by transcription factors - Output format and interpretation

**See**: `references/basic_inference.md`

**Use the ready-to-run script**: `scripts/basic_grn_inference.py` for standard inference tasks: ```bash python scripts/basic_grn_inference.py expression_data.tsv output_network.tsv --tf-file tfs.txt --seed 777 --limit 5000 ```

### 2. Algorithm Selection

Arboreto provides two algorithms:

**GRNBoost2 (Recommended)**: - Fast gradient boosting-based inference - Optimized for large datasets (10k+ observations) - Default choice for most analyses

**GENIE3**: - Random Forest-based inference - Original multiple regression approach - Use for comparison or validation

Quick comparison: ```python from arboreto.algo import grnboost2, genie3

# Fast, recommended network_grnboost = grnboost2(expression_data=matrix)

# Classic algorithm network_genie3 = genie3(expression_data=matrix) ```

**For detailed algorithm comparison, parameters, and selection guidance**: `references/algorithms.md`

### 3. Distributed Computing

Scale inference from local multi-core to cluster environments:

**Local (default)** - Uses all available cores automatically: ```python network = grnboost2(expression_data=matrix) ```

**Custom local client** - Control resources: ```python from distributed import LocalCluster, Client

local_cluster = LocalCluster(n_workers=10, memory_limit='8GB') client = Client(local_cluster)

network = grnboost2(expression_data=matrix, client_or_address=client)

client.close() local_cluster.close() ```

**Cluster computing** - Connect to remote Dask scheduler: ```python from distributed import Client

client = Client('tcp://scheduler:8786') network = grnboost2(expression_data=matrix, client_or_address=client) ```

**For cluster setup, performance optimization, and large-scale workflows**: `references/distributed_computing.md`

## Installation

```bash uv pip install arboreto ```

Conda (Bioconda):

```bash conda install -c bioconda arboreto ```

**Dependencies** (from upstream `requirements.txt`): `dask[complete]`, `distributed`, `numpy`, `pandas`, `scikit-learn`, `scipy`

**Input formats**: pandas DataFrame, dense `numpy.ndarray`, or sparse `scipy.sparse.csc_matrix` (rows = observations, columns = genes). For array/matrix inputs, pass `gene_names` explicitly.

## Common Use Cases

### Single-Cell RNA-seq Analysis ```python import pandas as pd from arboreto.algo import grnboost2

if __name__ == '__main__': # Load single-cell expression matrix (cells x genes) sc_data = pd.read_csv('scrna_counts.tsv', sep='\t')

# Infer cell-type-specific regulatory network network = grnboost2(expression_data=sc_data, seed=42)

# Filter high-confidence links high_confidence = network[network['importance'] > 0.5] high_confidence.to_csv('grn_high_confidence.tsv', sep='\t', index=False) ```

### Bulk RNA-seq with TF Filtering ```python from arboreto.utils import load_tf_names from arboreto.algo import grnboost2

if __name__ == '__main__': # Load data expression_data = pd.read_csv('rnaseq_tpm.tsv', sep='\t') tf_names = load_tf_names('human_tfs.txt')

# Infer with TF restriction network = grnboost2( expression_data=expression_data, tf_names=tf_names, seed=123 )

network.to_csv('tf_target_network.tsv', sep='\t', index=False) ```

### Comparative Analysis (Multiple Conditions) ```python from arboreto.algo import grnboost2

if __name__ == '__main__': # Infer networks for different conditions conditions = ['control', 'treatment_24h', 'treatment_48h']

for condition in conditions: data = pd.read_csv(f'{condition}_expression.tsv', sep='\t') network = grnboost2(expression_data=data, seed=42) network.to_csv(f'{condition}_network.tsv', sep='\t', index=False) ```

## Output Interpretation

Arboreto returns a DataFrame with regulatory links:

| Column | Description | |--------|-------------| | `TF` | Transcription factor (regulator) | | `target` | Target gene | | `importance` | Regulatory importance score (higher = stronger) |

**Filtering strategy**: - `limit=N` at inference time (return top N links globally) - Post-hoc importance threshold (e.g., > 0.5) - Top links per target via `groupby('target')` - Statistical significance testing (permutation tests, external tools)

## Integration with pySCENIC

Arboreto powers the GRN inference step in [pySCENIC](https://github.com/aertslab/pySCENIC). pySCENIC 0.11+ passes sparse expression matrices to `grnboost2` / `genie3`; pySCENIC 0.12+ defaults to `arboreto_with_multiprocessing.py` (no Dask) for compatibility — use standalone arboreto when you need Dask scaling.

```python # Standalone: infer co-expression modules before pySCENIC cisTarget pruning from arboreto.algo import grnboost2

network = grnboost2(expression_data=expression_df, tf_names=tf_list, limit=5000)

# Downstream: pySCENIC ctx pruning, regulon definition, AUCell (see pySCENIC docs) ```

Convert AnnData to a DataFrame for arboreto directly:

```python expression_df = adata.to_df() # cells x genes ```

## Reproducibility

Always set a seed for reproducible results: ```python network = grnboost2(expression_data=matrix, seed=777) ```

Run multiple seeds for robustness analysis: ```python from distributed import LocalCluster, Client

if __name__ == '__main__': client = Client(LocalCluster())

seeds = [42, 123, 777] networks = []

for seed in seeds: net = grnboost2(expression_data=matrix, client_or_address=client, seed=seed) networks.append(net)

# Consensus: links recurring across runs (example: mean importance per TF-target pair) import pandas as pd combined = pd.concat(networks) consensus = ( combined.groupby(['TF', 'target'], as_index=False)['importance'] .mean() .query('importance > 0.5') ) ```

## Troubleshooting

**Memory errors**: Reduce dataset size by filtering low-variance genes or use distributed computing

**Slow performance**: Use GRNBoost2 instead of GENIE3, enable distributed client, filter TF list

**Dask errors**: Ensure `if __name__ == '__main__':` guard is present in scripts (required on Windows/macOS with spawn-based multiprocessing)

**Empty results**: Check data format (genes as columns), verify TF names match column names in the expression matrix

**Sparse data**: Use `scipy.sparse.csc_matrix` and pass matching `gene_names`; supported since arboreto 0.1.6 / pySCENIC 0.11

Technical details

Version
1.0.0
License
BSD-3-Clause license
Last updated
Aug 20, 2026
Published
Aug 20, 2026

Decision snapshot

Primary pick

100
Ready
Adopt
Stage

33,974 GitHub stars

Audit

Install review

Install and adoption review

88
Needs review
Security
81/100
Maintenance
100/100
Install
92/100
Open full auditView eval report

Agent-proven evidence

Agent-proven evidence

Outcome reports after resolve, review, install, and one narrow run.

0
Proven
Needs first agent runAuto-install: review firstLast: Unknown
Success rate
Recent failure
Outcomes
0
Output quality
Failed
0
Not relevant
0
Installs
0
Risk blocked
0
Setup needed
0
Production
0

No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.

Install

Add to agent workflow

Free and open source. Review the report before installing into production agents.

Growth loop

Share kit

X

Scenario-led draft for arboreto, ready for a manual X post.

Curator note
arboreto: Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GR...

34.0K stars

https://www.openagentskill.com/skills/k-dense-ai-arboreto?ref=x
Open X draft
Optional reply with install command
Listing + install path for arboreto:
https://www.openagentskill.com/skills/k-dense-ai-arboreto?ref=x

Install: npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto

Listing source

Registry indexed

Claimable

This listing was indexed from public sources and is not marked official until a maintainer claim is approved.

Creator
K-Dense-AI
Indexed by
OpenAgentSkill community index

Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.

Claim this skill

Owner claim

Claim this skill listing

This Registry indexed listing is attributed to K-Dense-AI but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.

Creator backlink kit

Add the evidence badges to your README

Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/k-dense-ai-arboreto?metric=listed&label=Listed)](https://www.openagentskill.com/skills/k-dense-ai-arboreto)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/k-dense-ai-arboreto?metric=trust&label=Trust)](https://www.openagentskill.com/skills/k-dense-ai-arboreto)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/k-dense-ai-arboreto?metric=audit&label=Audit)](https://www.openagentskill.com/skills/k-dense-ai-arboreto/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/k-dense-ai-arboreto?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/k-dense-ai-arboreto)

Author

K

K-Dense-AI

@k-dense-ai

Platform fit

Health signals

GitHub stars
34.0K
Quality score
55/100
Last GitHub push
Aug 20, 2026
Framework hints
Unknown
OpenAgentSkill views
8
Install copies
0
Outbound clicks
0

Community signal

Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.

Trust & safety

Sandbox only

75
  • GitHub adoption34K GitHub starsPASS
  • Stars/forks activity34K stars, 3.3K forks; issue activity unavailable in current metadataPASS
  • Recent maintenance2d since pushPASS
  • License clarityBSD-3-Clause licensePASS
  • README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
  • Dependency/runtime riskcommand execution surface, external package install surfaceINFO