arboreto
Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Suppo
供给资产档案
编程与开发 Agent
代码审查、仓库分析、测试、CI、GitHub、DevOps 与开发工作流 Skill。
场景
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
适配 Agent
Claude Code + CLI + Codex
适用于 Codex、Claude Code、Cursor、CLI 或自定义 Agent。
安装
就绪
npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto
维护状态
新鲜
距上次推送 2 天
风险
需审查
Permission surface may require sandboxing
GitHub 质量
34K
92/100 质量 · 83/100 信任
覆盖标签
审查说明
Permission surface may require sandboxing · Quality score needs review
Agent 采用评分卡
一眼查看信任、审计与安装准备度
这些分数综合公开仓库元数据、OpenAgentSkill 审查信号、维护新鲜度与安装准备度。它用于候选筛选,不替代人工审查。
质量
优秀高置信候选,具有较强的采用度与健康维护信号。
信任
仅限沙盒有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。
审计
需审查对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。
OpenAgentSkill 信任评分 v5
安装前需人工审查
仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。
Stars
34K 个 GitHub Stars
仓库活跃度
34K 个 Star,3.3K 个 Fork
维护状态
距上次推送 2 天
许可证
BSD-3-Clause license
安装
npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto
安装安全性
标准软件包或运行时安装路径
权限范围
shell or command execution, filesystem or document access
Agent 结果
暂未有 Agent 结果数据
文档
README/SKILL.md 上下文充分
风险摘要
生产前审查
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Permission surface: shell or command execution, filesystem or document access
安装准备度
安装路径可用
- 安装路径可用
- 仓库证据可用
- 已声明许可证
- 暂无 Agent 验证结果证据
Agent 可读元数据
这个 Skill 的机器可读决策数据。
使用此区块或内嵌 JSON 判断 Agent 是否应安装该 Skill、选择替代方案,或先请求人工审查。
适用任务
- 工作流自动化 工作流
- Claude Code 团队
- 重视 GitHub 采用信号的团队
- Move data between tools
适用 Agent
安装决策
- 命令
- npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto
- 策略
- 审查
- 人工审查
- 是
信任与风险
- 信任
- 75/100
- 审计
- 88/100
- 风险级别
- 需审查
结果闭环
- 端点
- /api/agent/outcome
- 事件 ID
- resolve
- 结果
- 5
不适用场景
- 需要厂商支持 SLA 的团队
- 没有内部安全审查的高合规环境
- 当前元数据中未发现重大风险信号
- 高风险权限提示:Shell 或命令执行
- Permission surface may require sandboxing
Agent 安全 v2
56/100 · 安装前审查
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
高
Shell 或命令执行
Skill 元数据引用了终端、CLI、Shell、子进程或命令执行工作流。
中
网络访问
Skill 可能访问远程页面、API、仓库或外部服务。
中
文件系统访问
Skill 可能读取或写入项目文件、文档、生成产物或本地工作区状态。
中
数据库访问
Skill 可能检查 Schema、查询数据库或处理持久化存储。
- 高风险权限提示:Shell 或命令执行
- Permission surface may require sandboxing
安装目标
在你的 Agent 工作流中安装此 Skill
通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install k-dense-ai-arboretoAgent 解析计划
让 Agent 在安装前验证匹配度。
Resolve API 返回首选 Skill、替代方案、安全策略、审计说明、安装目标和可直接执行的提示词,无需抓取此页面。
打开 JSON
/api/agent/resolve?task=Use%20arboreto%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve 文本
/api/agent/resolve?task=Use%20arboreto%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
安装交接
/api/skills/k-dense-ai-arboreto/install
Agent 应检查
- 从 Resolve API 检查任务匹配与替代方案。
- 检查审计评分、信任评分和安全策略警告。
- 检查 Codex、Claude Code、Cursor 或 CLI 的安装目标兼容性。
复制提示词
Task: Use arboreto in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20arboreto%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/k-dense-ai-arboreto/install
Install command: npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent 交接
把安装路径交给 Agent,而不是再给一个目录页。
通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。
安装交接
/api/skills/k-dense-ai-arboreto/install
LLM 文本格式
/api/skills/k-dense-ai-arboreto/install?format=text
寻找替代方案
/api/skills/search?q=arboreto&limit=3
Agent 提示词
Use arboreto for this task. Review https://www.openagentskill.com/api/skills/k-dense-ai-arboreto/install, then install with: npx skills add K-Dense-AI/scientific-agent-skills --skill arboretoRegistry 元数据
用于自动选择 Skill 的 Agent 可读档案。
本页通过 Registry API 提供相同的决策、信任、审计、场景和安装信号,让 Agent 无需抓取界面即可排序。
Agent 决策面板
适合 工作流自动化 的首选
将其作为优先候选,再在你的 Agent 环境中验证 README 与安装路径。
栈中角色
首选
主要匹配
工作流自动化
信任标签
可用于生产
安装路径
命令已就绪
适用场景
- 工作流自动化 工作流
- Claude Code 团队
- 重视 GitHub 采用信号的团队
证据
- 33,974 个 GitHub Stars
- 仓库近期活跃
- 已提供安装命令或 GitHub 仓库
- 92/100 质量档案
- 8 个 OpenAgentSkill 交互事件
先审查
- 当前元数据中未发现重大风险信号
实施路径
- 1在沙盒 Agent 中安装它,并端到端完成一次工作流自动化任务。
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
信任档案
仅限沙盒
有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。
GitHub 采用度
通过34K 个 GitHub Stars
Star/Fork 活跃度
通过34K 个 Star,3.3K 个 Fork; 当前元数据中没有议题活跃度信息
近期维护
通过距上次推送 2 天
许可证清晰度
通过BSD-3-Clause license
积极信号
- AI 审查已通过
- 安装路径可用
- 仓库证据可用
- 近期维护的仓库
- Large GitHub adoption signal
- 安装命令未发现明显高风险模式
- 结果闭环已就绪,但需要首次真实 Agent 运行
安装前审查
- Quality score needs review
- Permission surface needs review: shell or command execution, filesystem or document access
- Permission surface: shell or command execution, filesystem or document access
- 暂未有真实 Agent 结果报告
- 无人值守安装前需要人工审查
建议操作
仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。
质量档案
优秀 适用于 Agent 工作流的候选
高置信候选,具有较强的采用度与健康维护信号。
工作流匹配
在这些场景使用此 Skill
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Manage repositories
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Operate local tools
Local desktop
I need my agent to operate local files and desktop apps in a repeatable workflow.
工作流匹配
加入完整工作流
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
Browser QA agent
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
替代方案短名单
安装前对比
可能适合该任务的相近 Skill。
D3
Bring data to life with SVG, Canvas and HTML. :bar_chart::chart_with_upwards_trend::tada:
Echarts
Apache ECharts is a powerful, interactive charting and data visualization library for browser
Data Science For Beginners
10 Weeks, 20 Lessons, Data Science for All!
Sequelize
Feature-rich ORM for modern Node.js and TypeScript, it supports PostgreSQL (with JSON and JSONB support), MySQL, MariaDB, SQLite, MS SQL Server, Snowflake, Oracle DB, DB2 and DB2 for IBM i.
概览
--- name: arboreto description: Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets. license: BSD-3-Clause license metadata: version: "1.0" skill-author: K-Dense Inc. ---
# Arboreto
## Overview
Arboreto is a Python library from [Aerts Lab](https://github.com/aertslab/arboreto) for inferring gene regulatory networks (GRNs) from gene expression data. It parallelizes tree-based ensemble regression (GRNBoost2, GENIE3) with [Dask](https://distributed.dask.org/) across local cores or remote clusters.
**Core capability**: Identify which transcription factors (TFs) regulate which target genes based on expression patterns across observations (cells, samples, conditions).
**Upstream**: PyPI **0.1.6** (2021-02-09, latest). Docs: [arboreto.readthedocs.io](https://arboreto.readthedocs.io/en/latest/). Primary downstream consumer: [pySCENIC](https://github.com/aertslab/pySCENIC).
## Quick Start
Install arboreto: ```bash uv pip install arboreto ```
Basic GRN inference: ```python import pandas as pd from arboreto.algo import grnboost2
if __name__ == '__main__': # Load expression data (genes as columns) expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
# Infer regulatory network network = grnboost2(expression_data=expression_matrix)
# Save results (TF, target, importance) network.to_csv('network.tsv', sep='\t', index=False, header=False) ```
**Critical**: Always use `if __name__ == '__main__':` guard because Dask spawns new processes.
## Core Capabilities
### 1. Basic GRN Inference
For standard GRN inference workflows including: - Input data preparation (Pandas DataFrame or NumPy array) - Running inference with GRNBoost2 or GENIE3 - Filtering by transcription factors - Output format and interpretation
**See**: `references/basic_inference.md`
**Use the ready-to-run script**: `scripts/basic_grn_inference.py` for standard inference tasks: ```bash python scripts/basic_grn_inference.py expression_data.tsv output_network.tsv --tf-file tfs.txt --seed 777 --limit 5000 ```
### 2. Algorithm Selection
Arboreto provides two algorithms:
**GRNBoost2 (Recommended)**: - Fast gradient boosting-based inference - Optimized for large datasets (10k+ observations) - Default choice for most analyses
**GENIE3**: - Random Forest-based inference - Original multiple regression approach - Use for comparison or validation
Quick comparison: ```python from arboreto.algo import grnboost2, genie3
# Fast, recommended network_grnboost = grnboost2(expression_data=matrix)
# Classic algorithm network_genie3 = genie3(expression_data=matrix) ```
**For detailed algorithm comparison, parameters, and selection guidance**: `references/algorithms.md`
### 3. Distributed Computing
Scale inference from local multi-core to cluster environments:
**Local (default)** - Uses all available cores automatically: ```python network = grnboost2(expression_data=matrix) ```
**Custom local client** - Control resources: ```python from distributed import LocalCluster, Client
local_cluster = LocalCluster(n_workers=10, memory_limit='8GB') client = Client(local_cluster)
network = grnboost2(expression_data=matrix, client_or_address=client)
client.close() local_cluster.close() ```
**Cluster computing** - Connect to remote Dask scheduler: ```python from distributed import Client
client = Client('tcp://scheduler:8786') network = grnboost2(expression_data=matrix, client_or_address=client) ```
**For cluster setup, performance optimization, and large-scale workflows**: `references/distributed_computing.md`
## Installation
```bash uv pip install arboreto ```
Conda (Bioconda):
```bash conda install -c bioconda arboreto ```
**Dependencies** (from upstream `requirements.txt`): `dask[complete]`, `distributed`, `numpy`, `pandas`, `scikit-learn`, `scipy`
**Input formats**: pandas DataFrame, dense `numpy.ndarray`, or sparse `scipy.sparse.csc_matrix` (rows = observations, columns = genes). For array/matrix inputs, pass `gene_names` explicitly.
## Common Use Cases
### Single-Cell RNA-seq Analysis ```python import pandas as pd from arboreto.algo import grnboost2
if __name__ == '__main__': # Load single-cell expression matrix (cells x genes) sc_data = pd.read_csv('scrna_counts.tsv', sep='\t')
# Infer cell-type-specific regulatory network network = grnboost2(expression_data=sc_data, seed=42)
# Filter high-confidence links high_confidence = network[network['importance'] > 0.5] high_confidence.to_csv('grn_high_confidence.tsv', sep='\t', index=False) ```
### Bulk RNA-seq with TF Filtering ```python from arboreto.utils import load_tf_names from arboreto.algo import grnboost2
if __name__ == '__main__': # Load data expression_data = pd.read_csv('rnaseq_tpm.tsv', sep='\t') tf_names = load_tf_names('human_tfs.txt')
# Infer with TF restriction network = grnboost2( expression_data=expression_data, tf_names=tf_names, seed=123 )
network.to_csv('tf_target_network.tsv', sep='\t', index=False) ```
### Comparative Analysis (Multiple Conditions) ```python from arboreto.algo import grnboost2
if __name__ == '__main__': # Infer networks for different conditions conditions = ['control', 'treatment_24h', 'treatment_48h']
for condition in conditions: data = pd.read_csv(f'{condition}_expression.tsv', sep='\t') network = grnboost2(expression_data=data, seed=42) network.to_csv(f'{condition}_network.tsv', sep='\t', index=False) ```
## Output Interpretation
Arboreto returns a DataFrame with regulatory links:
| Column | Description | |--------|-------------| | `TF` | Transcription factor (regulator) | | `target` | Target gene | | `importance` | Regulatory importance score (higher = stronger) |
**Filtering strategy**: - `limit=N` at inference time (return top N links globally) - Post-hoc importance threshold (e.g., > 0.5) - Top links per target via `groupby('target')` - Statistical significance testing (permutation tests, external tools)
## Integration with pySCENIC
Arboreto powers the GRN inference step in [pySCENIC](https://github.com/aertslab/pySCENIC). pySCENIC 0.11+ passes sparse expression matrices to `grnboost2` / `genie3`; pySCENIC 0.12+ defaults to `arboreto_with_multiprocessing.py` (no Dask) for compatibility — use standalone arboreto when you need Dask scaling.
```python # Standalone: infer co-expression modules before pySCENIC cisTarget pruning from arboreto.algo import grnboost2
network = grnboost2(expression_data=expression_df, tf_names=tf_list, limit=5000)
# Downstream: pySCENIC ctx pruning, regulon definition, AUCell (see pySCENIC docs) ```
Convert AnnData to a DataFrame for arboreto directly:
```python expression_df = adata.to_df() # cells x genes ```
## Reproducibility
Always set a seed for reproducible results: ```python network = grnboost2(expression_data=matrix, seed=777) ```
Run multiple seeds for robustness analysis: ```python from distributed import LocalCluster, Client
if __name__ == '__main__': client = Client(LocalCluster())
seeds = [42, 123, 777] networks = []
for seed in seeds: net = grnboost2(expression_data=matrix, client_or_address=client, seed=seed) networks.append(net)
# Consensus: links recurring across runs (example: mean importance per TF-target pair) import pandas as pd combined = pd.concat(networks) consensus = ( combined.groupby(['TF', 'target'], as_index=False)['importance'] .mean() .query('importance > 0.5') ) ```
## Troubleshooting
**Memory errors**: Reduce dataset size by filtering low-variance genes or use distributed computing
**Slow performance**: Use GRNBoost2 instead of GENIE3, enable distributed client, filter TF list
**Dask errors**: Ensure `if __name__ == '__main__':` guard is present in scripts (required on Windows/macOS with spawn-based multiprocessing)
**Empty results**: Check data format (genes as columns), verify TF names match column names in the expression matrix
**Sparse data**: Use `scipy.sparse.csc_matrix` and pass matching `gene_names`; supported since arboreto 0.1.6 / pySCENIC 0.11
技术详情
- 版本
- 1.0.0
- 许可证
- BSD-3-Clause license
- 最近更新
- 2026年8月20日
- 发布时间
- 2026年8月20日
决策摘要
首选
33,974 个 GitHub Stars
Agent 验证证据
Agent 验证证据
来自解析、审查、安装和一次小范围运行后的结果报告。
- 成功率
- —
- 近期失败
- —
- 结果
- 0
- 输出质量
- —
- 失败
- 0
- 不相关
- 0
- 安装次数
- 0
- 风险拦截
- 0
- 需要配置
- 0
- 生产环境
- 0
暂时没有 Agent 结果数据。首次 Agent 执行可以通过 /api/agent/outcome 报告成功、需要设置、风险拦截、失败或不相关。
增长闭环
分享工具包
为 arboreto 准备的场景化草稿,可手动发布到 X。
arboreto: Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GR... 34.0K stars https://www.openagentskill.com/skills/k-dense-ai-arboreto?ref=x
可选:带安装命令的回复
Listing + install path for arboreto: https://www.openagentskill.com/skills/k-dense-ai-arboreto?ref=x Install: npx skills add K-Dense-AI/scientific-agent-skills --skill arboreto
收录来源
Registry 收录
此列表来自公开来源,维护者认领获批前不会标记为官方。
- 创作者
- K-Dense-AI
- 收录方
- OpenAgentSkill 社区索引
归属链接指向公开仓库或创作者主页。创作者可认领列表以更新所有权信号。
认领此 Skill所有者认领
认领此 Skill 页面
这条 Registry 收录 列表归属于 K-Dense-AI,但尚未标记为官方。认领后可增加已验证所有者信号,使后续发布、安装和审计更新更值得信赖。
创作者外链工具包
将证据徽章加入你的 README
在开发者评估仓库的位置展示规范页面、当前信任与审计信号,以及真实的 Agent 验证证据。
[](https://www.openagentskill.com/skills/k-dense-ai-arboreto)
[](https://www.openagentskill.com/skills/k-dense-ai-arboreto)
[](https://www.openagentskill.com/skills/k-dense-ai-arboreto/audit)
[](https://www.openagentskill.com/skills/k-dense-ai-arboreto)作者
K-Dense-AI
@k-dense-ai
平台适配
健康信号
- GitHub Stars
- 34.0K
- 质量评分
- 55/100
- 最近 GitHub 推送
- 2026年8月20日
- 框架提示
- 未知
- OpenAgentSkill 浏览量
- 8
- 复制安装命令
- 0
- 跳转点击
- 0
社区信号
告诉我们这个 Skill 是否对你的 Agent 工作流有帮助。汇总反馈会持续改善排序。
信任与安全
仅限沙盒
- GitHub 采用度34K 个 GitHub Stars通过
- Star/Fork 活跃度34K 个 Star,3.3K 个 Fork; 当前元数据中没有议题活跃度信息通过
- 近期维护距上次推送 2 天通过
- 许可证清晰度BSD-3-Clause license通过
- README/SKILL.md 完整度元数据包含足够的用法与工作流上下文通过
- 依赖与运行时风险command execution surface, external package install surface信息
相关 Skill
D3
Bring data to life with SVG, Canvas and HTML. :bar_chart::chart_with_upwards_trend::tada:
113.1K StarsEcharts
Apache ECharts is a powerful, interactive charting and data visualization library for browser
66.6K StarsData Science For Beginners
10 Weeks, 20 Lessons, Data Science for All!
35.6K StarsSequelize
Feature-rich ORM for modern Node.js and TypeScript, it supports PostgreSQL (with JSON and JSONB support), MySQL, MariaDB, SQLite, MS SQL Server, Snowflake, Oracle DB, DB2 and DB2 for IBM i.
30.4K Stars