analyze-fasta

审查 · 61
已收录

Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus structured JSON for downstream chaining.

Verified installs0
Stars1.1K
版本1.0.0
质量78/100 ·
信任61/100 · 仅限沙盒
审计79/100 · 需审查

供给资产档案

研究与知识工作

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

浏览赛道

场景

研究 Agent

I need my agent to research a topic, compare sources, and produce a concise report.

适配 Agent

Claude Code + CLI + Codex

适用于 Codex、Claude Code、Cursor、CLI 或自定义 Agent。

安装

就绪

npx skills add ClawBio/ClawBio --skill analyze-fasta

维护状态

新鲜

今天有推送

风险

需审查

Dependency or permission surface needs review

GitHub 质量

1.1K

78/100 质量 · 69/100 信任

覆盖标签

研究研究 Agent生产力agent-skill

审查说明

Dependency or permission surface needs review · Permission surface may require sandboxing

Agent 采用评分卡

一眼查看信任、审计与安装准备度

这些分数综合公开仓库元数据、OpenAgentSkill 审查信号、维护新鲜度与安装准备度。它用于候选筛选,不替代人工审查。

质量

78

可靠的选择,值得加入生产工作流候选列表。

信任

仅限沙盒
61

有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。

审计

需审查
79

对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。

OpenAgentSkill 信任评分 v5

安装前需人工审查

仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

1.1K 个 GitHub Stars

仓库活跃度

1.1K 个 Star,257 个 Fork

维护状态

今天有推送

许可证

MIT

安装

npx skills add ClawBio/ClawBio --skill analyze-fasta

安装安全性

标准软件包或运行时安装路径

权限范围

secrets or environment access, shell or command execution

Agent 结果

暂未有 Agent 结果数据

文档

README/SKILL.md 上下文充分

风险摘要

生产前审查

  • The SKILL.md excerpt is truncated; full documentation may be incomplete, but the provided sections are clear.
  • Quality score needs review
  • Permission surface needs review: secrets or environment access, shell or command execution
  • Dependency/runtime risk: command execution surface, credential or environment access

安装准备度

安装路径可用

  • 安装路径可用
  • 仓库证据可用
  • 已声明许可证
  • 暂无 Agent 验证结果证据

Agent 可读元数据

这个 Skill 的机器可读决策数据。

使用此区块或内嵌 JSON 判断 Agent 是否应安装该 Skill、选择替代方案,或先请求人工审查。

打开 JSON

适用任务

  • 研究 Agent 工作流
  • Claude Code 团队
  • 重视 GitHub 采用信号的团队
  • 检索来源

适用 Agent

CodexClaude CodeCursorOpenAgentSkill CLICLI

安装决策

命令
npx skills add ClawBio/ClawBio --skill analyze-fasta
策略
阻止
人工审查

信任与风险

信任
61/100
审计
79/100
风险级别
需审查

结果闭环

端点
/api/agent/outcome
事件 ID
resolve
结果
5

安装命令

npx skills add ClawBio/ClawBio --skill analyze-fasta

不适用场景

  • 需要厂商支持 SLA 的团队
  • production agents without a repository review
  • The SKILL.md excerpt is truncated; full documentation may be incomplete, but the provided sections are clear.
  • 高风险权限提示:Shell or command execution, Secrets or environment access
  • Dependency or permission surface needs review

Agent 安全 v2

39/100 · 避免自动安装

Blocked for auto-install阻止

This skill should not be selected by an agent without explicit human security review.

Do not auto-install. Inspect the source, dependencies, and permission surface first.

通过 API 解析

Shell 或命令执行

Skill 元数据引用了终端、CLI、Shell、子进程或命令执行工作流。

网络访问

Skill 可能访问远程页面、API、仓库或外部服务。

文件系统访问

Skill 可能读取或写入项目文件、文档、生成产物或本地工作区状态。

Secrets or environment access

Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.

  • 高风险权限提示:Shell or command execution, Secrets or environment access
  • Dependency or permission surface needs review

安装目标

在你的 Agent 工作流中安装此 Skill

通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install clawbio-analyze-fasta

Agent 解析计划

让 Agent 在安装前验证匹配度。

Resolve API 返回首选 Skill、替代方案、安全策略、审计说明、安装目标和可直接执行的提示词,无需抓取此页面。

打开文本计划

Agent 应检查

  • 从 Resolve API 检查任务匹配与替代方案。
  • 检查审计评分、信任评分和安全策略警告。
  • 检查 Codex、Claude Code、Cursor 或 CLI 的安装目标兼容性。

复制提示词

Task: Use analyze-fasta in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20analyze-fasta%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/clawbio-analyze-fasta/install
Install command: npx skills add ClawBio/ClawBio --skill analyze-fasta
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 交接

把安装路径交给 Agent,而不是再给一个目录页。

通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。

打开安装 API

Agent 提示词

Use analyze-fasta for this task. Review https://www.openagentskill.com/api/skills/clawbio-analyze-fasta/install, then install with: npx skills add ClawBio/ClawBio --skill analyze-fasta

Registry 元数据

用于自动选择 Skill 的 Agent 可读档案。

本页通过 Registry API 提供相同的决策、信任、审计、场景和安装信号,让 Agent 无需抓取界面即可排序。

打开 Manifest

适配 Agent

89/100

研究 Agent

平台

Claude Code

审计报告

需审查 · 79/100

对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。

查看审计报告查看评估报告

Agent 决策面板

适合 研究 Agent 的首选

将其作为优先候选,再在你的 Agent 环境中验证 README 与安装路径。

89
就绪度
采用
阶段

栈中角色

首选

主要匹配

研究 Agent

信任标签

可用于生产

安装路径

命令已就绪

适用场景

  • 研究 Agent 工作流
  • Claude Code 团队
  • 重视 GitHub 采用信号的团队

证据

  • 1,112 个 GitHub Stars
  • 仓库近期活跃
  • 已提供安装命令或 GitHub 仓库
  • 78/100 质量档案
  • 1 个 OpenAgentSkill 交互事件

先审查

  • The SKILL.md excerpt is truncated; full documentation may be incomplete, but the provided sections are clear.

实施路径

  1. 1在沙盒 Agent 中安装它,并端到端完成一次研究 Agent任务。
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

信任档案

仅限沙盒

有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。

61
OpenAgentSkill 信任评分

GitHub 采用度

通过

1.1K 个 GitHub Stars

Star/Fork 活跃度

通过

1.1K 个 Star,257 个 Fork; 当前元数据中没有议题活跃度信息

近期维护

通过

今天有推送

许可证清晰度

通过

MIT

积极信号

  • AI 审查已通过
  • 安装路径可用
  • 仓库证据可用
  • 近期维护的仓库
  • 有意义的 GitHub 采用信号
  • 安装命令未发现明显高风险模式
  • 结果闭环已就绪,但需要首次真实 Agent 运行

安装前审查

  • The SKILL.md excerpt is truncated; full documentation may be incomplete, but the provided sections are clear.
  • Quality score needs review
  • Permission surface needs review: secrets or environment access, shell or command execution
  • Dependency/runtime risk: command execution surface, credential or environment access
  • Permission surface: secrets or environment access, shell or command execution
  • 暂未有真实 Agent 结果报告
  • 无人值守安装前需要人工审查

建议操作

仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。

质量档案

适用于 Agent 工作流的候选

可靠的选择,值得加入生产工作流候选列表。

78
GitHub Stars
1.1K
新鲜度
今天
安装就绪
许可证
MIT
安装前审查: The SKILL.md excerpt is truncated; full documentation may be incomplete, but the provided sections are clear.

工作流匹配

在这些场景使用此 Skill

工作流匹配

加入完整工作流

替代方案短名单

安装前对比

可能适合该任务的相近 Skill。

对比全部

概览

--- name: analyze-fasta description: Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus structured JSON for downstream chaining. license: MIT metadata: version: "0.1.0" author: Santiago Rodriguez Salinas domain: genomics tags: - fasta - biopython - sequence-analysis - gc-content - orf - protein-properties - isoelectric-point - gravy inputs: - name: input type: file format: - fasta - fa - fna - faa description: Single FASTA file with one or more nucleotide or protein records required: true outputs: - name: report type: file format: - md description: Markdown report with summary table, per-sequence metrics, and disclaimer - name: result type: file format: - json description: Machine-readable analysis results (sequence type, per-record metrics, summary) - name: report_html type: file format: - html description: Standalone HTML rendering of the same report for visual inspection - name: reproducibility type: directory description: Directory with commands.sh and run.json describing the exact run dependencies: python: ">=3.10" packages: - biopython>=1.80 demo_data: - path: example_data/demo_nucleotide.fasta description: Synthetic ~720 bp nucleotide sequence with a small ORF (CC0, no real organism) - path: example_data/demo_protein.fasta description: Synthetic ~120 aa protein sequence (CC0, no real organism) endpoints: cli: python skills/analyze-fasta/analyze_fasta.py --input {input_file} --output {output_dir} openclaw: requires: bins: - python3 env: config: always: false emoji: "🧬" homepage: https://github.com/ClawBio/ClawBio os: - darwin - linux install: - kind: pip package: biopython bins: trigger_keywords: - fasta - analyze fasta - analiza fasta - sequence analysis - gc content - find orfs - orf finder - protein properties - isoelectric point - gravy index - protparam - molecular weight protein - molecular weight dna ---

# 🧬 analyze-fasta

You are **analyze-fasta**, a specialised ClawBio agent for single-FASTA inspection. Your role is to take a FASTA file (nucleotide or protein), auto-detect its type, compute the standard set of sequence-level metrics with Biopython, and produce a structured report that downstream skills can chain to.

## Trigger

**Fire this skill when the user says any of:** - "analyze this fasta" - "analiza este fasta" - "what's the GC content of this sequence" - "find ORFs in this sequence" - "compute pI / isoelectric point of this protein" - "GRAVY index" - "protein properties from this fasta" - "summarise this fasta" - "describe this sequence"

**Do NOT fire when:** - The user has FASTQ reads — route to `seq-wrangler` (alignment QC). - The user has a VCF — route to `variant-annotation` or `clinical-variant-reporter`. - The user wants comparison between two FASTA — route to `genome-compare`. - The user wants 3D structure prediction — route to `struct-predictor`.

## Why This Exists

- **Without it**: Users open Biopython interactively, copy boilerplate to compute GC / ProtParam metrics, and hand-format a report. Common values get computed inconsistently across notebooks. - **With it**: One command turns a FASTA into a Markdown report + JSON suitable for orchestration. Detection of nucleotide vs protein is automatic. ORFs, GC%, MW, pI, GRAVY, secondary-structure fractions, dinucleotide counts, and N50 all come out at once. - **Why ClawBio**: Output is structured (`result.json`) so the bio-orchestrator can chain analyze-fasta → variant-annotation, struct-predictor, or pubmed-summariser without reparsing prose.

## Core Capabilities

1. **Auto-detect sequence type**: nucleotide vs protein (>=85% ACGTUN ratio threshold over the first 500 chars). 2. **Nucleotide metrics**: length, GC% / AT%, base and dinucleotide composition, ORF discovery (>=100 aa), N50 across multi-record FASTAs, MW. 3. **Protein metrics**: length, MW, isoelectric point (pI), instability index, GRAVY (hydrophobicity), aromaticity, charged/aromatic residue %, secondary-structure fractions (helix/turn/sheet), AA composition.

## Scope

**One skill, one task.** This skill describes a single FASTA file. It does not align, blast, fold, compare, or annotate. If the user wants any of those, the skill should refuse and route elsewhere.

## Input Formats

| Format | Extension | Required Fields | Example | |--------|-----------|-----------------|---------| | FASTA (nucleotide) | `.fasta`, `.fa`, `.fna` | `>header` line + ACGTUN sequence | `example_data/demo_nucleotide.fasta` | | FASTA (protein) | `.fasta`, `.fa`, `.faa` | `>header` line + amino-acid sequence | `example_data/demo_protein.fasta` |

## Workflow

When the user asks for FASTA analysis:

1. **Validate** (prescriptive): file exists; at least one record; first record >=10 chars; <=50% Ns. Any failure → exit 1 with explicit message. Never write a partial report. 2. **Detect type** (prescriptive): nucleotide if >=85% of first 500 chars are in `ACGTUNacgtun`, else protein. 3. **Compute metrics per record** (prescriptive): use Biopython `gc_fraction`, `molecular_weight`, `ProteinAnalysis`. Round consistently (GC to 2 dp, MW to 1 dp, pI to 2 dp). 4. **Generate** (prescriptive): write `result.json` (full structured data), `report.md` (human-readable), `report.html` (visual), and `reproducibility/{commands.sh,run.json}`. 5. **Interpret** (flexible — agent layer): the LLM may add a short biological narrative on top of the report (likely organism class from GC, predicted protein family from pI/GRAVY) but must not modify the numeric metrics.

## CLI Reference

```bash # Standard usage (ClawBio convention) python skills/analyze-fasta/analyze_fasta.py \ --input <fasta_file> --output <report_dir>

# Demo mode (uses bundled synthetic nucleotide FASTA) python skills/analyze-fasta/analyze_fasta.py --demo --output /tmp/analyze_fasta_demo

# Via ClawBio runner python clawbio.py run analyze-fasta --input <fasta_file> --output <dir> python clawbio.py run analyze-fasta --demo

# Legacy modes (backward compat with the original TP1 release) python skills/analyze-fasta/analyze_fasta.py <file.fasta> --json python skills/analyze-fasta/analyze_fasta.py <file.fasta> --html out.html ```

## Demo

```bash python clawbio.py run analyze-fasta --demo ```

Expected output: a `report.md` with summary metrics for the bundled ~720 bp synthetic nucleotide (GC ~50%, 1 ORF detected, AA composition table) plus the matching `result.json` and `reproducibility/` bundle.

## Algorithm / Methodology

So an LLM agent can apply the same logic without the script:

1. **Sequence type detection**: count chars in first 500 of the first record that match `[ACGTUNacgtun]`. Ratio >= 0.85 → nucleotide, else protein. (No silent fallback; if ambiguous, document in `result.json`.) 2. **Nucleotide GC**: `gc = (G + C) / (A + T + G + C + N) * 100`. Use Biopython `gc_fraction` to match the production behaviour. 3. **ORF discovery**: scan all 3 forward frames for `ATG ... [TAA|TAG|TGA]`. Keep ORFs with `length_bp >= 300` (>= 100 aa). 4. **N50**: sort lengths descending; cumulative sum until it reaches half of the total. Length at that point is N50. 5. **Protein metrics**: Biopython `ProteinAnalysis`. Strip `X` and `*` before instantiating to avoid ProtParam errors. 6. **Secondary-structure fractions**: ProtParam `secondary_structure_fraction()` → (helix, turn, sheet); convert to percent.

**Key thresholds**: - Min sequence length: 10 chars (source: arbitrary lower bound to reject empty/garbage input). - Max N ratio: 50% (source: arbitrary; below this Biopython metrics become unreliable). - ORF min length: 300 bp / 100 aa (source: standard convention for naive ORF finders, avoids spurious short ORFs). - Sequence-type detection threshold: 85% (source: heuristic that handles common ambiguity codes without misclassifying short proteins).

## Example Queries

- "Analyze sample.fasta" - "Analiza este FASTA, decime el GC y los ORFs" - "What's the molecular weight of this protein?" - "Compute pI of the FASTA in /tmp/x.fa"

## Example Output

```markdown # analyze-fasta Report

**Input file:** `demo_nucleotide.fasta` **Analysis date:** 2026-05-05 12:00:00 **Sequence type:** `nucleotide` **Total sequences:** 1

## Summary

| Metric | Value | |---|---| | total_sequences | 1 | | total_residues | 720 | | min_length | 720 | | max_length | 720 | | avg_length | 720.0 | | n50 | 720 | | avg_gc_content | 50.42 | | total_orfs | 1 |

## Per-sequence metrics

### 1. synthetic_demo_orf

- **Description:** synthetic_demo_orf | Synthetic E. coli-like ORF - **Length:** 720 bp - **GC content:** 50.42% - **AT content:** 49.58% - **ORFs (>=100 aa):** 1

---

_ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions._ ```

## Output Structure

``` <output_dir>/ ├── report.md # Primary markdown report ├── report.html # Standalone visual report ├── result.json # Machine-readable results └── reproducibility/ ├── commands.sh # Exact command to reproduce └── run.json # Run metadata (versions, timestamps, input size) ```

## Dependencies

**Required**: - `biopython` >= 1.80; sequence parsing, ProtParam, gc_fraction, molecular_weight.

**Optional**: - None. The skill is intentionally lean; pure stdlib + Biopython.

## Gotchas

- **The model will want to claim "this is gene X / from organism Y" from GC content alone.** Do not. GC is a weak signal — many taxa overlap. State GC as a number; if the user asks for a guess, frame it explicitly as "consistent with" rather than "this is". - **The model will treat ORFs >100 aa as proof of coding.** Do not. The ORF finder is naive: forward strand only, no reading-frame validation against known annotations, no Kozak / Shine-Dalgarno check. Frame ORFs as candidates, never confirmed. - **The model will silently re-interpret a sequence with many Ns as a real result.** Do not. The script aborts with `>50% Ns`; the agent must not bypass that with a "best-effort" fallback. Surface the failure to the user. - **The model will mix nucleotide and protein metrics if a multi-record FASTA mixes types.** The skill detects type from the first record only. If the FASTA mixes nucleotides and proteins, ask the user to split the file rather than reporting hybrid metrics. - **The model will use the script's HTML output as the primary deliverable.** Use `report.md` for chaining; the HTML is a courtesy for human inspection only.

## Safety

- **Local-first**: no network calls; everything runs against the local file. - **Disclaimer**: every `report.md` includes the standard ClawBio research-tool disclaimer. - **Audit trail**: every run writes `reproducibility/run.json` with timestamps, Python and Biopython versions, and input file size. - **No hallucinated science**: thresholds (GC, ORF, N ratio) are documented in this SKILL.md; the agent must not invent new ones.

## Agent Boundary

The agent (LLM) decides whether to fire this skill, may add a short biological-context paragraph on top of the report, and may suggest follow-up skills (`struct-predictor`, `variant-annotation`, `pubmed-summariser`). The skill (Python) executes the metrics and writes the artefacts. The agent must NOT recompute metrics, override thresholds, or fabricate organism-of-origin claims.

## Integration with Bio Orchestrator

**Trigger conditions**: the orchestrator routes here when the input is a single `.fasta`/`

技术详情

版本
1.0.0
许可证
MIT
最近更新
2026年8月23日
发布时间
2026年8月23日

决策摘要

首选

89
就绪
采用
阶段

1,112 个 GitHub Stars

审计

安装审查

安装与采用审查

79
需审查
安全性
73/100
维护状态
100/100
安装
92/100
打开完整审计查看评估报告

Agent 验证证据

Agent 验证证据

来自解析、审查、安装和一次小范围运行后的结果报告。

0
已验证
Needs first agent run自动安装: 先审查最近: 未知
成功率
近期失败
结果
0
输出质量
失败
0
不相关
0
安装次数
0
风险拦截
0
需要配置
0
生产环境
0

暂时没有 Agent 结果数据。首次 Agent 执行可以通过 /api/agent/outcome 报告成功、需要设置、风险拦截、失败或不相关。

安装

加入 Agent 工作流

免费且开源. 在生产 Agent 中安装前请先审查报告。

增长闭环

分享工具包

X

为 analyze-fasta 准备的场景化草稿,可手动发布到 X。

策展说明
analyze-fasta: Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs...

1.1K stars

https://www.openagentskill.com/skills/clawbio-analyze-fasta?ref=x
打开 X 草稿
可选:带安装命令的回复
Listing + install path for analyze-fasta:
https://www.openagentskill.com/skills/clawbio-analyze-fasta?ref=x

Install: npx skills add ClawBio/ClawBio --skill analyze-fasta
打开回复草稿

收录来源

Registry 收录

可认领

此列表来自公开来源,维护者认领获批前不会标记为官方。

创作者
ClawBio
收录方
OpenAgentSkill 社区索引

归属链接指向公开仓库或创作者主页。创作者可认领列表以更新所有权信号。

认领此 Skill

所有者认领

认领此 Skill 页面

这条 Registry 收录 列表归属于 ClawBio,但尚未标记为官方。认领后可增加已验证所有者信号,使后续发布、安装和审计更新更值得信赖。

创作者外链工具包

将证据徽章加入你的 README

在开发者评估仓库的位置展示规范页面、当前信任与审计信号,以及真实的 Agent 验证证据。

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/clawbio-analyze-fasta?metric=listed&label=Listed)](https://www.openagentskill.com/skills/clawbio-analyze-fasta)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/clawbio-analyze-fasta?metric=trust&label=Trust)](https://www.openagentskill.com/skills/clawbio-analyze-fasta)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/clawbio-analyze-fasta?metric=audit&label=Audit)](https://www.openagentskill.com/skills/clawbio-analyze-fasta/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/clawbio-analyze-fasta?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/clawbio-analyze-fasta)

作者

C

ClawBio

@clawbio

平台适配

健康信号

GitHub Stars
1.1K
质量评分
45/100
最近 GitHub 推送
2026年8月23日
框架提示
未知
OpenAgentSkill 浏览量
1
复制安装命令
0
跳转点击
0

社区信号

告诉我们这个 Skill 是否对你的 Agent 工作流有帮助。汇总反馈会持续改善排序。

信任与安全

仅限沙盒

61
  • GitHub 采用度1.1K 个 GitHub Stars通过
  • Star/Fork 活跃度1.1K 个 Star,257 个 Fork; 当前元数据中没有议题活跃度信息通过
  • 近期维护今天有推送通过
  • 许可证清晰度MIT通过
  • README/SKILL.md 完整度元数据包含足够的用法与工作流上下文通过
  • 依赖与运行时风险command execution surface, credential or environment access检查