vibe-science

审查 · 59
已收录

Scientific research engine with adversarial review, tree search, and serendipity detection. Use when: exploring hypotheses, validating findings against literature, running computational experiments with quality gates, or hunting for unexpected discoveries. Do NOT use for simple Q

Verified installs0
Stars16
版本1.0.0
质量59/100 · 有潜力
信任59/100 · Do not auto-install
审计74/100 · 需审查

供给资产档案

研究与知识工作

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

浏览赛道

场景

研究 Agent

I need my agent to research a topic, compare sources, and produce a concise report.

适配 Agent

Claude Code + CLI + Codex

适用于 Codex、Claude Code、Cursor、CLI 或自定义 Agent。

安装

就绪

npx skills add th3vib3coder/vibe-science --skill vibe-science

维护状态

新鲜

距上次推送 3 天

风险

需审查

Financial research output is not financial advice; require human review before any live investment decision

GitHub 质量

16

59/100 质量 · 67/100 信任

覆盖标签

研究研究 Agentagent-skill

审查说明

Financial research output is not financial advice; require human review before any live investment decision · No explicit safe operating boundaries or security considerations are documented in the provided SKILL.md excerpt.

Agent 采用评分卡

一眼查看信任、审计与安装准备度

这些分数综合公开仓库元数据、OpenAgentSkill 审查信号、维护新鲜度与安装准备度。它用于候选筛选,不替代人工审查。

质量

有潜力
59

有用的候选项,但采用前应与替代方案比较。

信任

Do not auto-install
59

Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.

审计

需审查
74

对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。

OpenAgentSkill 信任评分 v5

安装前需人工审查

Choose a stronger alternative or inspect the source manually before any install attempt.

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

16 个 GitHub Stars

仓库活跃度

16 个 Star,0 个 Fork

维护状态

距上次推送 3 天

许可证

Apache-2.0

安装

npx skills add th3vib3coder/vibe-science --skill vibe-science

安装安全性

标准软件包或运行时安装路径

权限范围

filesystem or document access, database access

Agent 结果

暂未有 Agent 结果数据

文档

Usable metadata, review docs

风险摘要

生产前审查

  • No explicit safe operating boundaries or security considerations are documented in the provided SKILL.md excerpt.
  • Financial research output is not financial advice; require human review before any live investment decision.
  • Low GitHub adoption signal
  • Quality score needs review

安装准备度

安装路径可用

  • 安装路径可用
  • 仓库证据可用
  • 已声明许可证
  • 暂无 Agent 验证结果证据

Agent 可读元数据

这个 Skill 的机器可读决策数据。

使用此区块或内嵌 JSON 判断 Agent 是否应安装该 Skill、选择替代方案,或先请求人工审查。

打开 JSON

适用任务

  • 研究 Agent 工作流
  • Claude Code 团队
  • builders willing to evaluate younger projects
  • 检索来源

适用 Agent

CodexClaude CodeCursorOpenAgentSkill CLICLI

安装决策

命令
npx skills add th3vib3coder/vibe-science --skill vibe-science
策略
审查
人工审查

信任与风险

信任
59/100
审计
74/100
风险级别
需审查

结果闭环

端点
/api/agent/outcome
事件 ID
resolve
结果
5

安装命令

npx skills add th3vib3coder/vibe-science --skill vibe-science

不适用场景

  • 需要厂商支持 SLA 的团队
  • production agents without a repository review
  • Low GitHub adoption signal
  • No explicit safe operating boundaries or security considerations are documented in the provided SKILL.md excerpt.
  • Financial research output is not financial advice; require human review before any live investment decision

Agent 安全 v2

54/100 · 避免自动安装

实验性审查

Sparse or mixed signals. Useful for discovery, but not for autonomous installation.

Test manually in an isolated workspace and compare against safer alternatives.

通过 API 解析

网络访问

Skill 可能访问远程页面、API、仓库或外部服务。

文件系统访问

Skill 可能读取或写入项目文件、文档、生成产物或本地工作区状态。

数据库访问

Skill 可能检查 Schema、查询数据库或处理持久化存储。

  • Financial research output is not financial advice; require human review before any live investment decision

安装目标

在你的 Agent 工作流中安装此 Skill

通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install th3vib3coder-vibe-science

Agent 解析计划

让 Agent 在安装前验证匹配度。

Resolve API 返回首选 Skill、替代方案、安全策略、审计说明、安装目标和可直接执行的提示词,无需抓取此页面。

打开文本计划

Agent 应检查

  • 从 Resolve API 检查任务匹配与替代方案。
  • 检查审计评分、信任评分和安全策略警告。
  • 检查 Codex、Claude Code、Cursor 或 CLI 的安装目标兼容性。

复制提示词

Task: Use vibe-science in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20vibe-science%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/th3vib3coder-vibe-science/install
Install command: npx skills add th3vib3coder/vibe-science --skill vibe-science
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 交接

把安装路径交给 Agent,而不是再给一个目录页。

通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。

打开安装 API

Agent 提示词

Use vibe-science for this task. Review https://www.openagentskill.com/api/skills/th3vib3coder-vibe-science/install, then install with: npx skills add th3vib3coder/vibe-science --skill vibe-science

Registry 元数据

用于自动选择 Skill 的 Agent 可读档案。

本页通过 Registry API 提供相同的决策、信任、审计、场景和安装信号,让 Agent 无需抓取界面即可排序。

打开 Manifest

适配 Agent

61/100

研究 Agent

平台

Claude Code

审计报告

需审查 · 74/100

对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。

查看审计报告查看评估报告

Agent 决策面板

Fallback candidate for Research agents

先用此 Skill 做原型验证,并保留备选方案。

61
就绪度
原型验证
阶段

栈中角色

备选候选

主要匹配

研究 Agent

信任标签

先做原型验证

安装路径

命令已就绪

适用场景

  • 研究 Agent 工作流
  • Claude Code 团队
  • builders willing to evaluate younger projects

证据

  • 仓库近期活跃
  • 已提供安装命令或 GitHub 仓库
  • 59/100 质量档案
  • 10 个 OpenAgentSkill 交互事件

先审查

  • Low GitHub adoption signal
  • No explicit safe operating boundaries or security considerations are documented in the provided SKILL.md excerpt.

实施路径

  1. 1在沙盒 Agent 中安装它,并端到端完成一次研究 Agent任务。
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

信任档案

Do not auto-install

Trust Score v5 found insufficient evidence for agent installation. Treat this as discovery material, not an executable recommendation.

59
OpenAgentSkill 信任评分

GitHub 采用度

修复

16 个 GitHub Stars

Star/Fork 活跃度

修复

16 个 Star,0 个 Fork; 当前元数据中没有议题活跃度信息

近期维护

通过

距上次推送 3 天

许可证清晰度

通过

Apache-2.0

积极信号

  • AI 审查已通过
  • 安装路径可用
  • 仓库证据可用
  • 近期维护的仓库
  • 安装命令未发现明显高风险模式
  • 结果闭环已就绪,但需要首次真实 Agent 运行

安装前审查

  • No explicit safe operating boundaries or security considerations are documented in the provided SKILL.md excerpt.
  • Financial research output is not financial advice; require human review before any live investment decision.
  • Low GitHub adoption signal
  • Quality score needs review
  • GitHub adoption: 16 GitHub stars
  • Stars/forks activity: 16 stars, 0 forks; issue activity unavailable in current metadata
  • 暂未有真实 Agent 结果报告
  • 无人值守安装前需要人工审查

建议操作

Choose a stronger alternative or inspect the source manually before any install attempt.

质量档案

有潜力 适用于 Agent 工作流的候选

有用的候选项,但采用前应与替代方案比较。

59
GitHub Stars
16
新鲜度
3 天前
安装就绪
许可证
Apache-2.0
安装前审查: Low GitHub adoption signal · No explicit safe operating boundaries or security considerations are documented in the provided SKILL.md excerpt.

工作流匹配

在这些场景使用此 Skill

工作流匹配

加入完整工作流

替代方案短名单

安装前对比

可能适合该任务的相近 Skill。

对比全部

概览

--- name: vibe-science description: "Scientific research engine with adversarial review, tree search, and serendipity detection. Use when: exploring hypotheses, validating findings against literature, running computational experiments with quality gates, or hunting for unexpected discoveries. Do NOT use for simple Q&A, code editing, or non-research tasks." skill-author: th3vib3coder license: Apache-2.0 ---

# Vibe Science v5.0 — IUDEX

> Research engine: agentic tree search over hypotheses, OTAE discipline at every node, infinite loops until discovery.

---

## WHY THIS SKILL EXISTS — READ THIS FIRST

This section is not optional. It is not a preamble. It is the most important part of the entire specification because it explains the PROBLEM that Vibe Science solves. Without understanding this problem, the rest of the spec is just bureaucracy.

### The Problem: AI Agents Are Dangerous in Science

An AI agent given a research task will:

1. **Optimize for completion, not truth.** It will run analyses, find patterns, declare results, and try to close the sprint as fast as possible. This is the agent's default disposition: shipping feels like success.

2. **Get excited by strong signals.** A p-value of 10⁻¹⁰⁰ feels like a discovery. An OR of 2.30 feels publishable. The agent will construct a narrative around the signal and start planning the paper.

3. **Not search for what kills its own claims.** The agent will not spontaneously search for "is this a known artifact?", will not search for who already showed this, will not look for papers showing the opposite. It confirms, it doesn't demolish.

4. **Not crystallize intermediate results.** The agent works in a context window that gets erased. Results that exist only in the conversation are lost. The agent says "I'll remember this" — it won't.

5. **Declare "done" prematurely.** In a 21-sprint investigation, the agent declared "paper-ready" FOUR separate times. Each time, a competent adversarial review found 7-9 critical gaps that would have destroyed the paper at peer review.

This is not a theoretical risk. This happened. Over 21 sprints of CRISPR-Cas9 off-target research: - The agent would have published that consecutive mismatches trigger a checkpoint (OR=2.30, p < 10⁻¹⁰⁰). **It was completely confounded** — propensity matching reversed the sign. - The agent would have published "bidirectional positional effects." **It was biologically impossible** — ALL mismatches reduce cleavage. - The agent would have published the regime switch as a strong finding. **Cohen's d was 0.07** — noise. - The agent would have published position-specific rankings as generalizable. **They don't generalize** between assays.

None of these claims were hallucinations. The data was real. The statistics were correct. The narratives were plausible. The problem was that the agent NEVER ASKED: "What if this is an artifact? Who has already shown this? What confounder would explain this away?"

### The Solution: Reviewer 2 as Disposition, Not Gate

Vibe Science exists to solve this problem. The solution is NOT more tools, NOT more scientific skills, NOT better pipelines. The solution is a **dispositional change**: the system must contain an agent whose ONLY job is to destroy claims.

This agent — Reviewer 2 — is not a quality gate that you pass. It is a co-pilot whose disposition is the OPPOSITE of the builder's:

| | Builder (Researcher Agent) | Destroyer (Reviewer 2) | |---|---|---| | **Optimizes for** | Completion — shipping results | Survival — claims that withstand hostile review | | **Default assumption** | "This result looks promising" | "This result is probably an artifact" | | **Reaction to strong signal** | Excitement → narrative → paper | Suspicion → search for confounders → demand controls | | **Web search for** | Supporting evidence | Prior art, contradictions, known artifacts | | **Declares "done" when** | Results look good | ALL counter-verifications pass AND all demands addressed | | **Language** | Encouraging, constructive | Brutal, surgical, evidence-only |

This asymmetry is not a bug — it is the entire architecture. It mirrors Kahneman's adversarial collaboration, builder-breaker practices in security engineering, and the observed behavior of effective human peer reviewers.

### What Reviewer 2 MUST Do at Every Intervention

Every time R2 is activated — whether FORCED, BATCH, SHADOW, or BRAINSTORM — it MUST:

1. **SEARCH BEFORE JUDGING.** Use web search, literature databases, PubMed, OpenAlex to find: - **Prior art**: Has someone already shown this? → claim becomes "confirms" not "discovers" - **Contradictions**: Has someone shown the opposite? → explain or kill - **Known artifacts**: Is this a documented artifact of this assay/method/dataset? - **Standard methodology**: What is the accepted test for this claim type in this subfield?

2. **DEMAND THE CONFOUNDER HARNESS.** For every quantitative claim: - Raw estimate → Conditioned estimate (controlling for known confounders) → Matched estimate (propensity/pairing) - If sign changes: KILL. If collapses >50%: DOWNGRADE. If survives: PROMOTABLE.

3. **REFUSE TO CLOSE.** Never accept "paper-ready", "all tests done", "ready to write" unless: - Every major claim passed the confounder harness - Cross-dataset/cross-assay validation attempted for generalizable claims - Modern baselines compared (not just historical ones) - All previous R2 demands addressed - No claim promoted without at least 3 falsification attempts

4. **TURN INCIDENTS INTO FRAMEWORKS.** When a flaw is caught (e.g., confounded claim), don't just fix that one instance. Demand the same check for ALL similar claims. Every incident becomes a protocol.

5. **CRYSTALLIZE EVERYTHING.** Demand that every result, every decision, every kill is written to a file. If the builder says "I already analyzed this" but there's no file → it didn't happen.

6. **ESCALATE, NEVER SOFTEN.** Each review pass must be MORE demanding than the last. If pass N found 5 issues, pass N+1 must look for issues that pass N missed. A review that finds fewer issues is suspicious.

### What Happens Without This

Without Rev2 as disposition (not just gate), the system produces: - Papers with confounded claims that survive internal review but are destroyed by the first competent peer reviewer - "Discoveries" that are already known artifacts in the field - Strong p-values on effects that disappear when you control for the obvious confounder

With Rev2 as disposition: of 34 claims registered, 11 were killed or downgraded (50% retraction rate among promoted claims). The most dangerous claim (OR=2.30, p < 10⁻¹⁰⁰) was caught in ONE sprint. Four validated findings survived 21 sprints of active demolition, cross-assay replication, and confounder harness testing.

### The Three Principles

1. **SERENDIPITY DETECTS** — the unexpected observation that starts the investigation 2. **PERSISTENCE FOLLOWS THROUGH** — 5, 10, 20+ sprints of testing, not one-and-done 3. **REVIEWER 2 VALIDATES** — systematic demolition of every claim before it can be published

All three are necessary. Serendipity without persistence is a footnote. Persistence without Rev2 is confirmation bias running for 20 sprints. Rev2 without serendipity misses the discoveries worth reviewing.

This is what Vibe Science must be. Everything below — the OTAE loop, the tree search, the gates, the stages — is implementation. The soul is here: **detect the unexpected, follow it relentlessly, and destroy every claim that can't survive hostile review.**

---

## CONSTITUTION (Immutable — Never Override)

**LAW 1: DATA-FIRST** — No thesis without evidence from data. If data doesn't exist, the claim is a HYPOTHESIS to test, not a finding. `NO DATA = NO GO.`

**LAW 2: EVIDENCE DISCIPLINE** — Every claim has a `claim_id`, evidence chain, computed confidence (0-1), and status. Claims without sources are hallucinations.

**LAW 3: GATES BLOCK** — Quality gates are hard stops, not suggestions. Pipeline cannot advance until gate passes. Fix first, re-gate, then continue. 27 gates total (8 schema-enforced in v5.0).

**LAW 4: REVIEWER 2 IS CO-PILOT** — R2 is not a gate you pass — it is a co-pilot you cannot fire. R2 can VETO any finding, REDIRECT any branch, FORCE re-investigation. Its demands are non-negotiable. R2 reviews brainstorm output, tree strategy, claims, and conclusions. No exceptions.

**LAW 5: SERENDIPITY IS THE MISSION** — Serendipity is not a side-effect — it is the primary engine of discovery. Actively hunt for the unexpected at every cycle. Serendipity Radar runs at every EVALUATE. Score >= 10 → QUEUE. Score >= 15 → INTERRUPT. A session with zero flags is suspicious.

**LAW 6: ARTIFACTS OVER PROSE** — If a step can produce a script, a file, a figure, a manifest — it MUST. Prose descriptions of what "should" happen are insufficient.

**LAW 7: FRESH CONTEXT RESILIENCE** — The system MUST be resumable from `STATE.md` + `TREE-STATE.json` alone. All context lives in files, never in chat history.

**LAW 8: EXPLORE BEFORE EXPLOIT** — Minimum 3 draft nodes before any is promoted. Exploration ratio >= 20% at T3. A tree with one branch is a list — lists miss discoveries.

**LAW 9: CONFOUNDER HARNESS** — Every quantitative claim MUST pass: raw → conditioned → matched. Sign change = **ARTIFACT** (killed). Collapse >50% = **CONFOUNDED** (downgraded). Survives = **ROBUST** (promotable). `NO HARNESS = NO CLAIM.`

**LAW 10: CRYSTALLIZE OR LOSE** — Every result, decision, pivot, kill MUST be written to a persistent file. The context window is a buffer that gets erased — it is NOT memory. `IF IT'S NOT IN A FILE, IT DOESN'T EXIST.`

> Full constitution with role-specific constraints: `references/constitution.md`

---

## v5.0 INNOVATIONS — IUDEX

v5.0 makes R2 structurally unbypassable. Based on Huang et al. (ICLR 2024): LLMs cannot self-correct reasoning without external feedback.

| Innovation | What | Protocol | Gate | |-----------|------|----------|------| | Seeded Fault Injection (SFI) | Orchestrator injects known faults before FORCED R2 reviews. R2 must catch them. | `references/seeded-fault-injection.md` | V0: RMS >= 0.80, FAR <= 0.10 | | Judge Agent (R3) | Meta-reviewer scores R2's quality on 6-dimension rubric | `references/judge-agent.md` | J0: total >= 12/18, no dim = 0 | | Blind-First Pass (BFP) | R2 sees claims without justifications first, breaks anchoring | `references/blind-first-pass.md` | — | | Schema-Validated Gates (SVG) | 8 critical gates enforce structure via JSON Schema | `references/schema-validation.md` | — | | Circuit Breaker | Same objection x 3 rounds → DISPUTED. Frozen, not killed. | `references/circuit-breaker.md` | — | | R2 Salvagente | Killed claims (INSUFFICIENT/CONFOUNDED/PREMATURE) must produce serendipity seed | `references/serendipity-engine.md` | — | | Confidence formula | E x D x (R_eff x C_eff x K_eff)^(1/3) with hard veto + dynamic floor | `references/evidence-engine.md` | — | | Agent Permission Model | R2 writes verdicts, orchestrator writes ledger. Separation of powers. | `references/constitution.md` | — |

---

## When to Use

- Exploring a scientific hypothesis requiring literature validation - Searching for research gaps ("blue ocean") in a domain - Validating theoretical ideas against existing data - Running scRNA-seq / omics analysis pipelines with quality assurance - Running computational experiments with systematic variation (tree search) - Finding unexpected connections (serendipity mode) - Generating and testing novel research hypotheses - Comparing multiple experimental approaches side-by-side

---

## SESSION INITIALIZATION

### Announce at Start

Display this banner, then the session info:

``` . * . * . * * . * . . * . . * . * . . *

██╗ ██╗██╗██████╗ ███████╗ ██║ ██║██║██╔══██╗██╔════╝ ██║ ██║██║██████╔╝█████╗ ╚██╗ ██

技术详情

版本
1.0.0
许可证
Apache-2.0
最近更新
2026年8月20日
发布时间
2026年8月20日

决策摘要

备选候选

61
就绪
原型验证
阶段

仓库近期活跃

审计

安装审查

安装与采用审查

74
需审查
安全性
79/100
维护状态
100/100
安装
92/100
打开完整审计查看评估报告

Agent 验证证据

Agent 验证证据

来自解析、审查、安装和一次小范围运行后的结果报告。

0
已验证
Needs first agent run自动安装: 先审查最近: 未知
成功率
近期失败
结果
0
输出质量
失败
0
不相关
0
安装次数
0
风险拦截
0
需要配置
0
生产环境
0

暂时没有 Agent 结果数据。首次 Agent 执行可以通过 /api/agent/outcome 报告成功、需要设置、风险拦截、失败或不相关。

安装

加入 Agent 工作流

免费且开源. 在生产 Agent 中安装前请先审查报告。

增长闭环

分享工具包

X

为 vibe-science 准备的场景化草稿,可手动发布到 X。

策展说明
vibe-science: Scientific research engine with adversarial review, tree search, and serendipity detection. U...

16 stars

https://www.openagentskill.com/skills/th3vib3coder-vibe-science?ref=x
打开 X 草稿
可选:带安装命令的回复
Listing + install path for vibe-science:
https://www.openagentskill.com/skills/th3vib3coder-vibe-science?ref=x

Install: npx skills add th3vib3coder/vibe-science --skill vibe-science
打开回复草稿

收录来源

Registry 收录

可认领

此列表来自公开来源,维护者认领获批前不会标记为官方。

创作者
th3vib3coder
收录方
OpenAgentSkill 社区索引

归属链接指向公开仓库或创作者主页。创作者可认领列表以更新所有权信号。

认领此 Skill

所有者认领

认领此 Skill 页面

这条 Registry 收录 列表归属于 th3vib3coder,但尚未标记为官方。认领后可增加已验证所有者信号,使后续发布、安装和审计更新更值得信赖。

创作者外链工具包

将证据徽章加入你的 README

在开发者评估仓库的位置展示规范页面、当前信任与审计信号,以及真实的 Agent 验证证据。

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/th3vib3coder-vibe-science?metric=listed&label=Listed)](https://www.openagentskill.com/skills/th3vib3coder-vibe-science)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/th3vib3coder-vibe-science?metric=trust&label=Trust)](https://www.openagentskill.com/skills/th3vib3coder-vibe-science)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/th3vib3coder-vibe-science?metric=audit&label=Audit)](https://www.openagentskill.com/skills/th3vib3coder-vibe-science/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/th3vib3coder-vibe-science?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/th3vib3coder-vibe-science)

作者

T

th3vib3coder

@th3vib3coder

平台适配

健康信号

GitHub Stars
16
质量评分
32/100
最近 GitHub 推送
2026年8月19日
框架提示
未知
OpenAgentSkill 浏览量
10
复制安装命令
0
跳转点击
0

社区信号

告诉我们这个 Skill 是否对你的 Agent 工作流有帮助。汇总反馈会持续改善排序。

信任与安全

Do not auto-install

59
  • GitHub 采用度16 个 GitHub Stars修复
  • Star/Fork 活跃度16 个 Star,0 个 Fork; 当前元数据中没有议题活跃度信息修复
  • 近期维护距上次推送 3 天通过
  • 许可证清晰度Apache-2.0通过
  • README/SKILL.md 完整度公开元数据需要更完整的 README/SKILL.md 上下文信息
  • 依赖与运行时风险公开元数据中未发现主要依赖风险提示通过