agent-harness

审查 · 73
已收录

Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until eve

Verified installs0
Stars24.8K
版本1.0.0
质量91/100 · 优秀
信任73/100 · 仅限沙盒
审计87/100 · 需审查

供给资产档案

研究与知识工作

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

浏览赛道

场景

研究 Agent

I need my agent to research a topic, compare sources, and produce a concise report.

适配 Agent

Claude Code + OpenAI Agents + CLI

适用于 Codex、Claude Code、Cursor、CLI 或自定义 Agent。

安装

就绪

npx skills add alirezarezvani/claude-skills --skill agent-harness

维护状态

新鲜

今天有推送

风险

需审查

Financial research output is not financial advice; require human review before any live investment decision

GitHub 质量

25K

91/100 质量 · 81/100 信任

覆盖标签

研究研究 Agentagent-skill

审查说明

Financial research output is not financial advice; require human review before any live investment decision · The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.

Agent 采用评分卡

一眼查看信任、审计与安装准备度

这些分数综合公开仓库元数据、OpenAgentSkill 审查信号、维护新鲜度与安装准备度。它用于候选筛选,不替代人工审查。

质量

优秀
91

高置信候选,具有较强的采用度与健康维护信号。

信任

仅限沙盒
73

有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。

审计

需审查
87

对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。

OpenAgentSkill 信任评分 v5

安装前需人工审查

仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。

CodexClaude CodeCursorOpenAgentSkill CLI

Stars

25K 个 GitHub Stars

仓库活跃度

25K 个 Star,3.5K 个 Fork

维护状态

今天有推送

许可证

MIT

安装

npx skills add alirezarezvani/claude-skills --skill agent-harness

安装安全性

标准软件包或运行时安装路径

权限范围

shell or command execution, filesystem or document access

Agent 结果

暂未有 Agent 结果数据

文档

README/SKILL.md 上下文充分

风险摘要

生产前审查

  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • Financial research output is not financial advice; require human review before any live investment decision.
  • Quality score needs review

安装准备度

安装路径可用

  • 安装路径可用
  • 仓库证据可用
  • 已声明许可证
  • 暂无 Agent 验证结果证据

Agent 可读元数据

这个 Skill 的机器可读决策数据。

使用此区块或内嵌 JSON 判断 Agent 是否应安装该 Skill、选择替代方案,或先请求人工审查。

打开 JSON

适用任务

  • 研究 Agent 工作流
  • Claude Code 团队
  • 重视 GitHub 采用信号的团队
  • 检索来源

适用 Agent

CodexClaude CodeCursorOpenAgentSkill CLIOpenAI AgentsCLI

安装决策

命令
npx skills add alirezarezvani/claude-skills --skill agent-harness
策略
审查
人工审查

信任与风险

信任
73/100
审计
87/100
风险级别
需审查

结果闭环

端点
/api/agent/outcome
事件 ID
resolve
结果
5

安装命令

npx skills add alirezarezvani/claude-skills --skill agent-harness

不适用场景

  • 需要厂商支持 SLA 的团队
  • production agents without a repository review
  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • 暂未有 OpenAgentSkill 使用反馈数据
  • 高风险权限提示:Shell 或命令执行

Agent 安全 v2

59/100 · 安装前审查

已审查并附权限说明审查

可用候选,但 Agent 在安装前应展示权限与审计说明。

在真实工作区安装前需要人工批准。

通过 API 解析

Shell 或命令执行

Skill 元数据引用了终端、CLI、Shell、子进程或命令执行工作流。

网络访问

Skill 可能访问远程页面、API、仓库或外部服务。

文件系统访问

Skill 可能读取或写入项目文件、文档、生成产物或本地工作区状态。

  • 高风险权限提示:Shell 或命令执行
  • Financial research output is not financial advice; require human review before any live investment decision

安装目标

在你的 Agent 工作流中安装此 Skill

通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install alirezarezvani-agent-harness

Agent 解析计划

让 Agent 在安装前验证匹配度。

Resolve API 返回首选 Skill、替代方案、安全策略、审计说明、安装目标和可直接执行的提示词,无需抓取此页面。

打开文本计划

Agent 应检查

  • 从 Resolve API 检查任务匹配与替代方案。
  • 检查审计评分、信任评分和安全策略警告。
  • 检查 Codex、Claude Code、Cursor 或 CLI 的安装目标兼容性。

复制提示词

Task: Use agent-harness in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install
Install command: npx skills add alirezarezvani/claude-skills --skill agent-harness
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 交接

把安装路径交给 Agent,而不是再给一个目录页。

通过公开安装端点获取命令、安全清单、目标提示词和该 Skill 的规范链接。

打开安装 API

Agent 提示词

Use agent-harness for this task. Review https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install, then install with: npx skills add alirezarezvani/claude-skills --skill agent-harness

Registry 元数据

用于自动选择 Skill 的 Agent 可读档案。

本页通过 Registry API 提供相同的决策、信任、审计、场景和安装信号,让 Agent 无需抓取界面即可排序。

打开 Manifest

适配 Agent

100/100

研究 Agent

平台

Claude Code, OpenAI Agents

审计报告

需审查 · 87/100

对安装准备度、安全元数据、维护情况与采用风险的机器可读审查。

查看审计报告查看评估报告

Agent 决策面板

适合 研究 Agent 的首选

将其作为优先候选,再在你的 Agent 环境中验证 README 与安装路径。

100
就绪度
采用
阶段

栈中角色

首选

主要匹配

研究 Agent

信任标签

可用于生产

安装路径

命令已就绪

适用场景

  • 研究 Agent 工作流
  • Claude Code 团队
  • 重视 GitHub 采用信号的团队

证据

  • 24,795 个 GitHub Stars
  • 仓库近期活跃
  • 已提供安装命令或 GitHub 仓库
  • 91/100 质量档案

先审查

  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • 暂未有 OpenAgentSkill 使用反馈数据

实施路径

  1. 1在沙盒 Agent 中安装它,并端到端完成一次研究 Agent任务。
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

信任档案

仅限沙盒

有用但信任信号不足或混杂的候选项。在结果闭环证明任务匹配前,请保持在隔离工作区内使用。

73
OpenAgentSkill 信任评分

GitHub 采用度

通过

25K 个 GitHub Stars

Star/Fork 活跃度

通过

25K 个 Star,3.5K 个 Fork; 当前元数据中没有议题活跃度信息

近期维护

通过

今天有推送

许可证清晰度

通过

MIT

积极信号

  • AI 审查已通过
  • 安装路径可用
  • 仓库证据可用
  • 近期维护的仓库
  • Large GitHub adoption signal
  • 安装命令未发现明显高风险模式
  • 结果闭环已就绪,但需要首次真实 Agent 运行

安装前审查

  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • Financial research output is not financial advice; require human review before any live investment decision.
  • Quality score needs review
  • 暂未有真实 Agent 结果报告
  • 无人值守安装前需要人工审查

建议操作

仅在沙盒中运行,并在用于真实工作前比较接近的替代方案。

质量档案

优秀 适用于 Agent 工作流的候选

高置信候选,具有较强的采用度与健康维护信号。

91
GitHub Stars
25K
新鲜度
今天
安装就绪
许可证
MIT
安装前审查: The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.

工作流匹配

在这些场景使用此 Skill

工作流匹配

加入完整工作流

替代方案短名单

安装前对比

可能适合该任务的相近 Skill。

对比全部

概览

--- name: agent-harness description: "Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived. Use when you want an agent or subagent to pick up a goal and drive it to a verified close across one of this repo's 18 domains ('run this goal through the engineering harness', 'set up an agentic loop for marketing work', 'make the finance domain self-verifying'). NOT for authoring Claude Code Workflow-tool .js scripts (workflow-builder), N-agent tournaments on one task (agenthub), single-file metric optimization (autoresearch-agent), or discovering published loop recipes (loop-library)." ---

# Agent Harness

You are a harness operator, not a hero. The loop — not your optimism — decides when work is done. Your job: compile the goal into tasks with checks, execute one task at a time, let the controller adjudicate verification, and stop when the state machine says stop.

## The contract

``` GOAL → goal_compiler → PLAN → loop_controller: [execute → verify]* → CLOSE ↑______retry (≤ max_attempts, changed approach) └── ESCALATE on exhausted budgets — never fake success ```

Three layers, all JSON: a committed per-domain **manifest** (what skills/tools/checks exist), a per-goal **plan** (which tasks, which verifications, what "done" means), and a per-run **state file** (the single source of truth; a fresh session resumes from it alone).

## Quick start

```bash # 0. Pick the domain manifest (18 committed under assets/harnesses/, e.g. engineering-team.json) ls assets/harnesses/

# 1. Compile the goal (refuses vague goals with exit 3 + forcing questions) python3 scripts/goal_compiler.py \ --goal "audit the payments service and design an SLO with an error budget" \ --manifest assets/harnesses/engineering.json --out plan.json

# 2. Initialize the loop state python3 scripts/loop_controller.py init --plan plan.json --state .agent-harness/state.json

# 3. Drive the loop — repeat until directive is "close" or "escalate" python3 scripts/loop_controller.py next --state .agent-harness/state.json # → {"action": "execute", "task": "T1", ...}: open the task's skill (SKILL.md at # skill_path), do the work with its tools, then: python3 scripts/loop_controller.py record --state .agent-harness/state.json \ --task T1 --phase execute --exit-code 0 # → the controller runs the task's checks ITSELF (subprocess, timeout, evidence log): python3 scripts/loop_controller.py verify --state .agent-harness/state.json --task T1 --cwd <repo-root>

# 4. Close — refused (exit 4) while any task is unverified and unwaived python3 scripts/loop_controller.py close --state .agent-harness/state.json ```

Regenerate a manifest after skills change (diff-stable, CI-checkable):

```bash python3 scripts/harness_manifest_builder.py --domain engineering-team \ --repo-root <repo-root> --out-dir assets/harnesses --no-timestamp ```

## Hard rules

1. **Never adjudicate your own verification.** `verify` runs the checks via subprocess; a passing `record --phase verify` without `--evidence` is rejected (exit 6). You do not get to declare a task verified. 2. **Never modify a gate you are judged by.** Check commands come from the manifest/plan. Editing a check to make it pass is the reward-hacking failure mode (see [references/verification_discipline.md](references/verification_discipline.md)) — same invariant as autoresearch-agent's locked evaluator. 3. **One task at a time, writes serialized.** Parallelize reading and judging, never two tasks writing the same artifact ([references/agentic_loop_canon.md](references/agentic_loop_canon.md)). 4. **Retry means a changed approach.** Same command + same input = same failure. The retry directive says so; honor it. 5. **Budgets are terminal states, not suggestions.** `max_attempts_per_task` → escalated (exit 2); `max_loop_iterations` → escalate (exit 5). Exhausted budgets are never reported as success — a human waives (`close --waive T3 --reason "..."`), you don't. 6. **Fresh context beats long context.** Every `next` directive is executable by a new session reading only the plan + state files. Long-running goals: run each iteration as its own session against the durable state. 7. **State lives in `.agent-harness/`** — never in `.agenthub/`, `.autoresearch/`, or `docs/TC/` (those belong to sibling skills). 8. **Plan and state files are a trust boundary.** `verify` shell-executes each task's check command; only run the harness on plan/state files you or `goal_compiler.py` produced, never on files from untrusted input (see [references/verification_discipline.md](references/verification_discipline.md)).

## Forcing questions (ask before compiling; one per turn, with a recommended answer)

| # | Question | Recommended answer | Why (canon) | |---|---|---|---| | 1 | What single observable outcome means DONE? | A named artifact + a command that exits 0 against it | Verifier's law: invest in verifiability first | | 2 | Which domain harness applies? | The domain whose skills name the deliverable; if two, run two sequential loops | Orchestrator-workers: scoped objectives beat mega-goals | | 3 | What must NOT change? | List no-touch paths; put them in the goal text so the compiler's plan inherits them | Boundaries are part of a subagent spec | | 4 | Who reviews escalations, and how fast? | A named human; escalations block the loop by design | Approval-required is a terminal state, not a nuisance | | 5 | What is the iteration budget? | Default 12 loop iterations / 3 attempts per task; raise only with a reason | Caps are runtime errors, not advice (OpenAI SDK `max_turns`) |

## Exit codes (branch on these mechanically)

| Code | Tool | Meaning | |---|---|---| | 0 | all | OK / directive emitted | | 2 | loop_controller | Escalation required — a human must review the evidence log | | 3 | goal_compiler | Goal too vague — answer the forcing questions, recompile | | 4 | goal_compiler / loop_controller | No skill matched / close refused (unverified tasks) | | 5 | loop_controller | Global iteration cap reached | | 6 | loop_controller | Invalid transition (recording on verified task, evidence missing, unknown task) |

## Verifiable success

- `python3 scripts/harness_manifest_builder.py --sample`, `scripts/goal_compiler.py --sample`, and `scripts/loop_controller.py --sample` all exit 0. - A vague goal (`--goal "make it better"`) exits 3 and prints forcing questions. - `loop_controller.py close` on a state with an unverified task exits 4. - The demo loop in `loop_controller.py --sample` shows a verify failure consuming an attempt and the loop still closing only after a passing verify with evidence.

## Related skills

- **workflow-builder**: authoring deterministic `.js` scripts for Claude Code's Workflow tool. NOT for goal-to-close loop state (this skill). - **agenthub**: N parallel agents competing on ONE task in git worktrees. Use it *inside* a harness task that wants competing attempts. - **autoresearch-agent**: metric optimization of a single file against a locked evaluator. Use it when a task's done_when is "metric improves". - **tc-tracker**: per-code-change lifecycle records. Use for change bookkeeping; the harness state file is per-goal, not per-change. - **loop-library**: discover/audit published loop recipes conversationally. This skill is the executable enforcement of that vocabulary. - **ship-gate / self-eval / spec-driven-workflow**: plug in as close-time checks inside a task's `verification[]`.

See [references/domain_harness_design.md](references/domain_harness_design.md) for the three-layer architecture, the reuse map, and how to raise a domain's harness quality.

技术详情

版本
1.0.0
许可证
MIT
最近更新
2026年8月22日
发布时间
2026年8月22日

决策摘要

首选

100
就绪
采用
阶段

24,795 个 GitHub Stars

审计

安装审查

安装与采用审查

87
需审查
安全性
78/100
维护状态
100/100
安装
92/100
打开完整审计查看评估报告

Agent 验证证据

Agent 验证证据

来自解析、审查、安装和一次小范围运行后的结果报告。

0
已验证
Needs first agent run自动安装: 先审查最近: 未知
成功率
近期失败
结果
0
输出质量
失败
0
不相关
0
安装次数
0
风险拦截
0
需要配置
0
生产环境
0

暂时没有 Agent 结果数据。首次 Agent 执行可以通过 /api/agent/outcome 报告成功、需要设置、风险拦截、失败或不相关。

安装

加入 Agent 工作流

免费且开源. 在生产 Agent 中安装前请先审查报告。

增长闭环

分享工具包

X

为 agent-harness 准备的场景化草稿,可手动发布到 X。

策展说明
agent-harness: Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiabl...

24.8K stars

https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x
打开 X 草稿
可选:带安装命令的回复
Listing + install path for agent-harness:
https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x

Install: npx skills add alirezarezvani/claude-skills --skill agent-harness
打开回复草稿

收录来源

Registry 收录

可认领

此列表来自公开来源,维护者认领获批前不会标记为官方。

收录方
OpenAgentSkill 社区索引

归属链接指向公开仓库或创作者主页。创作者可认领列表以更新所有权信号。

认领此 Skill

所有者认领

认领此 Skill 页面

这条 Registry 收录 列表归属于 alirezarezvani,但尚未标记为官方。认领后可增加已验证所有者信号,使后续发布、安装和审计更新更值得信赖。

创作者外链工具包

将证据徽章加入你的 README

在开发者评估仓库的位置展示规范页面、当前信任与审计信号,以及真实的 Agent 验证证据。

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=listed&label=Listed)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=trust&label=Trust)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=audit&label=Audit)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)

作者

A

alirezarezvani

@alirezarezvani

健康信号

GitHub Stars
24.8K
质量评分
54/100
最近 GitHub 推送
2026年8月22日
框架提示
未知
OpenAgentSkill 浏览量
0
复制安装命令
0
跳转点击
0

社区信号

告诉我们这个 Skill 是否对你的 Agent 工作流有帮助。汇总反馈会持续改善排序。

信任与安全

仅限沙盒

73
  • GitHub 采用度25K 个 GitHub Stars通过
  • Star/Fork 活跃度25K 个 Star,3.5K 个 Fork; 当前元数据中没有议题活跃度信息通过
  • 近期维护今天有推送通过
  • 许可证清晰度MIT通过
  • README/SKILL.md 完整度元数据包含足够的用法与工作流上下文通过
  • 依赖与运行时风险命令执行范围信息