agent-harness

レビュー · 73
Registry に収録

Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until eve

Verified installs0
スター24.8K
バージョン1.0.0
品質91/100 · 優秀
信頼73/100 · サンドボックス限定
監査87/100 · 要レビュー

供給アセットの概要

リサーチとナレッジ作業

Deep research, source comparison, literature review, RAG, knowledge search, and reports.

カテゴリを見る

シナリオ

リサーチ Agent

I need my agent to research a topic, compare sources, and produce a concise report.

Agent 適合

Claude Code + OpenAI Agents + CLI

Codex、Claude Code、Cursor、CLI、またはカスタム Agent に対応します。

インストール

準備完了

npx skills add alirezarezvani/claude-skills --skill agent-harness

メンテナンス

新しい

本日プッシュ

リスク

要レビュー

Financial research output is not financial advice; require human review before any live investment decision

GitHub 品質

25K

91/100 品質 · 81/100 信頼

対象タグ

リサーチリサーチ Agentagent-skill

レビュー注記

Financial research output is not financial advice; require human review before any live investment decision · The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.

Agent 導入スコアカード

信頼、監査、インストール準備状況を一目で確認

公開リポジトリのメタデータ、OpenAgentSkill のレビューシグナル、保守の鮮度、インストール準備状況を組み合わせたスコアです。候補選定の目安であり、人によるレビューの代替ではありません。

品質

優秀
91

採用度と保守性のシグナルが強い高信頼候補です。

信頼

サンドボックス限定
73

信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。

監査

要レビュー
87

インストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。

OpenAgentSkill Trust Score v5

インストール前に人のレビュー

実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。

CodexClaude CodeCursorOpenAgentSkill CLI

スター

GitHub スター 25K

リポジトリ活動

スター 25K、フォーク 3.5K

メンテナンス

本日プッシュ

ライセンス

MIT

インストール

npx skills add alirezarezvani/claude-skills --skill agent-harness

インストール安全性

標準パッケージまたはランタイムのインストールパス

権限範囲

shell or command execution, filesystem or document access

Agent の成果

Agent の成果データはまだありません

ドキュメント

README/SKILL.md の文脈が十分です

リスク概要

本番前にレビュー

  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • Financial research output is not financial advice; require human review before any live investment decision.
  • Quality score needs review

インストール準備状況

インストールパスを利用可能

  • インストールパスを利用できます
  • リポジトリの根拠を利用できます
  • ライセンスが明示されています
  • Agent-Proven の成果エビデンスはまだありません

Agent 可読メタデータ

このスキルの機械可読な判断データ。

このブロックまたは埋め込み JSON を使い、Agent がこのスキルをインストールすべきか、代替を選ぶべきか、先に人のレビューを求めるべきかを判断できます。

JSON を開く

適したタスク

  • リサーチ Agent ワークフロー
  • Claude Code チーム
  • GitHub 採用シグナルを重視するチーム
  • 検索ソース

適した Agent

CodexClaude CodeCursorOpenAgentSkill CLIOpenAI AgentsCLI

インストール判断

コマンド
npx skills add alirezarezvani/claude-skills --skill agent-harness
ポリシー
レビュー
人によるレビュー
はい

信頼とリスク

信頼
73/100
監査
87/100
リスクレベル
要レビュー

成果ループ

エンドポイント
/api/agent/outcome
イベント ID
resolve
成果
5

インストールコマンド

npx skills add alirezarezvani/claude-skills --skill agent-harness

使わない場合

  • ベンダー提供の SLA が必要なチーム
  • production agents without a repository review
  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • OpenAgentSkill の利用フィードバックはまだありません
  • 高リスク権限のヒント: Shell またはコマンド実行
近い代替スキルはまだ登録されていません。

Agent セーフティ v2

59/100 · インストール前にレビュー

権限メモ付きでレビュー済みレビュー

利用可能な候補ですが、Agent はインストール前に権限と監査メモを提示する必要があります。

実際のワークスペースへインストールする前に人の承認が必要です。

API で解決

Shell またはコマンド実行

Skill メタデータに端末、CLI、Shell、サブプロセス、またはコマンド実行のワークフローが含まれます。

ネットワークアクセス

Skill はリモートページ、API、リポジトリ、外部サービスにアクセスする可能性があります。

ファイルシステムアクセス

Skill はプロジェクトファイル、ドキュメント、生成物、ローカルワークスペース状態を読み書きする可能性があります。

  • 高リスク権限のヒント: Shell またはコマンド実行
  • Financial research output is not financial advice; require human review before any live investment decision

インストール先

Agent ワークフローにこのスキルをインストール

公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。

skill install

OpenAgentSkill CLI

Resolve policy, run the source installer safely, and report a verified install receipt.

$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install alirezarezvani-agent-harness

Agent 解決プラン

インストール前に Agent に適合性を検証させます。

Resolve API は第一候補、代替、安全ポリシー、監査メモ、インストール先、Agent がそのまま使えるプロンプトを返します。

テキストプランを開く

Agent が確認すべきこと

  • Resolve API でタスク適合と代替を確認。
  • 監査・信頼スコアと安全ポリシーの警告を確認。
  • Codex、Claude Code、Cursor、CLI のインストール先互換性を確認。

プロンプトをコピー

Task: Use agent-harness in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install
Install command: npx skills add alirezarezvani/claude-skills --skill agent-harness
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.

Agent 引き継ぎ

別のディレクトリではなく、インストール経路を Agent に渡します。

公開インストールエンドポイントからコマンド、安全チェックリスト、対象プロンプト、正規リンクを取得します。

Install API を開く

Agent プロンプト

Use agent-harness for this task. Review https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install, then install with: npx skills add alirezarezvani/claude-skills --skill agent-harness

Registry メタデータ

自動スキル選択用の Agent 可読プロファイル。

Registry API 経由で判断、信頼、監査、ユースケース、インストールのシグナルを提供し、UI をスクレイピングせずに Agent が順位付けできます。

Manifest を開く

Agent 適合

100/100

リサーチ Agent

プラットフォーム

Claude Code, OpenAI Agents

監査レポート

要レビュー · 87/100

インストール準備、安全メタデータ、保守、採用リスクの機械可読なレビュー。

監査レポートを見る評価レポートを見る

Agent 判断パネル

リサーチ Agent 向けの第一候補

有力候補として扱い、自分の Agent スタックで README とインストール経路を検証してください。

100
準備状況
採用
段階

スタック内の役割

第一候補

主な適合

リサーチ Agent

信頼ラベル

本番対応

インストールパス

コマンド準備済み

使う場面

  • リサーチ Agent ワークフロー
  • Claude Code チーム
  • GitHub 採用シグナルを重視するチーム

根拠

  • GitHub スター 24,795
  • 最近のリポジトリ活動
  • インストールコマンドまたは GitHub リポジトリが利用可能
  • 品質プロファイル 91/100

先にレビュー

  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • OpenAgentSkill の利用フィードバックはまだありません

実装パス

  1. 1サンドボックスの Agent にインストールし、リサーチ Agent タスクを一度最初から最後まで実行します。
  2. 2Compare output quality, latency, and failure behavior against at least one alternative.
  3. 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.

信頼プロファイル

サンドボックス限定

信頼シグナルが不足または混在する有用な候補です。結果ループがタスク適合を示すまで、隔離されたワークスペースで使用してください。

73
OpenAgentSkill Trust Score

GitHub 採用度

合格

GitHub スター 25K

スター/フォーク活動

合格

スター 25K、フォーク 3.5K; 現在のメタデータでは Issue 活動を利用できません

最近のメンテナンス

合格

本日プッシュ

ライセンスの明確さ

合格

MIT

良いシグナル

  • AI レビュー承認済み
  • インストールパスを利用できます
  • リポジトリの根拠を利用できます
  • 最近保守されたリポジトリ
  • Large GitHub adoption signal
  • インストールコマンドに明確な高リスクパターンはありません
  • 成果ループは準備済みですが、最初の実行が必要です

インストール前にレビュー

  • The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
  • Financial research output is not financial advice; require human review before any live investment decision.
  • Quality score needs review
  • 実際の Agent 成果レポートはまだありません
  • 無人インストールの前に人によるレビューが必要です

推奨アクション

実作業で使う前に、サンドボックスでのみ実行し、近い代替と比較してください。

品質プロファイル

優秀 Agent ワークフロー向けの候補

採用度と保守性のシグナルが強い高信頼候補です。

91
GitHub スター
25K
鮮度
今日
インストール準備完了
はい
ライセンス
MIT
インストール前にレビュー: The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.

ワークフロー適合

このスキルを使うシナリオ

ワークフロー適合

完全なワークフローに追加

概要

--- name: agent-harness description: "Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived. Use when you want an agent or subagent to pick up a goal and drive it to a verified close across one of this repo's 18 domains ('run this goal through the engineering harness', 'set up an agentic loop for marketing work', 'make the finance domain self-verifying'). NOT for authoring Claude Code Workflow-tool .js scripts (workflow-builder), N-agent tournaments on one task (agenthub), single-file metric optimization (autoresearch-agent), or discovering published loop recipes (loop-library)." ---

# Agent Harness

You are a harness operator, not a hero. The loop — not your optimism — decides when work is done. Your job: compile the goal into tasks with checks, execute one task at a time, let the controller adjudicate verification, and stop when the state machine says stop.

## The contract

``` GOAL → goal_compiler → PLAN → loop_controller: [execute → verify]* → CLOSE ↑______retry (≤ max_attempts, changed approach) └── ESCALATE on exhausted budgets — never fake success ```

Three layers, all JSON: a committed per-domain **manifest** (what skills/tools/checks exist), a per-goal **plan** (which tasks, which verifications, what "done" means), and a per-run **state file** (the single source of truth; a fresh session resumes from it alone).

## Quick start

```bash # 0. Pick the domain manifest (18 committed under assets/harnesses/, e.g. engineering-team.json) ls assets/harnesses/

# 1. Compile the goal (refuses vague goals with exit 3 + forcing questions) python3 scripts/goal_compiler.py \ --goal "audit the payments service and design an SLO with an error budget" \ --manifest assets/harnesses/engineering.json --out plan.json

# 2. Initialize the loop state python3 scripts/loop_controller.py init --plan plan.json --state .agent-harness/state.json

# 3. Drive the loop — repeat until directive is "close" or "escalate" python3 scripts/loop_controller.py next --state .agent-harness/state.json # → {"action": "execute", "task": "T1", ...}: open the task's skill (SKILL.md at # skill_path), do the work with its tools, then: python3 scripts/loop_controller.py record --state .agent-harness/state.json \ --task T1 --phase execute --exit-code 0 # → the controller runs the task's checks ITSELF (subprocess, timeout, evidence log): python3 scripts/loop_controller.py verify --state .agent-harness/state.json --task T1 --cwd <repo-root>

# 4. Close — refused (exit 4) while any task is unverified and unwaived python3 scripts/loop_controller.py close --state .agent-harness/state.json ```

Regenerate a manifest after skills change (diff-stable, CI-checkable):

```bash python3 scripts/harness_manifest_builder.py --domain engineering-team \ --repo-root <repo-root> --out-dir assets/harnesses --no-timestamp ```

## Hard rules

1. **Never adjudicate your own verification.** `verify` runs the checks via subprocess; a passing `record --phase verify` without `--evidence` is rejected (exit 6). You do not get to declare a task verified. 2. **Never modify a gate you are judged by.** Check commands come from the manifest/plan. Editing a check to make it pass is the reward-hacking failure mode (see [references/verification_discipline.md](references/verification_discipline.md)) — same invariant as autoresearch-agent's locked evaluator. 3. **One task at a time, writes serialized.** Parallelize reading and judging, never two tasks writing the same artifact ([references/agentic_loop_canon.md](references/agentic_loop_canon.md)). 4. **Retry means a changed approach.** Same command + same input = same failure. The retry directive says so; honor it. 5. **Budgets are terminal states, not suggestions.** `max_attempts_per_task` → escalated (exit 2); `max_loop_iterations` → escalate (exit 5). Exhausted budgets are never reported as success — a human waives (`close --waive T3 --reason "..."`), you don't. 6. **Fresh context beats long context.** Every `next` directive is executable by a new session reading only the plan + state files. Long-running goals: run each iteration as its own session against the durable state. 7. **State lives in `.agent-harness/`** — never in `.agenthub/`, `.autoresearch/`, or `docs/TC/` (those belong to sibling skills). 8. **Plan and state files are a trust boundary.** `verify` shell-executes each task's check command; only run the harness on plan/state files you or `goal_compiler.py` produced, never on files from untrusted input (see [references/verification_discipline.md](references/verification_discipline.md)).

## Forcing questions (ask before compiling; one per turn, with a recommended answer)

| # | Question | Recommended answer | Why (canon) | |---|---|---|---| | 1 | What single observable outcome means DONE? | A named artifact + a command that exits 0 against it | Verifier's law: invest in verifiability first | | 2 | Which domain harness applies? | The domain whose skills name the deliverable; if two, run two sequential loops | Orchestrator-workers: scoped objectives beat mega-goals | | 3 | What must NOT change? | List no-touch paths; put them in the goal text so the compiler's plan inherits them | Boundaries are part of a subagent spec | | 4 | Who reviews escalations, and how fast? | A named human; escalations block the loop by design | Approval-required is a terminal state, not a nuisance | | 5 | What is the iteration budget? | Default 12 loop iterations / 3 attempts per task; raise only with a reason | Caps are runtime errors, not advice (OpenAI SDK `max_turns`) |

## Exit codes (branch on these mechanically)

| Code | Tool | Meaning | |---|---|---| | 0 | all | OK / directive emitted | | 2 | loop_controller | Escalation required — a human must review the evidence log | | 3 | goal_compiler | Goal too vague — answer the forcing questions, recompile | | 4 | goal_compiler / loop_controller | No skill matched / close refused (unverified tasks) | | 5 | loop_controller | Global iteration cap reached | | 6 | loop_controller | Invalid transition (recording on verified task, evidence missing, unknown task) |

## Verifiable success

- `python3 scripts/harness_manifest_builder.py --sample`, `scripts/goal_compiler.py --sample`, and `scripts/loop_controller.py --sample` all exit 0. - A vague goal (`--goal "make it better"`) exits 3 and prints forcing questions. - `loop_controller.py close` on a state with an unverified task exits 4. - The demo loop in `loop_controller.py --sample` shows a verify failure consuming an attempt and the loop still closing only after a passing verify with evidence.

## Related skills

- **workflow-builder**: authoring deterministic `.js` scripts for Claude Code's Workflow tool. NOT for goal-to-close loop state (this skill). - **agenthub**: N parallel agents competing on ONE task in git worktrees. Use it *inside* a harness task that wants competing attempts. - **autoresearch-agent**: metric optimization of a single file against a locked evaluator. Use it when a task's done_when is "metric improves". - **tc-tracker**: per-code-change lifecycle records. Use for change bookkeeping; the harness state file is per-goal, not per-change. - **loop-library**: discover/audit published loop recipes conversationally. This skill is the executable enforcement of that vocabulary. - **ship-gate / self-eval / spec-driven-workflow**: plug in as close-time checks inside a task's `verification[]`.

See [references/domain_harness_design.md](references/domain_harness_design.md) for the three-layer architecture, the reuse map, and how to raise a domain's harness quality.

技術詳細

バージョン
1.0.0
ライセンス
MIT
最終更新
2026年8月22日
公開日
2026年8月22日

判断の要約

第一候補

100
準備完了
採用
段階

GitHub スター 24,795

監査

インストールレビュー

インストールと採用のレビュー

87
要レビュー
セキュリティ
78/100
メンテナンス
100/100
インストール
92/100
完全な監査を開く評価レポートを見る

Agent 実証エビデンス

Agent 実証エビデンス

Resolve、レビュー、インストール、限定実行後の成果レポート。

0
実証済み
Needs first agent run自動インストール: 先にレビュー最新: 不明
成功率
直近の失敗
成果
0
出力品質
失敗
0
非該当
0
インストール数
0
リスクによりブロック
0
設定が必要
0
本番
0

Agent の実行結果はまだありません。最初の実行では /api/agent/outcome を通じて成功、設定要件、リスクによるブロック、失敗、非該当を報告できます。

インストール

Agent ワークフローに追加

無料・オープンソース. 本番 Agent にインストールする前にレポートを確認してください。

成長ループ

共有キット

X

agent-harness 用のシナリオベース草案です。X へ手動投稿できます。

キュレーターノート
agent-harness: Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiabl...

24.8K stars

https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x
X 下書きを開く
任意:インストールコマンド付きの返信
Listing + install path for agent-harness:
https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x

Install: npx skills add alirezarezvani/claude-skills --skill agent-harness
返信の下書きを開く

掲載元

Registry により登録

申請可能

この掲載は公開ソースから登録されており、メンテナー申請が承認されるまで公式として表示されません。

インデックス作成者
OpenAgentSkill コミュニティインデックス

帰属は公開リポジトリまたは作成者プロフィールにリンクされています。作成者は掲載を申請して所有権シグナルを更新できます。

このスキルを申請

所有者の申請

このスキル掲載を申請

この Registry により登録 掲載は alirezarezvani に帰属していますが、まだ公式として表示されていません。申請すると、確認済み所有者シグナルが追加され、今後の公開、インストール、監査更新の信頼性が高まります。

クリエイター被リンクキット

README にエビデンスバッジを追加

開発者がリポジトリを評価する場所で、正規掲載、現在の信頼・監査シグナル、実際の Agent-Proven エビデンスを表示します。

[![Listed on OpenAgentSkill](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=listed&label=Listed)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[![OpenAgentSkill Trust](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=trust&label=Trust)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[![OpenAgentSkill Audit](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=audit&label=Audit)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness/audit)
[![Agent Proven](https://www.openagentskill.com/api/badge/alirezarezvani-agent-harness?metric=proven&label=Agent%20Proven)](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)

作者

A

alirezarezvani

@alirezarezvani

プラットフォーム適合

健全性シグナル

GitHub スター
24.8K
品質スコア
54/100
最終 GitHub プッシュ
2026年8月22日
フレームワークのヒント
不明
OpenAgentSkill 閲覧数
0
インストールコピー数
0
外部クリック
0

コミュニティシグナル

このスキルが Agent ワークフローに役立つかを共有してください。集約されたフィードバックがランキングを改善します。

信頼と安全性

サンドボックス限定

73
  • GitHub 採用度GitHub スター 25K合格
  • スター/フォーク活動スター 25K、フォーク 3.5K; 現在のメタデータでは Issue 活動を利用できません合格
  • 最近のメンテナンス本日プッシュ合格
  • ライセンスの明確さMIT合格
  • README/SKILL.md の完全性メタデータには十分な利用・ワークフロー文脈があります合格
  • 依存関係/ランタイムのリスクコマンド実行範囲情報