agent-harness
Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until eve
공급 자산 프로필
리서치 및 지식 작업
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
시나리오
리서치 Agent
I need my agent to research a topic, compare sources, and produce a concise report.
Agent 적합도
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI 또는 맞춤형 Agent에 적합합니다.
설치
준비됨
npx skills add alirezarezvani/claude-skills --skill agent-harness
유지보수
최신
오늘 푸시됨
위험
검토 필요
Financial research output is not financial advice; require human review before any live investment decision
GitHub 품질
25K
91/100 품질 · 81/100 신뢰
커버리지 태그
검토 메모
Financial research output is not financial advice; require human review before any live investment decision · The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
Agent 채택 스코어카드
신뢰, 감사, 설치 준비 상태를 한눈에 확인하세요
이 점수는 공개 저장소 메타데이터, OpenAgentSkill 검토 신호, 유지보수 최신성, 설치 준비 상태를 결합합니다. 후보 선정 신호일 뿐, 사람의 검토를 대체하지 않습니다.
품질
우수강한 채택 및 유지보수 신호를 갖춘 신뢰도 높은 추천입니다.
신뢰
샌드박스 전용신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.
감사
검토 필요설치 준비 상태, 보안 메타데이터, 유지보수 및 채택 위험에 대한 기계 판독형 검토입니다.
OpenAgentSkill 신뢰 점수 v5
설치 전 사람 검토
실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.
스타
GitHub 스타 25K
저장소 활동
스타 25K, 포크 3.5K
유지보수
오늘 푸시됨
라이선스
MIT
설치
npx skills add alirezarezvani/claude-skills --skill agent-harness
설치 안전성
표준 패키지 또는 런타임 설치 경로
권한 범위
shell or command execution, filesystem or document access
Agent 결과
아직 Agent 결과 데이터가 없습니다
문서
README/SKILL.md 맥락이 충분합니다
위험 요약
프로덕션 전 검토
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- Financial research output is not financial advice; require human review before any live investment decision.
- Quality score needs review
설치 준비 상태
설치 경로 사용 가능
- 설치 경로를 사용할 수 있습니다
- 저장소 근거를 사용할 수 있습니다
- 라이선스가 명시되었습니다
- 아직 Agent 검증 결과 근거가 없습니다
Agent 읽기용 메타데이터
이 스킬의 기계 판독형 의사결정 데이터.
이 블록 또는 포함된 JSON을 사용해 Agent가 이 스킬을 설치할지, 대안을 고를지, 먼저 사람의 검토를 요청할지 판단할 수 있습니다.
적합한 작업
- 리서치 Agent 워크플로
- Claude Code 팀
- GitHub 채택 신호를 중시하는 팀
- 검색 소스
적합한 Agent
설치 결정
- 명령어
- npx skills add alirezarezvani/claude-skills --skill agent-harness
- 정책
- 검토
- 사람 검토
- 예
신뢰와 위험
- 신뢰
- 73/100
- 감사
- 87/100
- 위험 수준
- 검토 필요
결과 루프
- 엔드포인트
- /api/agent/outcome
- 이벤트 ID
- resolve
- 결과
- 5
설치 명령어
npx skills add alirezarezvani/claude-skills --skill agent-harness사용하지 말아야 할 경우
- 벤더 지원 SLA가 필요한 팀
- production agents without a repository review
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다
- 고위험 권한 힌트: Shell 또는 명령 실행
Agent 안전 v2
59/100 · 설치 전 검토
사용 가능한 후보이지만 Agent는 설치 전에 권한과 감사 메모를 표시해야 합니다.
실제 작업 공간에 설치하기 전에 사람의 승인이 필요합니다.
높음
Shell 또는 명령 실행
Skill 메타데이터가 터미널, CLI, Shell, 하위 프로세스 또는 명령 실행 워크플로를 참조합니다.
중간
네트워크 접근
Skill은 원격 페이지, API, 저장소 또는 외부 서비스에 접근할 수 있습니다.
중간
파일 시스템 접근
Skill은 프로젝트 파일, 문서, 생성 산출물 또는 로컬 작업 공간 상태를 읽거나 쓸 수 있습니다.
- 고위험 권한 힌트: Shell 또는 명령 실행
- Financial research output is not financial advice; require human review before any live investment decision
설치 대상
Agent 워크플로에 이 스킬 설치
공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install alirezarezvani-agent-harnessAgent 해결 계획
설치 전에 Agent가 적합성을 검증하게 하세요.
Resolve API는 최우선 스킬, 대안, 안전 정책, 감사 메모, 설치 대상 및 Agent가 페이지를 스크래핑하지 않고 사용할 수 있는 프롬프트를 반환합니다.
JSON 열기
/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve 텍스트
/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
설치 핸드오프
/api/skills/alirezarezvani-agent-harness/install
Agent가 확인할 항목
- Resolve API에서 작업 적합도와 대안을 확인합니다.
- 감사 점수, 신뢰 점수 및 안전 정책 경고를 확인합니다.
- Codex, Claude Code, Cursor 또는 CLI의 설치 대상 호환성을 확인합니다.
프롬프트 복사
Task: Use agent-harness in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install
Install command: npx skills add alirezarezvani/claude-skills --skill agent-harness
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent 핸드오프
또 다른 디렉터리 페이지 대신 설치 경로를 Agent에게 제공합니다.
공개 설치 엔드포인트에서 명령어, 안전 체크리스트, 대상 프롬프트와 정규 링크를 가져옵니다.
설치 핸드오프
/api/skills/alirezarezvani-agent-harness/install
LLM 텍스트 형식
/api/skills/alirezarezvani-agent-harness/install?format=text
대안 찾기
/api/skills/search?q=agent-harness&limit=3
Agent 프롬프트
Use agent-harness for this task. Review https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install, then install with: npx skills add alirezarezvani/claude-skills --skill agent-harnessRegistry 메타데이터
자동 스킬 선택을 위한 Agent 읽기용 프로필.
Registry API를 통해 동일한 결정, 신뢰, 감사, 사용 사례, 설치 신호를 제공하므로 Agent가 UI를 스크래핑하지 않고도 순위를 매길 수 있습니다.
Agent 결정 패널
리서치 Agent용 우선 추천
우선 후보로 사용하되, 자신의 Agent 환경에서 README와 설치 경로를 검증하세요.
스택 내 역할
우선 추천
주요 적합도
리서치 Agent
신뢰 라벨
프로덕션 준비 완료
설치 경로
명령어 준비됨
사용 시점
- 리서치 Agent 워크플로
- Claude Code 팀
- GitHub 채택 신호를 중시하는 팀
근거
- GitHub 스타 24,795
- 최근 저장소 활동
- 설치 명령 또는 GitHub 저장소를 사용할 수 있습니다
- 품질 프로필 91/100
먼저 검토
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- 아직 OpenAgentSkill 사용 피드백 데이터가 없습니다
구현 경로
- 1샌드박스 Agent에 설치하고 리서치 Agent 작업을 처음부터 끝까지 한 번 실행하세요.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
신뢰 프로필
샌드박스 전용
신뢰 신호가 부족하거나 혼재된 유용한 후보입니다. 결과 루프가 작업 적합성을 입증할 때까지 격리된 작업 공간에서 사용하세요.
GitHub 채택도
통과GitHub 스타 25K
스타/포크 활동
통과스타 25K, 포크 3.5K; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다
최근 유지보수
통과오늘 푸시됨
라이선스 명확성
통과MIT
긍정 신호
- AI 검토 승인됨
- 설치 경로를 사용할 수 있습니다
- 저장소 근거를 사용할 수 있습니다
- 최근 유지보수된 저장소
- Large GitHub adoption signal
- 설치 명령에서 뚜렷한 고위험 패턴이 발견되지 않았습니다
- 결과 루프는 준비되었지만 첫 실제 Agent 실행이 필요합니다
설치 전 검토
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- Financial research output is not financial advice; require human review before any live investment decision.
- Quality score needs review
- 아직 실제 Agent 결과 보고서가 없습니다
- 무인 설치 전에 사람 검토가 필요합니다
권장 작업
실제 작업에 사용하기 전 샌드박스에서만 실행하고 유사 대안과 비교하세요.
품질 프로필
우수 Agent 워크플로용 후보
강한 채택 및 유지보수 신호를 갖춘 신뢰도 높은 추천입니다.
워크플로 적합도
이 스킬을 사용할 시나리오
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Manage repositories
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
워크플로 적합도
완전한 워크플로에 추가
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
대안 후보
설치 전 비교
이 작업에 적합할 수 있는 유사 스킬입니다.
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
Academic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
GPT Researcher
Run autonomous deep research over web and local sources
DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
개요
--- name: agent-harness description: "Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived. Use when you want an agent or subagent to pick up a goal and drive it to a verified close across one of this repo's 18 domains ('run this goal through the engineering harness', 'set up an agentic loop for marketing work', 'make the finance domain self-verifying'). NOT for authoring Claude Code Workflow-tool .js scripts (workflow-builder), N-agent tournaments on one task (agenthub), single-file metric optimization (autoresearch-agent), or discovering published loop recipes (loop-library)." ---
# Agent Harness
You are a harness operator, not a hero. The loop — not your optimism — decides when work is done. Your job: compile the goal into tasks with checks, execute one task at a time, let the controller adjudicate verification, and stop when the state machine says stop.
## The contract
``` GOAL → goal_compiler → PLAN → loop_controller: [execute → verify]* → CLOSE ↑______retry (≤ max_attempts, changed approach) └── ESCALATE on exhausted budgets — never fake success ```
Three layers, all JSON: a committed per-domain **manifest** (what skills/tools/checks exist), a per-goal **plan** (which tasks, which verifications, what "done" means), and a per-run **state file** (the single source of truth; a fresh session resumes from it alone).
## Quick start
```bash # 0. Pick the domain manifest (18 committed under assets/harnesses/, e.g. engineering-team.json) ls assets/harnesses/
# 1. Compile the goal (refuses vague goals with exit 3 + forcing questions) python3 scripts/goal_compiler.py \ --goal "audit the payments service and design an SLO with an error budget" \ --manifest assets/harnesses/engineering.json --out plan.json
# 2. Initialize the loop state python3 scripts/loop_controller.py init --plan plan.json --state .agent-harness/state.json
# 3. Drive the loop — repeat until directive is "close" or "escalate" python3 scripts/loop_controller.py next --state .agent-harness/state.json # → {"action": "execute", "task": "T1", ...}: open the task's skill (SKILL.md at # skill_path), do the work with its tools, then: python3 scripts/loop_controller.py record --state .agent-harness/state.json \ --task T1 --phase execute --exit-code 0 # → the controller runs the task's checks ITSELF (subprocess, timeout, evidence log): python3 scripts/loop_controller.py verify --state .agent-harness/state.json --task T1 --cwd <repo-root>
# 4. Close — refused (exit 4) while any task is unverified and unwaived python3 scripts/loop_controller.py close --state .agent-harness/state.json ```
Regenerate a manifest after skills change (diff-stable, CI-checkable):
```bash python3 scripts/harness_manifest_builder.py --domain engineering-team \ --repo-root <repo-root> --out-dir assets/harnesses --no-timestamp ```
## Hard rules
1. **Never adjudicate your own verification.** `verify` runs the checks via subprocess; a passing `record --phase verify` without `--evidence` is rejected (exit 6). You do not get to declare a task verified. 2. **Never modify a gate you are judged by.** Check commands come from the manifest/plan. Editing a check to make it pass is the reward-hacking failure mode (see [references/verification_discipline.md](references/verification_discipline.md)) — same invariant as autoresearch-agent's locked evaluator. 3. **One task at a time, writes serialized.** Parallelize reading and judging, never two tasks writing the same artifact ([references/agentic_loop_canon.md](references/agentic_loop_canon.md)). 4. **Retry means a changed approach.** Same command + same input = same failure. The retry directive says so; honor it. 5. **Budgets are terminal states, not suggestions.** `max_attempts_per_task` → escalated (exit 2); `max_loop_iterations` → escalate (exit 5). Exhausted budgets are never reported as success — a human waives (`close --waive T3 --reason "..."`), you don't. 6. **Fresh context beats long context.** Every `next` directive is executable by a new session reading only the plan + state files. Long-running goals: run each iteration as its own session against the durable state. 7. **State lives in `.agent-harness/`** — never in `.agenthub/`, `.autoresearch/`, or `docs/TC/` (those belong to sibling skills). 8. **Plan and state files are a trust boundary.** `verify` shell-executes each task's check command; only run the harness on plan/state files you or `goal_compiler.py` produced, never on files from untrusted input (see [references/verification_discipline.md](references/verification_discipline.md)).
## Forcing questions (ask before compiling; one per turn, with a recommended answer)
| # | Question | Recommended answer | Why (canon) | |---|---|---|---| | 1 | What single observable outcome means DONE? | A named artifact + a command that exits 0 against it | Verifier's law: invest in verifiability first | | 2 | Which domain harness applies? | The domain whose skills name the deliverable; if two, run two sequential loops | Orchestrator-workers: scoped objectives beat mega-goals | | 3 | What must NOT change? | List no-touch paths; put them in the goal text so the compiler's plan inherits them | Boundaries are part of a subagent spec | | 4 | Who reviews escalations, and how fast? | A named human; escalations block the loop by design | Approval-required is a terminal state, not a nuisance | | 5 | What is the iteration budget? | Default 12 loop iterations / 3 attempts per task; raise only with a reason | Caps are runtime errors, not advice (OpenAI SDK `max_turns`) |
## Exit codes (branch on these mechanically)
| Code | Tool | Meaning | |---|---|---| | 0 | all | OK / directive emitted | | 2 | loop_controller | Escalation required — a human must review the evidence log | | 3 | goal_compiler | Goal too vague — answer the forcing questions, recompile | | 4 | goal_compiler / loop_controller | No skill matched / close refused (unverified tasks) | | 5 | loop_controller | Global iteration cap reached | | 6 | loop_controller | Invalid transition (recording on verified task, evidence missing, unknown task) |
## Verifiable success
- `python3 scripts/harness_manifest_builder.py --sample`, `scripts/goal_compiler.py --sample`, and `scripts/loop_controller.py --sample` all exit 0. - A vague goal (`--goal "make it better"`) exits 3 and prints forcing questions. - `loop_controller.py close` on a state with an unverified task exits 4. - The demo loop in `loop_controller.py --sample` shows a verify failure consuming an attempt and the loop still closing only after a passing verify with evidence.
## Related skills
- **workflow-builder**: authoring deterministic `.js` scripts for Claude Code's Workflow tool. NOT for goal-to-close loop state (this skill). - **agenthub**: N parallel agents competing on ONE task in git worktrees. Use it *inside* a harness task that wants competing attempts. - **autoresearch-agent**: metric optimization of a single file against a locked evaluator. Use it when a task's done_when is "metric improves". - **tc-tracker**: per-code-change lifecycle records. Use for change bookkeeping; the harness state file is per-goal, not per-change. - **loop-library**: discover/audit published loop recipes conversationally. This skill is the executable enforcement of that vocabulary. - **ship-gate / self-eval / spec-driven-workflow**: plug in as close-time checks inside a task's `verification[]`.
See [references/domain_harness_design.md](references/domain_harness_design.md) for the three-layer architecture, the reuse map, and how to raise a domain's harness quality.
기술 세부 사항
- 버전
- 1.0.0
- 라이선스
- MIT
- 최근 업데이트
- 2026년 8월 22일
- 게시일
- 2026년 8월 22일
결정 스냅샷
우선 추천
GitHub 스타 24,795
Agent 검증 증거
Agent 검증 증거
Resolve, 검토, 설치 및 한 번의 제한된 실행 후 결과 보고서입니다.
- 성공률
- —
- 최근 실패
- —
- 결과
- 0
- 출력 품질
- —
- 실패
- 0
- 관련 없음
- 0
- 설치
- 0
- 위험 차단
- 0
- 설정 필요
- 0
- 프로덕션
- 0
아직 Agent 결과 데이터가 없습니다. 첫 실행은 /api/agent/outcome을 통해 성공, 설정 필요, 위험 차단, 실패 또는 비관련 결과를 보고할 수 있습니다.
성장 루프
공유 키트
agent-harness용 시나리오 기반 초안입니다. X에 수동으로 게시할 수 있습니다.
agent-harness: Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiabl... 24.8K stars https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x
선택 사항: 설치 명령이 포함된 답글
Listing + install path for agent-harness: https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x Install: npx skills add alirezarezvani/claude-skills --skill agent-harness
등록 출처
Registry 색인
이 등록은 공개 소스에서 색인되었으며 유지보수자 소유권 주장이 승인될 때까지 공식으로 표시되지 않습니다.
- 색인 주체
- OpenAgentSkill 커뮤니티 인덱스
귀속은 공개 저장소 또는 제작자 프로필에 연결됩니다. 제작자는 등록을 주장하여 소유권 신호를 업데이트할 수 있습니다.
이 스킬 소유권 주장소유자 소유권 주장
이 스킬 등록 소유권 주장
이 Registry 색인 등록은 alirezarezvani에게 귀속되어 있지만 아직 공식으로 표시되지 않았습니다. 소유권을 주장하면 확인된 소유자 신호가 추가되어 이후 출시, 설치 및 감사 업데이트를 더 신뢰할 수 있습니다.
크리에이터 백링크 키트
README에 증거 배지 추가
개발자가 저장소를 평가하는 위치에 정규 등록, 현재 신뢰 및 감사 신호, 실제 Agent-Proven 증거를 표시합니다.
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness/audit)
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)작성자
alirezarezvani
@alirezarezvani
플랫폼 적합도
상태 신호
- GitHub 스타
- 24.8K
- 품질 점수
- 54/100
- 최근 GitHub 푸시
- 2026년 8월 22일
- 프레임워크 힌트
- 알 수 없음
- OpenAgentSkill 조회수
- 0
- 설치 명령 복사
- 0
- 외부 클릭
- 0
커뮤니티 신호
이 스킬이 Agent 워크플로에 유용한지 알려 주세요. 집계된 피드백은 시간이 지날수록 순위를 개선합니다.
신뢰와 안전
샌드박스 전용
- GitHub 채택도GitHub 스타 25K통과
- 스타/포크 활동스타 25K, 포크 3.5K; 현재 메타데이터에서 이슈 활동을 확인할 수 없습니다통과
- 최근 유지보수오늘 푸시됨통과
- 라이선스 명확성MIT통과
- README/SKILL.md 완성도메타데이터에 충분한 사용 및 워크플로 맥락이 포함되어 있습니다통과
- 의존성/런타임 위험명령 실행 범위정보
관련 스킬
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
53.5K 스타Academic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
38.4K 스타GPT Researcher
Run autonomous deep research over web and local sources
28.0K 스타DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
19.8K 스타