agent-harness
Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until eve
Supply asset profile
Research and knowledge work
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add alirezarezvani/claude-skills --skill agent-harness
Maintenance
fresh
Pushed today
Risk
Needs review
Financial research output is not financial advice; require human review before any live investment decision
GitHub quality
25K
91/100 Quality · 81/100 Trust
Coverage tags
Review notes
Financial research output is not financial advice; require human review before any live investment decision · The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
Agent adoption scorecard
Trust, audit, and install readiness at a glance
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Human review before install
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
25K GitHub stars
Repo activity
25K stars, 3.5K forks
Maintenance
Pushed today
License
MIT
Install
npx skills add alirezarezvani/claude-skills --skill agent-harness
Install safety
standard package or runtime install path
Permission surface
shell or command execution, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Review before production
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- Financial research output is not financial advice; require human review before any live investment decision.
- Quality score needs review
Install readiness
Install path available
- Install path is available
- Repository evidence is available
- License is declared
- No Agent Proven outcome evidence yet
Agent-readable metadata
Machine-readable decision data for this skill.
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
- Research agents workflows
- Claude Code teams
- teams that value GitHub adoption signals
- Search sources
Suited agents
Install decision
- Command
- npx skills add alirezarezvani/claude-skills --skill agent-harness
- Policy
- review
- Human review
- yes
Trust and risk
- Trust
- 73/100
- Audit
- 87/100
- Risk level
- Needs review
Outcome loop
- Endpoint
- /api/agent/outcome
- Event ID
- resolve
- Outcomes
- 5
Install command
npx skills add alirezarezvani/claude-skills --skill agent-harnessDo not use when
- teams that need a vendor-supported SLA
- production agents without a repository review
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- No OpenAgentSkill engagement data yet
- High-risk permission hints: Shell or command execution
Alternative
Last30days Skill
53.5K Stars
npx skills add mvanhorn/last30days-skill -g
Alternative
Academic Research Skills
38.4K Stars
npx skills add Imbad0202/academic-research-skills
Alternative
GPT Researcher
28.0K Stars
npx skills add assafelovic/gpt-researcher
Alternative
DeepResearch
19.8K Stars
npx skills add Alibaba-NLP/DeepResearch
Agent safety v2
59/100 · Review before install
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
high
Shell or command execution
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Network access
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Filesystem access
Skill may read or write project files, documents, generated artifacts, or local workspace state.
- High-risk permission hints: Shell or command execution
- Financial research output is not financial advice; require human review before any live investment decision
Install targets
Install this skill in your agent workflow
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install alirezarezvani-agent-harnessAgent resolve plan
Let an agent verify fit before installing.
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/alirezarezvani-agent-harness/install
Agent should check
- Task fit and alternatives from Resolve API.
- Audit score, trust score, and safety policy warnings.
- Install target compatibility for Codex, Claude Code, Cursor, or CLI.
Copy prompt
Task: Use agent-harness in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20agent-harness%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install
Install command: npx skills add alirezarezvani/claude-skills --skill agent-harness
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Give an agent the install path, not another directory page.
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/alirezarezvani-agent-harness/install
LLM text format
/api/skills/alirezarezvani-agent-harness/install?format=text
Find alternatives
/api/skills/search?q=agent-harness&limit=3
Agent prompt
Use agent-harness for this task. Review https://www.openagentskill.com/api/skills/alirezarezvani-agent-harness/install, then install with: npx skills add alirezarezvani/claude-skills --skill agent-harnessRegistry metadata
Agent-readable profile for automatic skill selection.
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/alirezarezvani-agent-harness
LLM text
/api/registry/manifest/alirezarezvani-agent-harness?format=text
Install alias
/api/registry/install/alirezarezvani-agent-harness
Recommend
/api/registry/recommend?task=Use%20agent-harness%20in%20an%20agent%20workflow&limit=3
Agent fit
Research agents
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
Needs review · 87/100
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Primary pick for Research agents
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Research agents
Trust label
Production-ready
Install path
Command ready
Use when
- Research agents workflows
- Claude Code teams
- teams that value GitHub adoption signals
Evidence
- 24,795 GitHub stars
- recent repository activity
- install command or GitHub repo available
- 91/100 quality profile
review first
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- No OpenAgentSkill engagement data yet
Implementation path
- 1Install it in a sandbox agent and run one Research agents task end to end.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
Trust profile
Sandbox only
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS25K GitHub stars
Stars/forks activity
PASS25K stars, 3.5K forks; issue activity unavailable in current metadata
Recent maintenance
PASSPushed today
License clarity
PASSMIT
Good signals
- AI review approved
- Install path is available
- Repository evidence is available
- Recently maintained repository
- Large GitHub adoption signal
- Install command has no obvious high-risk pattern
- Outcome loop is ready but needs first real agent run
Review before install
- The skill relies on external scripts (goal_compiler.py, loop_controller.py, etc.) not fully reviewed in this excerpt; their security posture should be verified independently.
- Financial research output is not financial advice; require human review before any live investment decision.
- Quality score needs review
- No real agent outcome reports yet
- Human review required before unattended installation
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Excellent candidate for agent workflows
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Use this skill in these scenarios
Investigate faster
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
Manage repositories
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Workflow fit
Add it to a complete workflow
Find, compare, and synthesize
Research report agent
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Alternative shortlist
Compare before you install
Similar skills that may fit this task.
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
Academic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
GPT Researcher
Run autonomous deep research over web and local sources
DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Overview
--- name: agent-harness description: "Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived. Use when you want an agent or subagent to pick up a goal and drive it to a verified close across one of this repo's 18 domains ('run this goal through the engineering harness', 'set up an agentic loop for marketing work', 'make the finance domain self-verifying'). NOT for authoring Claude Code Workflow-tool .js scripts (workflow-builder), N-agent tournaments on one task (agenthub), single-file metric optimization (autoresearch-agent), or discovering published loop recipes (loop-library)." ---
# Agent Harness
You are a harness operator, not a hero. The loop — not your optimism — decides when work is done. Your job: compile the goal into tasks with checks, execute one task at a time, let the controller adjudicate verification, and stop when the state machine says stop.
## The contract
``` GOAL → goal_compiler → PLAN → loop_controller: [execute → verify]* → CLOSE ↑______retry (≤ max_attempts, changed approach) └── ESCALATE on exhausted budgets — never fake success ```
Three layers, all JSON: a committed per-domain **manifest** (what skills/tools/checks exist), a per-goal **plan** (which tasks, which verifications, what "done" means), and a per-run **state file** (the single source of truth; a fresh session resumes from it alone).
## Quick start
```bash # 0. Pick the domain manifest (18 committed under assets/harnesses/, e.g. engineering-team.json) ls assets/harnesses/
# 1. Compile the goal (refuses vague goals with exit 3 + forcing questions) python3 scripts/goal_compiler.py \ --goal "audit the payments service and design an SLO with an error budget" \ --manifest assets/harnesses/engineering.json --out plan.json
# 2. Initialize the loop state python3 scripts/loop_controller.py init --plan plan.json --state .agent-harness/state.json
# 3. Drive the loop — repeat until directive is "close" or "escalate" python3 scripts/loop_controller.py next --state .agent-harness/state.json # → {"action": "execute", "task": "T1", ...}: open the task's skill (SKILL.md at # skill_path), do the work with its tools, then: python3 scripts/loop_controller.py record --state .agent-harness/state.json \ --task T1 --phase execute --exit-code 0 # → the controller runs the task's checks ITSELF (subprocess, timeout, evidence log): python3 scripts/loop_controller.py verify --state .agent-harness/state.json --task T1 --cwd <repo-root>
# 4. Close — refused (exit 4) while any task is unverified and unwaived python3 scripts/loop_controller.py close --state .agent-harness/state.json ```
Regenerate a manifest after skills change (diff-stable, CI-checkable):
```bash python3 scripts/harness_manifest_builder.py --domain engineering-team \ --repo-root <repo-root> --out-dir assets/harnesses --no-timestamp ```
## Hard rules
1. **Never adjudicate your own verification.** `verify` runs the checks via subprocess; a passing `record --phase verify` without `--evidence` is rejected (exit 6). You do not get to declare a task verified. 2. **Never modify a gate you are judged by.** Check commands come from the manifest/plan. Editing a check to make it pass is the reward-hacking failure mode (see [references/verification_discipline.md](references/verification_discipline.md)) — same invariant as autoresearch-agent's locked evaluator. 3. **One task at a time, writes serialized.** Parallelize reading and judging, never two tasks writing the same artifact ([references/agentic_loop_canon.md](references/agentic_loop_canon.md)). 4. **Retry means a changed approach.** Same command + same input = same failure. The retry directive says so; honor it. 5. **Budgets are terminal states, not suggestions.** `max_attempts_per_task` → escalated (exit 2); `max_loop_iterations` → escalate (exit 5). Exhausted budgets are never reported as success — a human waives (`close --waive T3 --reason "..."`), you don't. 6. **Fresh context beats long context.** Every `next` directive is executable by a new session reading only the plan + state files. Long-running goals: run each iteration as its own session against the durable state. 7. **State lives in `.agent-harness/`** — never in `.agenthub/`, `.autoresearch/`, or `docs/TC/` (those belong to sibling skills). 8. **Plan and state files are a trust boundary.** `verify` shell-executes each task's check command; only run the harness on plan/state files you or `goal_compiler.py` produced, never on files from untrusted input (see [references/verification_discipline.md](references/verification_discipline.md)).
## Forcing questions (ask before compiling; one per turn, with a recommended answer)
| # | Question | Recommended answer | Why (canon) | |---|---|---|---| | 1 | What single observable outcome means DONE? | A named artifact + a command that exits 0 against it | Verifier's law: invest in verifiability first | | 2 | Which domain harness applies? | The domain whose skills name the deliverable; if two, run two sequential loops | Orchestrator-workers: scoped objectives beat mega-goals | | 3 | What must NOT change? | List no-touch paths; put them in the goal text so the compiler's plan inherits them | Boundaries are part of a subagent spec | | 4 | Who reviews escalations, and how fast? | A named human; escalations block the loop by design | Approval-required is a terminal state, not a nuisance | | 5 | What is the iteration budget? | Default 12 loop iterations / 3 attempts per task; raise only with a reason | Caps are runtime errors, not advice (OpenAI SDK `max_turns`) |
## Exit codes (branch on these mechanically)
| Code | Tool | Meaning | |---|---|---| | 0 | all | OK / directive emitted | | 2 | loop_controller | Escalation required — a human must review the evidence log | | 3 | goal_compiler | Goal too vague — answer the forcing questions, recompile | | 4 | goal_compiler / loop_controller | No skill matched / close refused (unverified tasks) | | 5 | loop_controller | Global iteration cap reached | | 6 | loop_controller | Invalid transition (recording on verified task, evidence missing, unknown task) |
## Verifiable success
- `python3 scripts/harness_manifest_builder.py --sample`, `scripts/goal_compiler.py --sample`, and `scripts/loop_controller.py --sample` all exit 0. - A vague goal (`--goal "make it better"`) exits 3 and prints forcing questions. - `loop_controller.py close` on a state with an unverified task exits 4. - The demo loop in `loop_controller.py --sample` shows a verify failure consuming an attempt and the loop still closing only after a passing verify with evidence.
## Related skills
- **workflow-builder**: authoring deterministic `.js` scripts for Claude Code's Workflow tool. NOT for goal-to-close loop state (this skill). - **agenthub**: N parallel agents competing on ONE task in git worktrees. Use it *inside* a harness task that wants competing attempts. - **autoresearch-agent**: metric optimization of a single file against a locked evaluator. Use it when a task's done_when is "metric improves". - **tc-tracker**: per-code-change lifecycle records. Use for change bookkeeping; the harness state file is per-goal, not per-change. - **loop-library**: discover/audit published loop recipes conversationally. This skill is the executable enforcement of that vocabulary. - **ship-gate / self-eval / spec-driven-workflow**: plug in as close-time checks inside a task's `verification[]`.
See [references/domain_harness_design.md](references/domain_harness_design.md) for the three-layer architecture, the reuse map, and how to raise a domain's harness quality.
Technical details
- Version
- 1.0.0
- License
- MIT
- Last updated
- Aug 22, 2026
- Published
- Aug 22, 2026
Decision snapshot
Primary pick
24,795 GitHub stars
Audit
Install review
Install and adoption review
- Security
- 78/100
- Maintenance
- 100/100
- Install
- 92/100
Agent-proven evidence
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
- Success rate
- —
- Recent failure
- —
- Outcomes
- 0
- Output quality
- —
- Failed
- 0
- Not relevant
- 0
- Installs
- 0
- Risk blocked
- 0
- Setup needed
- 0
- Production
- 0
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Add to agent workflow
Free and open source. Review the report before installing into production agents.
Growth loop
Share kit
Scenario-led draft for agent-harness, ready for a manual X post.
agent-harness: Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiabl... 24.8K stars https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x
Optional reply with install command
Listing + install path for agent-harness: https://www.openagentskill.com/skills/alirezarezvani-agent-harness?ref=x Install: npx skills add alirezarezvani/claude-skills --skill agent-harness
Listing source
Registry indexed
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
- Creator
- alirezarezvani
- Indexed by
- OpenAgentSkill community index
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
Claim this skill listing
This Registry indexed listing is attributed to alirezarezvani but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Add the evidence badges to your README
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness/audit)
[](https://www.openagentskill.com/skills/alirezarezvani-agent-harness)Author
alirezarezvani
@alirezarezvani
Tags
Platform fit
Health signals
- GitHub stars
- 24.8K
- Quality score
- 54/100
- Last GitHub push
- Aug 22, 2026
- Framework hints
- Unknown
- OpenAgentSkill views
- 0
- Install copies
- 0
- Outbound clicks
- 0
Community signal
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Trust & safety
Sandbox only
- GitHub adoption25K GitHub starsPASS
- Stars/forks activity25K stars, 3.5K forks; issue activity unavailable in current metadataPASS
- Recent maintenancePushed todayPASS
- License clarityMITPASS
- README/SKILL.md completenessMetadata includes enough usage and workflow contextPASS
- Dependency/runtime riskcommand execution surfaceINFO
Related skills
Last30days Skill
Research the last 30 days across Reddit, X, YouTube, Hacker News, Polymarket, GitHub, and the web, then synthesize a grounded brief for an AI agent.
53.5K StarsAcademic Research Skills
Academic Research Skills for Claude Code: research → write → review → revise → finalize
38.4K StarsGPT Researcher
Run autonomous deep research over web and local sources
28.0K StarsDeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
19.8K Stars