Creator ยท addyosmani
Last updated ยท Sep 1, 2026
Subjects every non-trivial decision to a fresh-context adversarial review before it stands. Use when correctness matters more than speed, when working in unfamiliar code, when stakes are high (production, security-sensitive logic, irreversible operations), or any time a confident
Sandbox only
Install targets
Codex install prompt
Install the "doubt-driven-development" agent skill from https://github.com/addyosmani/agent-skills/tree/main/skills/doubt-driven-development. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Subjects every non-trivial decision to a fresh-context adversarial review before it stands. Use when correctness matters more than speed, when working in unfamiliar code, when stakes are high (production, security-sensitive logic, irreversible operations), or any time a confident output would be cheaper to verify now than to debug later. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"addyosmani-doubt-driven-development","task":"Install doubt-driven-development","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
Testing and QA
I need my agent to test a web app, reproduce bugs, and verify fixes.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add addyosmani/agent-skills --skill doubt-driven-development
Maintenance
fresh
10d since push
Risk
Needs review
Dependency or permission surface needs review
GitHub quality
91K
95/100 Quality ยท 79/100 Trust
Coverage tags
Review notes
Dependency or permission surface needs review ยท Permission surface may require sandboxing
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
91K GitHub stars
Repo activity
91K stars, 9.8K forks
Maintenance
10d since push
License
MIT
Install
npx skills add addyosmani/agent-skills --skill doubt-driven-development
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add addyosmani/agent-skills --skill doubt-driven-developmentDo not use when
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
high
Skill metadata references credentials, tokens, environment variables, or secret-bearing workflows.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20doubt-driven-development%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20doubt-driven-development%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/addyosmani-doubt-driven-development/install
Agent should check
Copy prompt
Task: Use doubt-driven-development in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20doubt-driven-development%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/addyosmani-doubt-driven-development/install
Install command: npx skills add addyosmani/agent-skills --skill doubt-driven-development
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/addyosmani-doubt-driven-development/install
LLM text format
/api/skills/addyosmani-doubt-driven-development/install?format=text
Find alternatives
/api/skills/search?q=doubt-driven-development&limit=3
Agent prompt
Use doubt-driven-development for this task. Review https://www.openagentskill.com/api/skills/addyosmani-doubt-driven-development/install, then install with: npx skills add addyosmani/agent-skills --skill doubt-driven-developmentRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/addyosmani-doubt-driven-development
LLM text
/api/registry/manifest/addyosmani-doubt-driven-development?format=text
Install alias
/api/registry/install/addyosmani-doubt-driven-development
Recommend
/api/registry/recommend?task=Use%20doubt-driven-development%20in%20an%20agent%20workflow&limit=3
Agent fit
Testing and QA
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Testing and QA
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
PASS91K GitHub stars
Stars/forks activity
PASS91K stars, 9.8K forks; issue activity unavailable in current metadata
Recent maintenance
PASS10d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Verify behavior
I need my agent to test a web app, reproduce bugs, and verify fixes.
Build and ship code
I need a coding agent that can understand a repository, edit code, and review pull requests.
Search private knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Workflow fit
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Ingest, retrieve, and cite
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
๐ต๏ธโโ๏ธ Collect a dossier on a person by username from 3000+ sites
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical is the open-source platform for secrets, certificates, and privileged access management.
--- name: doubt-driven-development description: Subjects every non-trivial decision to a fresh-context adversarial review before it stands. Use when correctness matters more than speed, when working in unfamiliar code, when stakes are high (production, security-sensitive logic, irreversible operations), or any time a confident output would be cheaper to verify now than to debug later. ---
# Doubt-Driven Development
## Overview
A confident answer is not a correct one. Long sessions accumulate context that quietly turns assumptions into "facts" without anyone noticing. Doubt-driven development is the discipline of materializing a fresh-context reviewer โ biased to **disprove**, not approve โ before any non-trivial output stands.
This is not `/review`. `/review` is a verdict on a finished artifact. This is an in-flight posture: non-trivial decisions get cross-examined while course-correction is still cheap.
## When to Use
A decision is **non-trivial** when at least one of these is true:
- It introduces or modifies branching logic - It crosses a module or service boundary - It asserts a property the type system or compiler cannot verify (thread safety, idempotence, ordering, invariants) - Its correctness depends on context the future reader cannot see - Its blast radius is irreversible (production deploy, data migration, public API change)
Apply the skill when:
- About to make an architectural decision under uncertainty - About to commit non-trivial code - About to claim a non-obvious fact ("this is safe", "this scales", "this matches the spec") - Working in code you don't fully understand
**When NOT to use:**
- Mechanical operations (renaming, formatting, file moves) - Following a clear, unambiguous user instruction - Reading or summarizing existing code - One-line changes with obvious correctness - Pure tooling operations (running tests, listing files) - The user has explicitly asked for speed over verification
If you doubt every keystroke, you ship nothing. The skill applies only to non-trivial decisions as defined above.
## Loading Constraints
This skill is designed for the **main-session orchestrator**, where Step 3 (DOUBT, detailed below) can spawn a fresh-context reviewer.
- **Do NOT add this skill to a persona's `skills:` frontmatter.** A persona that follows Step 3 would spawn another persona โ the orchestration anti-pattern explicitly forbidden by `../../references/orchestration-patterns.md` ("personas do not invoke other personas"). - **If you find yourself applying this skill from inside a subagent context** (where Claude Code prevents nested subagent spawn): the preferred path is to surface to the user that doubt-driven cannot run nested and let the main session handle it. As a last resort only, a degraded self-questioning fallback exists โ rewrite ARTIFACT + CONTRACT as a fresh self-prompt with a hard mental separator from your prior reasoning, and walk Steps 1โ5. This is **not fresh-context review** (you carry your own context with you), so flag the result as degraded and prefer escalation whenever the user is reachable.
## The Process
Copy this checklist when applying the skill:
``` Doubt cycle: - [ ] Step 1: CLAIM โ wrote the claim + why-it-matters - [ ] Step 2: EXTRACT โ isolated artifact + contract, stripped reasoning - [ ] Step 3: DOUBT โ invoked fresh-context reviewer with adversarial prompt - [ ] Step 4: RECONCILE โ classified every finding against the artifact text - [ ] Step 5: STOP โ met stop condition (trivial findings, 3 cycles, or user override) ```
### Step 1: CLAIM โ Surface what stands
Name the decision in two or three lines:
``` CLAIM: "The new caching layer is thread-safe under the read-heavy workload described in the spec." WHY THIS MATTERS: a race here corrupts user data and is hard to detect in QA. ```
If you can't write the claim that compactly, you have a vibe, not a decision. Surface it before scrutinizing it.
### Step 2: EXTRACT โ Smallest reviewable unit
A fresh-context reviewer needs the **artifact** and the **contract**, not the journey.
- Code: the diff or the function โ not the whole file - Decision: the proposal in 3โ5 sentences plus the constraints it has to satisfy - Assertion: the claim plus the evidence that supposedly supports it (kept distinct from the Step 1 CLAIM block, which is the orchestrator's hypothesis under scrutiny)
Strip your reasoning. If you hand over conclusions, you'll get back validation of your conclusions. The unit must be small enough that a reviewer can hold it in mind in one read โ if it's a 500-line PR, decompose first.
### Step 3: DOUBT โ Invoke the fresh-context reviewer
The reviewer's prompt **must be adversarial**. Framing decides the answer.
``` Adversarial review. Find what is wrong with this artifact. Assume the author is overconfident. Look for: - Unstated assumptions - Edge cases not handled - Hidden coupling or shared state - Ways the contract could be violated - Existing conventions this might break - Failure modes under unexpected input
Do NOT validate. Do NOT summarize. Find issues, or state explicitly that you cannot find any after thorough examination.
ARTIFACT: <paste artifact> CONTRACT: <paste contract> ```
**Pass ARTIFACT + CONTRACT only. Do NOT pass the CLAIM.** Handing the reviewer your conclusion biases it toward agreement. The reviewer must independently determine whether the artifact satisfies the contract.
In Claude Code, the role-based reviewers in `agents/` start with isolated context by design and are usable here โ see `agents/` for the roster and per-domain match.
**The adversarial prompt above takes precedence over the persona's default response shape.** Personas like `code-reviewer` are written to produce balanced verdicts with both strengths and weaknesses; doubt-driven needs issues-only output. Paste the adversarial prompt verbatim into the invocation so it overrides the persona's default. If a persona's response shape can't be overridden cleanly, fall back to a generic subagent with the adversarial prompt.
#### Cross-model escalation
A single-model reviewer shares blind spots with the original author โ a colder, different-architecture model catches them. Doubt-driven is already opt-in for non-trivial decisions, so within that scope offering cross-model is part of the skill's value, not optional friction.
**Interactive sessions: always offer. Never silently skip.**
**Step 1: Ask the user**
After the single-model review in Step 3 above, but before RECONCILE, pause and ask:
> *"Single-model review complete. Want a cross-model second opinion? Options: Gemini CLI, Codex CLI, manual external review (you paste it elsewhere), or skip."*
This question is mandatory in every interactive doubt cycle โ even on artifacts that feel low-stakes. The user โ not the agent โ decides whether the cost is worth it. The agent's job is to surface the choice.
**Step 2: If the user picks a CLI โ verify, then invoke**
1. Check the tool is in PATH (`which gemini`, `which codex`). 2. Test it works (`gemini --version` or equivalent) before passing the full prompt โ a stale or broken binary may pass `which` but fail on real input. 3. Confirm the exact invocation with the user, including required flags, auth, and env vars (e.g., API keys). Implementations vary; never assume. 4. Pass ARTIFACT + CONTRACT + the adversarial prompt **only**. No session context, no CLAIM. 5. Mind shell escaping. If the artifact contains quotes, `$(...)`, or backticks, prefer stdin (`echo โฆ | gemini`) or a heredoc over inline `-p "โฆ"`. When in doubt, ask the user to confirm the invocation before running it. 6. Take the output into Step 4 (RECONCILE).
**Never interpolate the artifact into a shell-quoted argument.** Code, markdown, and review prompts routinely contain backticks, `$(...)`, and quote characters that will either truncate the prompt or execute embedded shell. Write the full prompt to a file and pipe it through stdin.
Example shapes (verify flags against your installed tool โ syntax differs across implementations and versions):
```bash # Write the adversarial prompt + ARTIFACT + CONTRACT to a temp file first. # Then pipe via stdin so shell metacharacters in the artifact stay inert.
# Codex (read-only sandbox keeps the CLI from writing to your workspace): codex exec --sandbox read-only -C <repo-path> - < /tmp/doubt-prompt.md
# Gemini ('--approval-mode plan' is read-only; '-p ""' triggers non-interactive # mode and the prompt is read from stdin): gemini --approval-mode plan -p "" < /tmp/doubt-prompt.md ```
A read-only sandbox is the load-bearing detail: a doubt artifact may itself contain instructions (intentional or accidental prompt injection) that the cross-model CLI would otherwise execute against your workspace.
**Step 3: If the CLI is unavailable or fails**
Surface the failure explicitly. Offer: run it manually, try a different tool, or skip. Do not silently fall back to single-model โ the user should know cross-model didn't happen.
**Step 4: If the user skips**
Acknowledge the skip in the output (*"Proceeding with single-model findings only"*) and continue to RECONCILE. Skipping is fine; silent skipping is not.
**Non-interactive contexts** (CI, `/loop`, autonomous-loop, scheduled runs):
- Cross-model is **skipped**, and the skip must be **announced** in the output: *"Cross-model skipped: non-interactive context."* - **Never invoke an external CLI without explicit user authorization** โ this is a load-bearing safety property.
Cross-model adds cost, latency, and tool fragility. The agent surfaces the choice every cycle; the user decides whether this artifact warrants it.
### Step 4: RECONCILE โ Fold findings back
The reviewer's output is data, not verdict. **You are still the orchestrator.** Re-read the artifact text against each finding before classifying โ rubber-stamping the reviewer is the same failure mode as ignoring it.
For each finding, classify in this **precedence order** (first matching class wins):
1. **Contract misread** โ reviewer flagged something specifically because the CONTRACT you provided was unclear or incomplete. Fix the contract first, re-classify on the next cycle. 2. **Valid + actionable** โ real issue requiring a change to the artifact. Change it, re-loop. 3. **Valid trade-off** โ issue is real but cost of fixing exceeds cost of accepting. Document the trade-off explicitly so the user sees it. 4. **Noise** โ reviewer flagged something that's actually correct under context the reviewer didn't have. Note it, move on, and ask: would adding that context to the contract have prevented the false flag?
A fresh reviewer can be wrong because it lacks context. Don't defer just because it's "fresh."
### Step 5: STOP โ Bounded loop, not recursion
Stop when:
- Next iteration returns only trivial or already-considered findings, **or** - 3 cycles completed (escalate to user, don't grind a fourth alone), **or** - User explicitly says "ship it"
If after 3 cycles the reviewer still surfaces substantive issues, the artifact may not be ready. Surface this to the user โ three unresolved cycles is information about the artifact, not a reason to keep looping.
If 3 cycles is "obviously insufficient" because the artifact is large: the artifact is too big โ return to Step 2 and decompose. Do not lift the bound.
## Common Rationalizations
| Rationalization | Reality | |---|---| | "I'm confident, skip the doubt step" | Confidence correlates poorly with correctness on novel problems. Moments of certainty are exactly when blind spots hide. | | "Spawning a reviewer is expensive" | Debugging a wrong commit in production is more expensive. The check is bounded; the bug isn't. | | "The reviewer will just nitpick" | Only if unscoped. Constrain the prompt to "issues that would make this fail under the contract." | | "I'll do doubt at the end with `/review`" | `/review` is a final gate. Doubt-driven catches wrong directions early
Source provenance
Decision snapshot
91,373 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for doubt-driven-development, ready for a manual X post.
doubt-driven-development: Subjects every non-trivial decision to a fresh-context adversarial review before it stands. U... 91.4K stars https://www.openagentskill.com/skills/addyosmani-doubt-driven-development?ref=x
Listing + install path for doubt-driven-development: https://www.openagentskill.com/skills/addyosmani-doubt-driven-development?ref=x Install: npx skills add addyosmani/agent-skills --skill doubt-driven-development
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to addyosmani but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/addyosmani-doubt-driven-development?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/addyosmani-doubt-driven-development?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/addyosmani-doubt-driven-development/audit)
[](https://www.openagentskill.com/skills/addyosmani-doubt-driven-development?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)addyosmani
@addyosmani
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
๐ต๏ธโโ๏ธ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K StarsPermission surface
secrets or environment access, shell or command execution
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness