Creator · wanshuiyin
Last updated · Sep 2, 2026
Audit experiment integrity before claiming results. Uses cross-model review (external reviewer backend) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user says \"审计实验\", \"check experiment integrity\", \"audit results
Review then install
Install targets
Codex install prompt
Install the "experiment-audit" agent skill from https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/tree/main/skills/experiment-audit. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Audit experiment integrity before claiming results. Uses cross-model review (external reviewer backend) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user says \"审计实验\", \"check experiment integrity\", \"audit results\", \"实验诚实度\", or after experiments complete before writing claims. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"wanshuiyin-experiment-audit","task":"Install experiment-audit","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Deep research, source comparison, literature review, RAG, knowledge search, and reports.
Scenario
Research agents
I need my agent to research a topic, compare sources, and produce a concise report.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-audit
Maintenance
fresh
13d since push
Risk
Safe to try
Quality score needs review
GitHub quality
16K
89/100 Quality · 85/100 Trust
Coverage tags
Review notes
Quality score needs review
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
ExcellentHigh-confidence pick with strong adoption and healthy maintenance signals.
Trust
Review then installGood shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
Audit
Safe to tryA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Use as the primary candidate after human or sandbox review.
Stars
16K GitHub stars
Repo activity
16K stars, 1.4K forks
Maintenance
13d since push
License
MIT
Install
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-audit
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-auditDo not use when
Agent safety v2
Usable candidate, but the agent should surface permission and audit notes before installation.
Require human approval before installing into a real workspace.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20experiment-audit%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20experiment-audit%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/wanshuiyin-experiment-audit/install
Agent should check
Copy prompt
Task: Use experiment-audit in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20experiment-audit%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/wanshuiyin-experiment-audit/install
Install command: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-audit
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/wanshuiyin-experiment-audit/install
LLM text format
/api/skills/wanshuiyin-experiment-audit/install?format=text
Find alternatives
/api/skills/search?q=experiment-audit&limit=3
Agent prompt
Use experiment-audit for this task. Review https://www.openagentskill.com/api/skills/wanshuiyin-experiment-audit/install, then install with: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-auditRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/wanshuiyin-experiment-audit
LLM text
/api/registry/manifest/wanshuiyin-experiment-audit?format=text
Install alias
/api/registry/install/wanshuiyin-experiment-audit
Recommend
/api/registry/recommend?task=Use%20experiment-audit%20in%20an%20agent%20workflow&limit=3
Agent fit
Research agents
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Role in stack
Primary pick
Primary fit
Research agents
Trust label
Production-ready
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Good shortlist signal, but the agent should review audit notes, install policy, and outcome evidence before running it.
GitHub adoption
PASS16K GitHub stars
Stars/forks activity
PASS16K stars, 1.4K forks; issue activity unavailable in current metadata
Recent maintenance
PASS13d since push
License clarity
PASSMIT
Good signals
Review before install
Recommended action
Use as the primary candidate after human or sandbox review.
Quality profile
High-confidence pick with strong adoption and healthy maintenance signals.
Workflow fit
Investigate faster
I need my agent to research a topic, compare sources, and produce a concise report.
Automate repeated work
I need my agent to automate a repeated workflow across tools and files.
Publish consistently
I need my agent to turn research and product updates into useful content drafts.
Workflow fit
Find, compare, and synthesize
A workflow for agents that gather sources, compare claims, summarize long material, and draft useful research briefs.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
Infisical is the open-source platform for secrets, certificates, and privileged access management.
--- name: experiment-audit description: "Audit experiment integrity before claiming results. Uses cross-model review (external reviewer backend) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user says \"审计实验\", \"check experiment integrity\", \"audit results\", \"实验诚实度\", or after experiments complete before writing claims." argument-hint: "[experiment-dir-or-results-path]" allowed-tools: Bash(*), Read, Write, Edit, Grep, Glob, mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_review__review, mcp__manual_review__review_reply ---
# Experiment Audit: Cross-Model Integrity Verification
> 🔒 **Do not wrap this skill in `/loop`, `/schedule`, or `CronCreate`.** It is > verdict-bearing — it judges experiment integrity. Re-running that verdict on a > timer adds no new signal, and a loop that accepts its own output to decide > when to stop crosses into self-acquittal (`acceptance-gate.md`). Schedule the > *external wait that precedes it* — experiments done → then audit **once**. See > [`shared-references/external-cadence.md`](../shared-references/external-cadence.md).
Audit experiment integrity for: **$ARGUMENTS**
## Why This Exists
LLM agents can produce fraudulent experimental results through: 1. **Fake ground truth** — creating synthetic "reference" from model outputs, then reporting high agreement as performance 2. **Score normalization** — dividing metrics by the model's own max to get 0.99+ 3. **Phantom results** — claiming numbers from files that don't exist or functions never called 4. **Insufficient scope** — reporting 2-scene pilots as "comprehensive evaluation"
These are NOT intentional deception — they are failure modes of optimizing agents that lack integrity constraints. This skill adds that constraint.
## Core Principle
**The executor collects file paths. The external reviewer backend reads code and judges integrity. The executor does NOT participate in integrity judgment.**
This follows `shared-references/reviewer-independence.md` and `shared-references/experiment-integrity.md`.
## Constants
- **REVIEWER_BACKEND = `codex`** — Default: Codex MCP (ultra). Override with `— reviewer: oracle-pro` for Oracle MCP, or `— reviewer: manual` for Manual Review MCP. If manual-review MCP is unavailable, stop and print the install command; do not fall back to Codex. See `shared-references/reviewer-routing.md`.
## Reviewer Calling Convention
When calling the reviewer, branch on REVIEWER_BACKEND:
**If REVIEWER_BACKEND = `codex`:** Use `mcp__codex__codex` for new review threads. Use `mcp__codex__codex-reply` for follow-up rounds (reuse threadId).
**If REVIEWER_BACKEND = `manual`:** Use `mcp__manual_review__review` for new review threads with: prompt: [exact same prompt that would go to Codex] config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true} Save the returned `threadId`. Use `mcp__manual_review__review_reply` for follow-up rounds with: threadId: [saved manual-review threadId] prompt: [follow-up prompt] config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true}
Prompt fidelity: the manual prompt must be exactly the same text that Codex would receive. Review tracing applies equally to both backends.
## Workflow
### Step 1: Collect Artifacts (Executor — Claude)
Locate and list these files WITHOUT reading or summarizing their content:
``` Scan project directory for: 1. Evaluation scripts: *eval*.py, *metric*.py, *test*.py, *benchmark*.py 2. Result files: *.json, *.csv in results/, outputs/, logs/ 3. Ground truth paths: look in eval scripts for data loading (dataset paths, GT references) 4. Experiment tracker: EXPERIMENT_TRACKER.md, EXPERIMENT_LOG.md 5. Paper claims: NARRATIVE_REPORT.md, paper/sections/*.tex, PAPER_PLAN.md 6. Config files: *.yaml, *.toml, *.json configs with metric definitions ``` A verdict-bearing manual response MUST begin with `Reviewer-Model: <exact-model-id>` — pass the model THIS session is actually running as in `executor_model`. Missing, unknown, or same-family identity cannot acquit; emit `REVIEW_UNAVAILABLE` rather than guessing. If the executor model cannot be named, manual review's cross-family claim is unprovable — say so in the report instead of asserting it.
**DO NOT summarize, interpret, or explain any file content.** Only collect paths.
### Step 2: Send to Reviewer
Based on the selected reviewer backend (see Reviewer Calling Convention), pass ONLY file paths and the audit checklist to the reviewer. The reviewer reads everything directly.
For `codex`, call `mcp__codex__codex` with: - `model: gpt-5.6-sol` - `config: {"model_reasoning_effort": "ultra"}` - `sandbox: read-only` - `cwd: [project directory]` - `prompt: [the exact full prompt below]`
For `manual`, call `mcp__manual_review__review` with: - `config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true}` - `prompt: [the exact full prompt below]`
Manual review cannot use Codex-only `model`, `sandbox`, or `cwd`; include the same file paths in the prompt so the user can inspect them.
Use this exact prompt for both backends:
``` You are an experiment integrity auditor. Start from the assumption that the evaluation is compromised somewhere — your job is to find where. Be adversarial. Trust nothing the author tells you — verify everything yourself. Read ALL files listed below and check for the following fraud patterns.
Files to read: - Evaluation scripts: [list paths] - Result files: [list paths] - Experiment tracker: [list paths] - Paper claims: [list paths] - Config files: [list paths]
## Audit Checklist
### A. Ground Truth Provenance For each evaluation script: 1. Where does "ground truth" / "reference" / "target" come from? 2. Is it loaded from the DATASET, or generated/derived from MODEL OUTPUTS? 3. If derived: is it explicitly labeled as proxy evaluation? 4. Are official eval scripts used when available for this benchmark? FAIL if: GT is derived from model outputs without explicit proxy labeling.
### B. Score Normalization For each metric computation: 1. Is any metric divided by max/min/mean of the model's OWN output? 2. Are raw scores reported alongside any normalized scores? 3. Are any scores suspiciously close to 1.0 or 100%? FAIL if: Normalization denominator comes from prediction statistics.
### C. Result File Existence For each claim in the paper/narrative: 1. Does the referenced result file actually exist? 2. Does the claimed metric key exist in that file? 3. Does the claimed NUMBER match what's in the file? 4. Is the experiment tracker status DONE (not TODO/IN_PROGRESS)? FAIL if: Claimed results reference nonexistent files or mismatched numbers.
### D. Dead Code Detection For each metric function defined in eval scripts: 1. Is it actually CALLED in any evaluation pipeline? 2. Does its output appear in any result file? WARN if: Metric functions exist but are never called.
### E. Scope Assessment 1. How many scenes/datasets/configurations were actually tested? 2. How many seeds/runs per configuration? 3. Does the paper use words like "comprehensive", "extensive", "robust"? 4. Is the actual scope sufficient for those claims? WARN if: Scope language exceeds actual evidence.
### F. Evaluation Type Classification Classify each evaluation as: - real_gt: uses dataset-provided ground truth - synthetic_proxy: uses model-generated reference - self_supervised_proxy: no GT by design - simulation_only: simulated environment - human_eval: human judges
## Output Format
For each check (A-F), report: - Status: PASS | WARN | FAIL - Evidence: exact file:line references - Details: what specifically was found
Overall verdict: PASS | WARN | FAIL Be thorough. Read every eval script line by line. ```
### Step 3: Parse and Write Report (Executor — Claude)
Parse the reviewer's response and write `EXPERIMENT_AUDIT.md`:
```markdown # Experiment Audit Report
**Date**: [today] **Auditor**: External reviewer backend, ultra reasoning (cross-model, read-only) **Project**: [project name]
## Overall Verdict: [PASS | WARN | FAIL]
## Integrity Status: [pass | warn | fail]
## Checks
### A. Ground Truth Provenance: [PASS|WARN|FAIL] [details + file:line evidence]
### B. Score Normalization: [PASS|WARN|FAIL] [details]
### C. Result File Existence: [PASS|WARN|FAIL] [details]
### D. Dead Code Detection: [PASS|WARN|FAIL] [details]
### E. Scope Assessment: [PASS|WARN|FAIL] [details]
### F. Evaluation Type: [real_gt | synthetic_proxy | ...] [classification + evidence]
## Action Items - [specific fixes if WARN or FAIL]
## Claim Impact - Claim 1: [supported | needs qualifier | unsupported] - Claim 2: ... ```
Also write `EXPERIMENT_AUDIT.json` for machine consumption:
```json { "date": "2026-04-10", "auditor": "external-reviewer-ultra", "overall_verdict": "warn", "integrity_status": "warn", "checks": { "gt_provenance": {"status": "pass", "details": "..."}, "score_normalization": {"status": "warn", "details": "..."}, "result_existence": {"status": "pass", "details": "..."}, "dead_code": {"status": "pass", "details": "..."}, "scope": {"status": "warn", "details": "..."}, "eval_type": "real_gt" }, "claims": [ {"id": "C1", "impact": "supported"}, {"id": "C2", "impact": "needs_qualifier"} ] } ```
### Step 4: Print Summary
``` 🔬 Experiment Audit Complete
GT Provenance: ✅ PASS — real dataset GT used Score Normalization: ⚠️ WARN — boundary metric uses self-reference Result Existence: ✅ PASS — all files exist, numbers match Dead Code: ✅ PASS — all metric functions called Scope: ⚠️ WARN — 2 scenes, paper says "comprehensive"
Overall: ⚠️ WARN See EXPERIMENT_AUDIT.md for details. ```
## Integration with Other Skills
### Automatic in /research-pipeline (advisory, never blocks)
When integrated into the pipeline, this skill runs automatically after `/experiment-bridge` and before `/auto-review-loop`:
``` /experiment-bridge → results ready ↓ /experiment-audit (automatic, advisory) ├── PASS → continue normally ├── WARN → print ⚠️ warning, continue, tag claims as [INTEGRITY: WARN] └── FAIL → print 🔴 alert, continue, tag claims as [INTEGRITY CONCERN] ↓ /auto-review-loop → proceeds with integrity tags visible to reviewer ```
**Never blocks the pipeline.** Even on FAIL, the pipeline continues — but claims carry visible integrity tags.
### Read by /result-to-claim (if exists)
``` if EXPERIMENT_AUDIT.json exists: read integrity_status attach to verdict: {claim_supported: "yes", integrity_status: "warn"} if integrity_status == "fail": downgrade verdict display: "yes [INTEGRITY CONCERN]" else: verdict as normal, integrity_status = "unavailable" mark as "provisional — no integrity audit" ```
### Read by /paper-write (if exists)
``` if EXPERIMENT_AUDIT.json exists AND integrity_status == "fail": add footnote to affected claims: "Note: integrity audit flagged concerns with this evaluation" ```
## Key Rules
- **Reviewer independence**: executor collects paths, reviewer judges. Period. - **Never block**: warn loudly, never halt the pipeline. - **File-as-switch**: no EXPERIMENT_AUDIT.md = skill was never run = zero impact on existing behavior. - **Cross-model**: the reviewer MUST be a different model family from the executor. - **Honest about limits**: the audit catches common patterns, not all possible fraud. It is a safety net, not a guarantee.
## Acknowledgements
Motivated by communit
Source provenance
Decision snapshot
15,641 GitHub stars
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for experiment-audit, ready for a manual X post.
experiment-audit: Audit experiment integrity before claiming results. Uses cross-model review (external reviewe... 15.6K stars https://www.openagentskill.com/skills/wanshuiyin-experiment-audit?ref=x
Listing + install path for experiment-audit: https://www.openagentskill.com/skills/wanshuiyin-experiment-audit?ref=x Install: npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-a...
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to wanshuiyin but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-audit?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-audit?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-audit/audit)
[](https://www.openagentskill.com/skills/wanshuiyin-experiment-audit?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)wanshuiyin
@wanshuiyin
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Review then install
Wazuh
Wazuh - The Open Source Security Platform. Unified XDR and SIEM protection for endpoints and cloud workloads.
16.3K StarsMaigret
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
32.9K StarsNuclei
Nuclei is a fast, customizable vulnerability scanner powered by the global security community and built on a simple YAML-based DSL, enabling collaboration to tackle trending vulnerabilities on the internet. It helps you find vulnerabilities in your applications, APIs, networks, DNS, and cloud configurations.
29.2K StarsInfisical
Infisical is the open-source platform for secrets, certificates, and privileged access management.
27.4K StarsPermission surface
shell or command execution, filesystem or document access
Agent outcomes
No agent outcome data yet
Docs
Strong README/SKILL.md context
Risk summary
Install readiness