Skill comparison
Use this as a shortlist, then open the skill detail page before adopting.
Decision summary
Strongest overall
experiment-audit
Use this as a leading candidate, then validate the README and install path in your own agent stack.
Fastest prototype
experiment-audit
Best first install candidate based on install readiness and adoption.
Freshest repo
experiment-audit
Most recent maintenance signal among this shortlist.
| Signal | experiment-audit Audit experiment integrity before claiming results. Uses cross-model review (external reviewer backend) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user says \"审计实验\", \"check experiment integrity\", \"audit results\", \"实验诚实度\", or after experiments complete before writing claims. |
|---|---|
| Quality | 89/100 Excellent |
| Decision verdict | 100/100 Production-ready Use this as a leading candidate, then validate the README and install path in your own agent stack. |
| Adoption | 16K stars Verified outcomes are shown on each skill page |
| Freshness | Aug 26, 2026 |
| Use-case fit | |
| Workflow fit | |
| Platform hints | Claude Code, OpenAI Agents |
| Warnings | No OpenAgentSkill engagement data yet |
| Best for |
| Research agents workflows · Claude Code teams · teams that value GitHub adoption signals |
| Not ideal for | teams that need a vendor-supported SLA · high-compliance environments without internal security review |
| OpenAgentSkill engagement | 0 views 0 install copies |
| Install | $ npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill experiment-audit |