Skill comparison
Use this as a shortlist, then open the skill detail page before adopting.
Decision summary
Strongest overall
design-ai-benchmarking
Prototype with this skill first; keep a fallback candidate ready.
Fastest prototype
design-ai-benchmarking
Best first install candidate based on install readiness and adoption.
Freshest repo
design-ai-benchmarking
Most recent maintenance signal among this shortlist.
| Signal | design-ai-benchmarking Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm definition, decoupled multi-dimensional rubrics with anchors, planted calibration probes, reviewer-panel construction, inter-rater reliability targets, LLM-as-judge versus human-as-judge adjudication, construct-independence guards, and a structured rating-export schema. Use before data collection on an AI-vs-expert evaluation. |
|---|---|
| Quality | 71/100 Strong |
| Decision verdict | 70/100 Prototype first Prototype with this skill first; keep a fallback candidate ready. |
| Adoption | 283 stars Verified outcomes are shown on each skill page |
| Freshness | Sep 6, 2026 |
| Use-case fit | |
| Workflow fit | |
| Platform hints | Claude Code |
| Warnings | No critical security issues detected; the skill is advisory and uses standard file tools only. · No OpenAgentSkill engagement data yet |
Skill comparison
Use this as a shortlist, then open the skill detail page before adopting.
Decision summary
Strongest overall
design-ai-benchmarking
Prototype with this skill first; keep a fallback candidate ready.
Fastest prototype
design-ai-benchmarking
Best first install candidate based on install readiness and adoption.
Freshest repo
design-ai-benchmarking
Most recent maintenance signal among this shortlist.
| Signal | design-ai-benchmarking Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm definition, decoupled multi-dimensional rubrics with anchors, planted calibration probes, reviewer-panel construction, inter-rater reliability targets, LLM-as-judge versus human-as-judge adjudication, construct-independence guards, and a structured rating-export schema. Use before data collection on an AI-vs-expert evaluation. |
|---|---|
| Quality | 71/100 Strong |
| Decision verdict | 70/100 Prototype first Prototype with this skill first; keep a fallback candidate ready. |
| Adoption | 283 stars Verified outcomes are shown on each skill page |
| Freshness | Sep 6, 2026 |
| Use-case fit | |
| Workflow fit | |
| Platform hints | Claude Code |
| Warnings | No critical security issues detected; the skill is advisory and uses standard file tools only. · No OpenAgentSkill engagement data yet |
| Best for | Research agents workflows · Claude Code teams · builders willing to evaluate younger projects |
| Not ideal for | teams that need a vendor-supported SLA · production agents without a repository review |
| OpenAgentSkill engagement | 0 views 0 install copies |
| Install | $ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking |
| Best for | Research agents workflows · Claude Code teams · builders willing to evaluate younger projects |
| Not ideal for | teams that need a vendor-supported SLA · production agents without a repository review |
| OpenAgentSkill engagement | 0 views 0 install copies |
| Install | $ npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking |