OpenAgentSkill Registry Manifest Skill: design-ai-benchmarking Slug: aperivue-design-ai-benchmarking Category: design-creative Description: Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm definition, decoupled multi-dimensional rubrics with anchors, planted calibration probes, reviewer-panel construction, inter-rater reliability targets, LLM-as-judge versus human-as-judge adjudication, construct-independence guards, and a structured rating-export schema. Use before data collection on an AI-vs-expert evaluation. Agent fit: - Decision: 70/100 Prototype first - Primary fit: Research agents - Role: Fallback candidate Supply profile: - Track: Research and knowledge work - Scenario: Research agents - Applicable agents: Claude Code, CLI, Codex, Cursor - Maintenance: Pushed today - Risk: Needs review Trust: - Trust score: 70/100 Manual review - Audit: 78/100 Needs review Attribution: - Status: Registry indexed - Source: github candidate review - Creator: Aperivue - Claim URL: https://www.openagentskill.com/skills/aperivue-design-ai-benchmarking#claim-this-skill Install: npx skills add Aperivue/medsci-skills --skill design-ai-benchmarking URLs: - Web: https://www.openagentskill.com/skills/aperivue-design-ai-benchmarking - API: https://www.openagentskill.com/api/agent/skills/aperivue-design-ai-benchmarking - Install API: https://www.openagentskill.com/api/skills/aperivue-design-ai-benchmarking/install - Repository: https://github.com/Aperivue/medsci-skills/tree/main/skills/design-ai-benchmarking