design-ai-benchmarking
Aperivue
Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm defini…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
1–4 / 4
Candidates in this shortlist, not the full registry. GitHub stars belong to repositories, not individual skills.
Results: 4
Aperivue
Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm defini…
agentscope-ai
Use when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing…
JoelLewis
Design, build, and optimize dashboards for RIA practice management with AUM tracking, revenue analytics, and KPI frameworks. Use when the user asks about tracking firm-l…
scdenney
Reviews an existing conjoint study for threats to inference and returns prioritized findings across five areas — design integrity (attributes, profile restrictions, task…