eval
agentagon
Create, repair or reuse an agent evaluation from agreed behaviors and scoring; validate sensitivity, independently review, freeze and prepare eval-only delivery.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 43
Results: 43
agentagon
Create, repair or reuse an agent evaluation from agreed behaviors and scoring; validate sensitivity, independently review, freeze and prepare eval-only delivery.
skyf0xx
Use for a periodic check-in or an honest audit of progress against GOAL.json's success criteria, including whether people involved are actually delivering. Scores each c…
onejune2018
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表…
darkrishabh
A test runner for agentskills.io-style AI agent skills
sangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
github
Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimi…
promptfoo
Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…
paperclipai
Add or extend a Paperclip Runner protocol evaluation definition, roster, assertion, or report fixture with provenance and narrow validation.
dotnet
A skill with no eval.yaml for testing discovery behavior.
awslabs
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). Use when the user says "is my dataset okay", "evaluate my data", "check my…
sangrokjung
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
paperclipai
Add or extend a Paperclip full-stack runner E2E workflow, fixture, matcher, or report evidence path for local or Daytona execution.
minhnv0807
Use when marketing numbers look bad and the user needs to know WHY — layered root-cause diagnosis across tracking, delivery, creative, landing page, offer, and audience…
eval-exec
NEO Emacs (WIP): GPU powered Emacs written in Rust with a modern display engine. Aiming for modern design & multi-threaded Elisp, 10x performance, zero-pause GC and 100%…
ConardLi
A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval
yaojingang
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.