No visual example yet
Explore the skillagent-evaluation
NeoLabHQ
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement qu…
OPENAGENTSKILL / DIRECTORY
Encuentra una habilidad para tu próxima tarea con Codex, Claude Code, Cursor y más.
7 Skills
Resultados: 7
No visual example yet
Explore the skillNeoLabHQ
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement qu…
No visual example yet
Explore the skillGiskard-AI
🐢 Open-Source Evaluation & Testing library for LLM Agents
No visual example yet
Explore the skillsangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
No visual example yet
Explore the skillAgentEvalHQ
AgentEval is the comprehensive .NET toolkit for AI agent evaluation—tool usage validation, RAG quality metrics, stochastic evaluation, and model comparison—built first f…
No visual example yet
Explore the skillgithub
Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimi…
No visual example yet
Explore the skilldotnet
Scaffolds new agent skills for the dotnet/skills repository. Use when creating a new skill, generating SKILL.md files, writing a skill description that the runtime will…
No visual example yet
Explore the skilldotnet
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…