No visual example yet
Explore the skillGiskard Oss
Giskard-AI
🐢 Open-Source Evaluation & Testing library for LLM Agents
OPENAGENTSKILL / DIRECTORY
次のタスクに合うスキルを。Codex、Claude Code、Cursor などのツールを探せます。
7 Skills
検索結果: 7
No visual example yet
Explore the skillGiskard-AI
🐢 Open-Source Evaluation & Testing library for LLM Agents
No visual example yet
Explore the skillAgentEvalHQ
AgentEval is the comprehensive .NET toolkit for AI agent evaluation—tool usage validation, RAG quality metrics, stochastic evaluation, and model comparison—built first f…
No visual example yet
Explore the skillgithub
Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimi…
No visual example yet
Explore the skilldotnet
Scaffolds new agent skills for the dotnet/skills repository. Use when creating a new skill, generating SKILL.md files, writing a skill description that the runtime will…
No visual example yet
Explore the skilldotnet
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…
No visual example yet
Explore the skillOwl-Listener
Facilitate a structured team critique — framing, feedback rules, and actionable outcomes. Use when running a session with people in the room. For a solo expert review, u…
No visual example yet
Explore the skillsangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…