OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Trouvez un skill pour votre prochaine tâche avec Codex, Claude Code, Cursor et plus encore.

Résultats · “evals”

34 Skills

Résultats: 34

Recherche

No visual example yet

Explore the skill

Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.

Prix non confirméRechercheClaude Code
132GitHub
Sécurité

No visual example yet

Explore the skill

Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…

Prix non confirméSécuritéClaude Code
132GitHub
Développement

No visual example yet

Explore the skill

Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…

Prix non confirméDéveloppementClaude CodeÀ vérifier avant usage
24,7 kGitHub
Développement

No visual example yet

Explore the skill

Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.

Prix non confirméDéveloppementClaude Code
132GitHub
Recherche

No visual example yet

Explore the skill

Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.

Prix non confirméRechercheClaude Code
132GitHub
Droit et conformité

No visual example yet

Explore the skill

Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.

Prix non confirméDroit et conformitéClaude Code
132GitHub
IA et connaissances

No visual example yet

Explore the skill

Future Agi

future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…

Prix non confirméIA et connaissancesÀ vérifier avant usage
1,2 kGitHub
IA et connaissances

No visual example yet

Explore the skill

Langfuse

langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…

Prix non confirméIA et connaissancesOpenAI AgentsLangChainÀ vérifier avant usage
29,4 kGitHub
Développement

No visual example yet

Explore the skill

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…

Prix non confirméDéveloppementClaude Code
5,3 kGitHub
Développement

No visual example yet

Explore the skill

Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…

Prix non confirméDéveloppementClaude CodeOpenAI Agents
979GitHub

Guides et comparaisons

Pour les développeurs