OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Encuentra una habilidad para tu próxima tarea con Codex, Claude Code, Cursor y más.

Resultados · “evals”

34 Skills

Resultados: 34

Búsqueda e investigación

No visual example yet

Explore the skill

Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.

Precio sin confirmarBúsqueda e investigaciónClaude Code
132GitHub
Seguridad

No visual example yet

Explore the skill

Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…

Precio sin confirmarSeguridadClaude Code
132GitHub
Desarrollo

No visual example yet

Explore the skill

Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…

Precio sin confirmarDesarrolloClaude CodeRevisar antes de usar
24,7 milGitHub
Desarrollo

No visual example yet

Explore the skill

Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.

Precio sin confirmarDesarrolloClaude Code
132GitHub
Búsqueda e investigación

No visual example yet

Explore the skill

Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.

Precio sin confirmarBúsqueda e investigaciónClaude Code
132GitHub
Legal y cumplimiento

No visual example yet

Explore the skill

Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.

Precio sin confirmarLegal y cumplimientoClaude Code
132GitHub
IA y conocimiento

No visual example yet

Explore the skill

Future Agi

future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…

Precio sin confirmarIA y conocimientoRevisar antes de usar
1,2 milGitHub
IA y conocimiento

No visual example yet

Explore the skill

Langfuse

langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…

Precio sin confirmarIA y conocimientoOpenAI AgentsLangChainRevisar antes de usar
29,4 milGitHub
Desarrollo

No visual example yet

Explore the skill

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…

Precio sin confirmarDesarrolloClaude Code
5,3 milGitHub
Desarrollo

No visual example yet

Explore the skill

Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…

Precio sin confirmarDesarrolloClaude CodeOpenAI Agents
979GitHub

Guías y comparativas

Para desarrolladores