OPENAGENTSKILL / DIRECTORY

AI Agent Skills

为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。

搜索结果 · “evals”

34 Skills

搜索结果: 34

搜索与研究

暂未收录效果图

查看技能说明

Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.

价格未确认搜索与研究Claude Code
132GitHub
安全

暂未收录效果图

查看技能说明

Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…

价格未确认安全Claude Code
132GitHub
开发与测试

暂未收录效果图

查看技能说明

Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…

价格未确认开发与测试Claude Code使用前请审核
2.5万GitHub
开发与测试

暂未收录效果图

查看技能说明

Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.

价格未确认开发与测试Claude Code
132GitHub
搜索与研究

暂未收录效果图

查看技能说明

Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.

价格未确认搜索与研究Claude Code
132GitHub
法律与合规

暂未收录效果图

查看技能说明

Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.

价格未确认法律与合规Claude Code
132GitHub
AI 与知识库

暂未收录效果图

查看技能说明

Future Agi

future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…

价格未确认AI 与知识库使用前请审核
1155GitHub
AI 与知识库

暂未收录效果图

查看技能说明

Langfuse

langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…

价格未确认AI 与知识库OpenAI AgentsLangChain使用前请审核
2.9万GitHub
开发与测试

暂未收录效果图

查看技能说明

Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…

价格未确认开发与测试Claude CodeOpenAI Agents
979GitHub

指南与对比

开发者入口