暂未收录效果图
查看技能说明evals-specify
tikalk
Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
34 Skills
搜索结果: 34
暂未收录效果图
查看技能说明tikalk
Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.
暂未收录效果图
查看技能说明tikalk
Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…
暂未收录效果图
查看技能说明vasilyu1983
Designs coding-agent observability and evals. Use when measuring traces, replay, checkpoint lineage, quality trajectories, tool grading, regression, or cost.
暂未收录效果图
查看技能说明magnus919
Design, run, review, or release framework- and vendor-neutral evaluations and observability for AI agents. Use when defining agent evals, datasets, graders, tra
暂未收录效果图
查看技能说明promptfoo
Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…
暂未收录效果图
查看技能说明tikalk
Refine, cluster, and accept draft criteria into the published goldset. Isolates 20% holdout split and publishes goldset.md + goldset.json.
暂未收录效果图
查看技能说明tikalk
Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.
暂未收录效果图
查看技能说明tikalk
Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.
暂未收录效果图
查看技能说明tikalk
Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.
暂未收录效果图
查看技能说明future-agi
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…
暂未收录效果图
查看技能说明GitHamza0206
OpenSource Production ready Customer service with built in Evals and monitoring
暂未收录效果图
查看技能说明langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…
暂未收录效果图
查看技能说明ombharatiya
AI system design guide for engineers building production AI systems and evals.
暂未收录效果图
查看技能说明dotnet
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…
暂未收录效果图
查看技能说明microsoft
Prepare and publish a new version of the waza azd extension. USE FOR: "publish extension", "release new version", "bump version", "prepare release", "update changelog",…
暂未收录效果图
查看技能说明getsentry
Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…