OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Temukan skill untuk tugas berikutnya dengan Codex, Claude Code, Cursor, dan lainnya.

Hasil · “evals”

34 Skills

Hasil: 34

Pencarian & riset

No visual example yet

Explore the skill

Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.

Harga belum dikonfirmasiPencarian & risetClaude Code
132GitHub
Lihat skill
Keamanan

No visual example yet

Explore the skill

Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…

Harga belum dikonfirmasiKeamananClaude Code
132GitHub
Lihat skill
Desain & UI

No visual example yet

Explore the skill

Design, run, review, or release framework- and vendor-neutral evaluations and observability for AI agents. Use when defining agent evals, datasets, graders, tra

Harga belum dikonfirmasiDesain & UIClaude Code
76GitHub
Lihat skill
Pengembangan

No visual example yet

Explore the skill

Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…

Harga belum dikonfirmasiPengembanganClaude CodeTinjau sebelum digunakan
24,7 rbGitHub
Lihat skill
Browser & otomatisasi

No visual example yet

Explore the skill

Refine, cluster, and accept draft criteria into the published goldset. Isolates 20% holdout split and publishes goldset.md + goldset.json.

Harga belum dikonfirmasiBrowser & otomatisasiClaude Code
132GitHub
Lihat skill
Pengembangan

No visual example yet

Explore the skill

Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.

Harga belum dikonfirmasiPengembanganClaude Code
132GitHub
Lihat skill
Pencarian & riset

No visual example yet

Explore the skill

Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.

Harga belum dikonfirmasiPencarian & risetClaude Code
132GitHub
Lihat skill
Hukum & kepatuhan

No visual example yet

Explore the skill

Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.

Harga belum dikonfirmasiHukum & kepatuhanClaude Code
132GitHub
Lihat skill
AI & pengetahuan

No visual example yet

Explore the skill

Future Agi

future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…

Harga belum dikonfirmasiAI & pengetahuanTinjau sebelum digunakan
1,2 rbGitHub
Lihat skill
AI & pengetahuan

No visual example yet

Explore the skill

Langfuse

langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…

Harga belum dikonfirmasiAI & pengetahuanOpenAI AgentsLangChainTinjau sebelum digunakan
29,4 rbGitHub
Lihat skill
Pengembangan

No visual example yet

Explore the skill

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…

Harga belum dikonfirmasiPengembanganClaude Code
5,3 rbGitHub
Lihat skill
Pengembangan

No visual example yet

Explore the skill

Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…

Harga belum dikonfirmasiPengembanganClaude CodeOpenAI Agents
979GitHub
Lihat skill

Panduan & perbandingan

Untuk pengembang