OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Finde den passenden Skill für deine nächste Aufgabe mit Codex, Claude Code, Cursor und mehr.

Ergebnisse · “evals”

34 Skills

Ergebnisse: 34

Suche & Recherche

No visual example yet

Explore the skill

Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.

Preis unbestätigtSuche & RechercheClaude Code
132GitHub
Skill ansehen
Sicherheit

No visual example yet

Explore the skill

Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…

Preis unbestätigtSicherheitClaude Code
132GitHub
Skill ansehen
Design & UI

No visual example yet

Explore the skill

Design, run, review, or release framework- and vendor-neutral evaluations and observability for AI agents. Use when defining agent evals, datasets, graders, tra

Preis unbestätigtDesign & UIClaude Code
76GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…

Preis unbestätigtEntwicklungClaude CodeVor Nutzung prüfen
24.743GitHub
Skill ansehen
Browser & Automation

No visual example yet

Explore the skill

Refine, cluster, and accept draft criteria into the published goldset. Isolates 20% holdout split and publishes goldset.md + goldset.json.

Preis unbestätigtBrowser & AutomationClaude Code
132GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.

Preis unbestätigtEntwicklungClaude Code
132GitHub
Skill ansehen
Suche & Recherche

No visual example yet

Explore the skill

Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.

Preis unbestätigtSuche & RechercheClaude Code
132GitHub
Skill ansehen
Recht & Compliance

No visual example yet

Explore the skill

Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.

Preis unbestätigtRecht & ComplianceClaude Code
132GitHub
Skill ansehen
KI & Wissen

No visual example yet

Explore the skill

Future Agi

future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…

Preis unbestätigtKI & WissenVor Nutzung prüfen
1155GitHub
Skill ansehen
KI & Wissen

No visual example yet

Explore the skill

Langfuse

langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…

Preis unbestätigtKI & WissenOpenAI AgentsLangChainVor Nutzung prüfen
29.437GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…

Preis unbestätigtEntwicklungClaude Code
5312GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…

Preis unbestätigtEntwicklungClaude CodeOpenAI Agents
979GitHub
Skill ansehen

Anleitungen & Vergleiche

Für Entwickler