OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Finde den passenden Skill für deine nächste Aufgabe mit Codex, Claude Code, Cursor und mehr.

Ergebnisse · “evals”

13 Skills

Ergebnisse: 13

Entwicklung

No visual example yet

Explore the skill

Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…

Preis unbestätigtEntwicklungClaude CodeVor Nutzung prüfen
24.743GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.

Preis unbestätigtEntwicklungClaude Code
132GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…

Preis unbestätigtEntwicklungClaude Code
5312GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…

Preis unbestätigtEntwicklungClaude CodeOpenAI Agents
979GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Create, run, diagnose, and iteratively improve Agent Skill evaluations (evals) with the skill-up CLI / 使用 skill-up CLI 创建、运行、诊断并持续改进 Agent Skill 评测. Use when the user as…

Preis unbestätigtEntwicklungClaude Code
846GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

An Agent Skills skill for developing DeepSeek Harness (DSH) plugins(开发 DSH 插件的 Agent Skill)——插件/服务/事件/工具/LLM 适配器/打包安装的标准。Works with Claude Code, Codex, DSH, VS Code Copi…

Preis unbestätigtEntwicklungClaude CodeVor Nutzung prüfen
34GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Audits instrumentation health of existing Arize traces. Runs deterministic checks over a bounded span sample (orphaned/uncategorized/duplicate spans, flat structure, bla…

Preis unbestätigtEntwicklungClaude Code
51GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Sync the Logic-Lens working-copy skills/ into the installed plugin cache so content-evals test the EDITED skill, not the last published one. ALWAYS run this after editin…

Preis unbestätigtEntwicklungClaude Code
24GitHub
Skill ansehen
Entwicklung

No visual example yet

Explore the skill

Godmode

thiientv

Production-grade Agent Skills for AI coding agents—composable workflows for planning, TDD, debugging, review, UI/UX, releases, incidents, and evals.

Preis unbestätigtEntwicklungVor Nutzung prüfen
90GitHub
Skill ansehen

Anleitungen & Vergleiche

Für Entwickler