OPENAGENTSKILL / DIRECTORY

AI Agent Skills

다음 작업에 맞는 스킬을 찾아보세요. Codex, Claude Code, Cursor 등을 지원합니다.

검색 결과 · “evals”

34 Skills

검색 결과: 34

검색 및 연구

No visual example yet

Explore the skill

Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.

가격 미확인검색 및 연구Claude Code
132GitHub
보안

No visual example yet

Explore the skill

Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…

가격 미확인보안Claude Code
132GitHub
개발 및 테스트

No visual example yet

Explore the skill

Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…

가격 미확인개발 및 테스트Claude Code사용 전 검토
2.5만GitHub
검색 및 연구

No visual example yet

Explore the skill

Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.

가격 미확인검색 및 연구Claude Code
132GitHub
AI 및 지식

No visual example yet

Explore the skill

Future Agi

future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…

가격 미확인AI 및 지식사용 전 검토
1.2천GitHub
AI 및 지식

No visual example yet

Explore the skill

Langfuse

langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…

가격 미확인AI 및 지식OpenAI AgentsLangChain사용 전 검토
2.9만GitHub
개발 및 테스트

No visual example yet

Explore the skill

Creates, optimizes, and iteratively refines agent prompts, system prompts, developer prompts, and reusable prompt templates. Use when asked to improve a prompt, optimize…

가격 미확인개발 및 테스트Claude CodeOpenAI Agents
979GitHub

가이드 및 비교

개발자용