OPENAGENTSKILL / DIRECTORY

AI Agent Skills

Trouvez un skill pour votre prochaine tâche avec Codex, Claude Code, Cursor et plus encore.

Résultats · “evaluations”

23 Skills

Résultats: 23

Recherche

No visual example yet

Explore the skill

This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also su…

Prix non confirméRechercheClaude Code
23GitHub
IA et connaissances

No visual example yet

Explore the skill

Langtrace

Scale3-Labs

Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular…

Prix non confirméIA et connaissancesOpenAI AgentsLangChainÀ vérifier avant usage
1,2 kGitHub
Développement

No visual example yet

Explore the skill

Opik

comet-ml

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

Prix non confirméDéveloppementLangChainÀ vérifier avant usage
19,8 kGitHub
IA et connaissances

No visual example yet

Explore the skill

Openlit

openlit

Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Inte…

Prix non confirméIA et connaissancesLangChainÀ vérifier avant usage
2,5 kGitHub
Développement

No visual example yet

Explore the skill

Create, run, diagnose, and iteratively improve Agent Skill evaluations (evals) with the skill-up CLI / 使用 skill-up CLI 创建、运行、诊断并持续改进 Agent Skill 评测. Use when the user as…

Prix non confirméDéveloppementClaude Code
846GitHub
Développement

No visual example yet

Explore the skill

Contoso Chat

Azure-Samples

This sample has the full End2End process of creating RAG application with Prompty and Azure AI Foundry. It includes GPT-4 LLM application code, evaluations, deployment a…

Prix non confirméDéveloppementOpenAI AgentsÀ vérifier avant usage
762GitHub
Développement

No visual example yet

Explore the skill

Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple graders wit…

Prix non confirméDéveloppementClaude Code
696GitHub
Autres skills

No visual example yet

Explore the skill

awesome-jev

AbdelStark

Discover and compare Jev and TypeSafe SDKs, integrations, applications, and evaluations. Use when selecting public resources for a typed-decision workflow or checking wh…

Prix non confirméAutres skillsClaude Code
566GitHub
Droit et conformité

No visual example yet

Explore the skill

Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.

Prix non confirméDroit et conformitéClaude Code
132GitHub
Navigateur et automatisation

No visual example yet

Explore the skill

Designs and runs reproducible evaluations for AI agents, prompts, tools, skills, and model-backed workflows using realistic datasets, isolated baselines, object

Prix non confirméNavigateur et automatisationClaude Code
94GitHub
Design et UI

No visual example yet

Explore the skill

Evaluate UI designs against usability heuristics, UX laws, interaction patterns, interaction design principles, information architecture, and content quality. Conduct he…

Prix non confirméDesign et UIClaude Code
54GitHub
Éducation

No visual example yet

Explore the skill

Evidence-honest teaching reflection for university professors. 6-agent team covering student-evaluation analysis (thematic, bias-caveated), mid-semester feedback, peer-o…

Prix non confirméÉducationClaude Code
34GitHub

Guides et comparaisons

Pour les développeurs