No visual example yet
Explore the skillanthropic-evaluations
dwmkerr
This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also su…
OPENAGENTSKILL / DIRECTORY
Trouvez un skill pour votre prochaine tâche avec Codex, Claude Code, Cursor et plus encore.
23 Skills
Résultats: 23
No visual example yet
Explore the skilldwmkerr
This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also su…
No visual example yet
Explore the skillScale3-Labs
Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular…
No visual example yet
Explore the skillcomet-ml
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
No visual example yet
Explore the skilllangwatch
The platform for LLM evaluations and AI agent testing
No visual example yet
Explore the skillopenlit
Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Inte…
No visual example yet
Explore the skillalibaba
Create, run, diagnose, and iteratively improve Agent Skill evaluations (evals) with the skill-up CLI / 使用 skill-up CLI 创建、运行、诊断并持续改进 Agent Skill 评测. Use when the user as…
No visual example yet
Explore the skillAzure-Samples
This sample has the full End2End process of creating RAG application with Prompty and Azure AI Foundry. It includes GPT-4 LLM application code, evaluations, deployment a…
No visual example yet
Explore the skillmgechev
Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple graders wit…
No visual example yet
Explore the skillAnacletoLAB
🍇 GRAPE is a Rust/Python Graph Representation Learning library for Predictions and Evaluations
No visual example yet
Explore the skillAbdelStark
Discover and compare Jev and TypeSafe SDKs, integrations, applications, and evaluations. Use when selecting public resources for a typed-decision workflow or checking wh…
No visual example yet
Explore the skilltikalk
Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.
No visual example yet
Explore the skillsoba-labs
Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, run…
No visual example yet
Explore the skillthiientv
Designs and runs reproducible evaluations for AI agents, prompts, tools, skills, and model-backed workflows using realistic datasets, isolated baselines, object
No visual example yet
Explore the skillmagnus919
Design, run, review, or release framework- and vendor-neutral evaluations and observability for AI agents. Use when defining agent evals, datasets, graders, tra
No visual example yet
Explore the skillcuellarfr
Evaluate UI designs against usability heuristics, UX laws, interaction patterns, interaction design principles, information architecture, and content quality. Conduct he…
No visual example yet
Explore the skillYujxZJCN
Evidence-honest teaching reflection for university professors. 6-agent team covering student-evaluation analysis (thematic, bias-caveated), mid-semester feedback, peer-o…