anthropic-evaluations
dwmkerr
This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also su…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Search retrieves candidates across the registry and ranks a bounded shortlist by task fit. This count is matching candidates, not the registry total. No suitable match? Try a specific tool or task.
1–16 / 18
Results: 18
dwmkerr
This skill should be used when the user asks to "create evals", "evaluate an agent", "build evaluation suite", or mentions agent testing, graders, or benchmarks. Also su…
Scale3-Labs
Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular…
comet-ml
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
openlit
Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Inte…
alibaba
Create, run, diagnose, and iteratively improve Agent Skill evaluations (evals) with the skill-up CLI / 使用 skill-up CLI 创建、运行、诊断并持续改进 Agent Skill 评测. Use when the user as…
Azure-Samples
This sample has the full End2End process of creating RAG application with Prompty and Azure AI Foundry. It includes GPT-4 LLM application code, evaluations, deployment a…
mgechev
Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple graders wit…
AnacletoLAB
🍇 GRAPE is a Rust/Python Graph Representation Learning library for Predictions and Evaluations
tikalk
Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.
soba-labs
Use this skill when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, run…
cuellarfr
Evaluate UI designs against usability heuristics, UX laws, interaction patterns, interaction design principles, information architecture, and content quality. Conduct he…
eclipse-langium
Guide for using the langium-ai (LAI) CLI to generate language descriptors, synthesize system prompts, run evaluations, and iteratively refine AI-powered tooling in Langi…
danicat
Comprehensive guide to develop and improve Agent Skill performance. Contains best practices for skill formatting (frontmatter and metadata), naming, descriptions, fine-t…
agentagon
Assess agent code, changes, traces and evaluations; report findings and prepare missing evals only with confirmed authorization.
agentagon
Onboard an agent application by recommending a useful first quality check, reusing or preparing evaluations, and measuring a baseline within agreed limits.
Owner-curated external sources. Not filtered by the scores or compatibility controls above; excluded from GitHub rankings and automatic installation.
No external entries match this query.