No visual example yet
Explore the skillagent-evals-and-observability
magnus919
>-
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
33 Skills
Results: 33
No visual example yet
Explore the skillmagnus919
>-
No visual example yet
Explore the skillpromptfoo
Write, refine, run, and QA non-redteam promptfoo eval suites after the target or provider already works: prompts, vars, test cases, assertions, model-graded rubrics, tra…
No visual example yet
Explore the skilltikalk
Refine, cluster, and accept draft criteria into the published goldset. Isolates 20% holdout split and publishes goldset.md + goldset.json.
No visual example yet
Explore the skilltikalk
Generate executable graders and configs from goldset. Generates Python graders / metrics and auto-runs unit tests to verify grader correctness.
No visual example yet
Explore the skilltikalk
Extract eval criteria from product specs and production failure traces (bottom-up error analysis). Writes proposed criteria to .adlc/drafts/evals/.
No visual example yet
Explore the skilltikalk
Analyze evaluation results and close the loop. Specification failures create local CDRs to fix agent rules; generalization failures go to evaluator backlog.
No visual example yet
Explore the skilltikalk
Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate se…
No visual example yet
Explore the skilltikalk
Run evaluations and validate evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Executes PromptFoo or pytest DeepEval.
No visual example yet
Explore the skillvasilyu1983
Designs coding-agent observability and evals. Use when measuring traces, replay, checkpoint lineage, quality trajectories, tool grading, regression, or cost.
No visual example yet
Explore the skillfuture-agi
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails.…
No visual example yet
Explore the skillGitHamza0206
OpenSource Production ready Customer service with built in Evals and monitoring
No visual example yet
Explore the skilllangfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK,…
No visual example yet
Explore the skillyannickYamo
Build AI-native products with agency-control tradeoffs, calibration loops, and eval strategies. Use when building AI agents, LLM features, or products where AI handles u…
No visual example yet
Explore the skillombharatiya
AI system design guide for engineers building production AI systems and evals.
No visual example yet
Explore the skilldotnet
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…
No visual example yet
Explore the skillmicrosoft
Prepare and publish a new version of the waza azd extension. USE FOR: "publish extension", "release new version", "bump version", "prepare release", "update changelog",…