No visual example yet
Explore the skillRagaAI Catalyst
raga-ai-hub
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…
OPENAGENTSKILL / DIRECTORY
다음 작업에 맞는 스킬을 찾아보세요. Codex, Claude Code, Cursor 등을 지원합니다.
11 Skills
검색 결과: 11
No visual example yet
Explore the skillraga-ai-hub
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…
No visual example yet
Explore the skillGiskard-AI
🐢 Open-Source Evaluation & Testing library for LLM Agents
No visual example yet
Explore the skillcoze-dev
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…
No visual example yet
Explore the skillablab
No visual example yet
Explore the skillArize-ai
AI Observability & Evaluation
No visual example yet
Explore the skillGoogleCloudPlatform
Ship AI Agents to Google Cloud in minutes, not months. Production-ready templates with built-in CI/CD, evaluation, and observability.
No visual example yet
Explore the skilldotnet
Scaffolds new agent skills for the dotnet/skills repository. Use when creating a new skill, generating SKILL.md files, writing a skill description that the runtime will…
No visual example yet
Explore the skilldotnet
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and ru…
No visual example yet
Explore the skillalibaba
A CLI evaluation framework to make your Agent Skill Up.
No visual example yet
Explore the skillprobabl-ai
Track your Data Science. Skore's open-source Python library accelerates ML model development with automated evaluation reports, smart methodological guidance, and compre…
No visual example yet
Explore the skillsangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…