Skill ディレクトリ

AI Agent のための再利用可能な Skill を見つける。

タスクで実際の GitHub Skill を検索し、利用前に Stars、Trust、監査、カテゴリ、インストール経路を確認できます。

すべての推奨は、リポジトリ、監査、インストール経路に明確につながっています。

検索結果: agent-evaluation

英語版ディレクトリ

A comprehensive set of 38 marketing skills and 5 commands for Claude Code covering SEO/GEO and influencer marketing with evaluation frameworks.

2.6K
Stars
86/100
信頼
カテゴリ: productivity監査

A Codex skill for generating minimal zine-style editorial poster prompts and images.

6.3K
Stars
83/100
信頼
カテゴリ: design-creative監査

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

4.2K
Stars
85/100
信頼
カテゴリ: robotics-iot監査

BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI workflow, RAG, Agent, Unified model management, Evaluation, SFT, Dataset Management, Enterprise-level System Management, Observability and more.

11K
Stars
87/100
信頼
カテゴリ: document-processing監査

AI Observability & Evaluation

10K
Stars
76/100
信頼
カテゴリ: development監査

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

4.5K
Stars
78/100
信頼
カテゴリ: development監査

28 eval-informed mental models and critical-thinking skills for Claude Code, GitHub Copilot, Codex, Cursor, and other Agent Skills-compatible tools

941
Stars
84/100
信頼
カテゴリ: utility監査

Ship AI Agents to Google Cloud in minutes, not months. Production-ready templates with built-in CI/CD, evaluation, and observability.

6.5K
Stars
86/100
信頼
カテゴリ: development監査

🐢 Open-Source Evaluation & Testing library for LLM Agents

5.7K
Stars
81/100
信頼
カテゴリ: development監査

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.

5.6K
Stars
86/100
信頼
カテゴリ: development監査

AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation

4.8K
Stars
84/100
信頼
カテゴリ: data監査

Evaluation and Tracking for LLM Experiments and AI Agents

3.4K
Stars
80/100
信頼
カテゴリ: development監査