AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Live registry search

Results for “evaluation

Exact name and slug matches are checked against the live registry before ranked alternatives.

Clear search

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-16 of 78 ranked candidates matching "evaluation"

Best blend of relevance, quality, freshness, and verified outcomes

1

Cloud Map Evaluation

PROMISING · 63TRUST · 73SAFE · REVIEWEDCODING

[RAL' 25 & IROS‘ 25] MapEval: Towards Unified, Robust and Efficient SLAM Map Evaluation Framework.

$ npx skills add JokerJohn/Cloud_Map_Evaluation
473 stars45 quality73 trustReviewed with permission notes5mo since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

c++robotics
by JokerJohnDetailsQuick view
2

Place Recognition Evaluation

NEEDS REVIEW · 36TRUST · 67SAFE · BLOCKEDAUTOMATION

Benchmarking and evaluation framework for place recognition methods, featuring SuperPoint+SuperGlue, LoGG3D-Net, Scan Context, DBoW2, MixVPR, STD

$ npx skills add 4ku/Place-recognition-evaluation
126 stars33 quality67 trustBlocked for auto-install2y since pushRisky

Scenario Workflow automation

CLI + Codex · 4 targets

c++robotics
by 4kuDetailsQuick view
3

VPR Methods Evaluation

PROMISING · 64TRUST · 77SAFE · REVIEWEDDESIGN

Wrapper for 10+ VPR models. Use any SOTA VPR model just by changing one parameter

$ npx skills add gmberton/VPR-methods-evaluation
197 stars43 quality77 trustReviewed with permission notes4mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythoncomputer-vision
by gmbertonDetailsQuick view
4

Evaluation Guidebook

VERIFIEDSTRONG · 76TRUST · 81SAFE · REVIEWEDDESIGN

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

$ npx skills add huggingface/evaluation-guidebook
2.1K stars54 quality81 trustReviewed with permission notes9mo since pushNeeds review

Scenario Design and creative

CLI + Codex · 4 targets

jupyter-notebookmachine-learning
by huggingfaceDetailsQuick view
5

VLMEvalKit

VERIFIEDEXCELLENT · 93TRUST · 88SAFE · REVIEWEDRESEARCH

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

$ npx skills add open-compass/VLMEvalKit
1 agent calls0% success4.2K stars69 quality88 trustReviewedClaude Code + OpenAI Agents2mo since pushSafe to try

Scenario Research agents

Claude Code + OpenAI Agents · 4 targets

pythoncomputer-vision
by open-compassDetailsQuick view
6

RagaAI Catalyst

VERIFIEDEXCELLENT · 97TRUST · 87SAFE · REVIEWEDCODING

Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…

$ npx skills add raga-ai-hub/RagaAI-Catalyst
16.1K stars62 quality87 trustReviewed6mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

pythonllmops
by raga-ai-hubDetailsQuick view
7

Agenta

VERIFIEDEXCELLENT · 97TRUST · 86SAFE · REVIEWEDCODING

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

$ npx skills add Agenta-AI/agenta
4.5K stars67 quality86 trustReviewed13d since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

typescriptllmops
by Agenta-AIDetailsQuick view
8

Giskard Oss

VERIFIEDEXCELLENT · 97TRUST · 87SAFE · REVIEWEDCODING

🐢 Open-Source Evaluation & Testing library for LLM Agents

$ npx skills add Giskard-AI/giskard-oss
5.7K stars65 quality87 trustReviewed1mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

pythonllmops
by Giskard-AIDetailsQuick view
9

Coze Loop

VERIFIEDEXCELLENT · 97TRUST · 89SAFE · REVIEWEDCODING

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…

$ npx skills add coze-dev/coze-loop
5.6K stars65 quality89 trustReviewed1mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

gollmops
by coze-devDetailsQuick view
10

AutoRAG

VERIFIEDEXCELLENT · 96TRUST · 87SAFE · REVIEWEDRESEARCH

AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation

$ npx skills add Marker-Inc-Korea/AutoRAG
4.8K stars64 quality87 trustReviewed2mo since pushSafe to try

Scenario RAG and knowledge

CLI + Codex · 4 targets

pythonrag
by Marker-Inc-KoreaDetailsQuick view
11

Trulens

VERIFIEDEXCELLENT · 95TRUST · 86SAFE · REVIEWEDCODING

Evaluation and Tracking for LLM Experiments and AI Agents

$ npx skills add truera/trulens
3.4K stars63 quality86 trustReviewed2mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

pythonllmops
by trueraDetailsQuick view
12

Evalscope

VERIFIEDEXCELLENT · 94TRUST · 88SAFE · REVIEWEDCODING

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

$ npx skills add modelscope/evalscope
3.0K stars63 quality88 trustReviewed2mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

pythonrag
by modelscopeDetailsQuick view
13

FinSight AI

VERIFIEDEXCELLENT · 90TRUST · 87SAFE · REVIEWEDRESEARCH

AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.

$ npx skills add juanjuandog/FinSight-AI
1.2K stars60 quality87 trustReviewed with permission notes3mo since pushNeeds review

Scenario RAG and knowledge

CLI + Codex · 4 targets

javarag
by juanjuandogDetailsQuick view
14

AB3DMOT

VERIFIEDPROMISING · 62TRUST · 76SAFE · EXPERIMENTALDESIGN

(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"

$ npx skills add xinshuoweng/AB3DMOT
1.8K stars50 quality76 trustExperimental2y since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonmachine-learning
by xinshuowengDetailsQuick view
15

Skills

VERIFIEDEXCELLENT · 90TRUST · 85SAFE · REVIEWEDCODING

Agent Skills for Langfuse, the open source LLM engineering platform for tracing, prompt management, and evaluation

$ npx skills add langfuse/skills
214 stars59 quality85 trustReviewedClaude Code30d since pushSafe to try

Scenario Coding agents

Claude Code + CLI · 4 targets

python
by langfuseDetailsQuick view
16

OpenJudge

STRONG · 80TRUST · 80SAFE · REVIEWEDRESEARCH

OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

$ npx skills add agentscope-ai/OpenJudge
775 stars54 quality80 trustReviewedClaude Code19d since pushSafe to try

Scenario RAG and knowledge

Claude Code + CLI · 4 targets

pythonai-agents
by agentscope-aiDetailsQuick view

Page 1

Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.

Try the agent resolve API