AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Live registry search

Results for “performance-evaluation

Exact name and slug matches are checked against the live registry before ranked alternatives.

Clear search

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-16 of 111 ranked candidates matching "performance-evaluation"

Best blend of relevance, quality, freshness, and verified outcomes

1

RagaAI Catalyst

VERIFIEDEXCELLENT · 97TRUST · 87SAFE · REVIEWEDCODING

Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hos…

$ npx skills add raga-ai-hub/RagaAI-Catalyst
16.1K stars62 quality87 trustReviewed6mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

pythonllmops
by raga-ai-hubDetailsQuick view
2

Evalscope

VERIFIEDEXCELLENT · 94TRUST · 88SAFE · REVIEWEDCODING

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

$ npx skills add modelscope/evalscope
3.0K stars63 quality88 trustReviewed2mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

pythonrag
by modelscopeDetailsQuick view
3

Promptfoo

VERIFIEDEXCELLENT · 100TRUST · 89SAFE · REVIEWEDCODING

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declara…

$ npx skills add promptfoo/promptfoo
22.2K stars71 quality89 trustReviewed with permission notesClaude Code + OpenAI Agents2mo since pushSafe to try

Scenario Testing and QA

Claude Code + OpenAI Agents · 4 targets

typescriptrag
by promptfooDetailsQuick view
4

Hdbscan

VERIFIEDEXCELLENT · 94TRUST · 86SAFE · REVIEWEDCODING

A high performance implementation of HDBSCAN clustering.

$ npx skills add scikit-learn-contrib/hdbscan
3.1K stars63 quality86 trustReviewed2mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

jupyter-notebookmachine-learning
by scikit-learn-contribDetailsQuick view
5

Vercel React Best Practices

VERIFIEDEXCELLENT · 100TRUST · 91SAFE · REVIEWEDCODING

React and Next.js performance guidance for writing, reviewing, and refactoring production UI code.

$ npx skills add vercel-labs/agent-skills --skill vercel-react-best-practices
30.3K stars75 quality91 trustReviewedClaude Code + OpenAI AgentsPushed todaySafe to try

Scenario Coding agents

Claude Code + OpenAI Agents · 4 targets

codexclaude-codecursorreactnext.js
by vercel-labsDetailsQuick view
6

Hyperswitch

VERIFIEDEXCELLENT · 100TRUST · 90SAFE · REVIEWEDFINANCE

Open source, composable payments platform | PCI compliant | SaaS and Self-host options | Enables connectivity to multiple payment, payout, fraud, vault and tokenization…

$ npx skills add juspay/hyperswitch
43.2K stars73 quality90 trustReviewed with permission notes2mo since pushNeeds review

Scenario Finance and quant

CLI + Codex · 4 targets

rustfinance
by juspayDetailsQuick view
7

Lighthouse

VERIFIEDEXCELLENT · 100TRUST · 88SAFE · REVIEWEDCODING

Automated auditing, performance metrics, and best practices for the web.

$ npx skills add GoogleChrome/lighthouse
30.4K stars72 quality88 trustReviewed2mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

javascriptdeveloper-tools
by GoogleChromeDetailsQuick view
8

VLMEvalKit

VERIFIEDEXCELLENT · 93TRUST · 88SAFE · REVIEWEDRESEARCH

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

$ npx skills add open-compass/VLMEvalKit
1 agent calls0% success4.2K stars69 quality88 trustReviewedClaude Code + OpenAI Agents2mo since pushSafe to try

Scenario Research agents

Claude Code + OpenAI Agents · 4 targets

pythoncomputer-vision
by open-compassDetailsQuick view
9

Agenta

VERIFIEDEXCELLENT · 97TRUST · 86SAFE · REVIEWEDCODING

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

$ npx skills add Agenta-AI/agenta
4.5K stars67 quality86 trustReviewed14d since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

typescriptllmops
by Agenta-AIDetailsQuick view
10

Giskard Oss

VERIFIEDEXCELLENT · 97TRUST · 87SAFE · REVIEWEDCODING

🐢 Open-Source Evaluation & Testing library for LLM Agents

$ npx skills add Giskard-AI/giskard-oss
5.7K stars65 quality87 trustReviewed1mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

pythonllmops
by Giskard-AIDetailsQuick view
11

Coze Loop

VERIFIEDEXCELLENT · 97TRUST · 89SAFE · REVIEWEDCODING

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from developmen…

$ npx skills add coze-dev/coze-loop
5.6K stars65 quality89 trustReviewed1mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

gollmops
by coze-devDetailsQuick view
12

AutoRAG

VERIFIEDEXCELLENT · 96TRUST · 87SAFE · REVIEWEDRESEARCH

AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation

$ npx skills add Marker-Inc-Korea/AutoRAG
4.8K stars64 quality87 trustReviewed2mo since pushSafe to try

Scenario RAG and knowledge

CLI + Codex · 4 targets

pythonrag
by Marker-Inc-KoreaDetailsQuick view
13

Trulens

VERIFIEDEXCELLENT · 95TRUST · 86SAFE · REVIEWEDCODING

Evaluation and Tracking for LLM Experiments and AI Agents

$ npx skills add truera/trulens
3.4K stars63 quality86 trustReviewed2mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

pythonllmops
by trueraDetailsQuick view
14

Tf Quant Finance

VERIFIEDEXCELLENT · 85TRUST · 84SAFE · REVIEWEDFINANCE

High-performance TensorFlow library for quantitative finance.

$ npx skills add google/tf-quant-finance
5.4K stars57 quality84 trustReviewed with permission notes6mo since pushNeeds review

Scenario Finance and quant

CLI + Codex · 4 targets

pythonfinance
by googleDetailsQuick view
15

Winscript

VERIFIEDEXCELLENT · 93TRUST · 88SAFE · REVIEWEDLEGAL

Open-source tool to build your Windows script from scratch. It includes debloat, privacy, performance & app installing scripts.

$ npx skills add flick9000/winscript
2.5K stars63 quality88 trustReviewed2mo since pushSafe to try

Scenario Legal and compliance

CLI + Codex · 4 targets

cssprivacy
by flick9000DetailsQuick view
16

Weld

VERIFIEDEXCELLENT · 90TRUST · 84SAFE · REVIEWEDDATA

High-performance runtime for data analytics applications

$ npx skills add weld-project/weld
3.0K stars59 quality84 trustReviewed4mo since pushSafe to try

Scenario Data analysis

CLI + Codex · 4 targets

rustmachine-learning
by weld-projectDetailsQuick view

Page 1

Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.

Try the agent resolve API