AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Live registry search

Results for “leaderboard

Exact name and slug matches are checked against the live registry before ranked alternatives.

Clear search

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-10 of 10 ranked candidates matching "leaderboard"

Best blend of relevance, quality, freshness, and verified outcomes

1

Leaderboard

NEEDS REVIEW · 43TRUST · 71SAFE · EXPERIMENTALDESIGN

SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.

$ npx skills add SpeechColab/Leaderboard
545 stars38 quality71 trustExperimental1y since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by SpeechColabDetailsQuick view
2

Llm Leaderboard

NEEDS REVIEW · 45TRUST · 70SAFE · EXPERIMENTALDESIGN

A joint community effort to create one central leaderboard for LLMs.

$ npx skills add LudwigStumpp/llm-leaderboard
306 stars36 quality70 trustExperimental2y since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonmachine-learning
by LudwigStumppDetailsQuick view
3

Tokscale

VERIFIEDEXCELLENT · 95TRUST · 82SAFE · EXPERIMENTALCODING

🛰️ A CLI tool for tracking token usage from OpenCode, Claude Code, 🦞OpenClaw (Clawdbot/Moltbot), Pi, Codex, Gemini, Cursor, AmpCode, Factory Droid, Kimi, and more! • �…

$ npx skills add junhoyeo/tokscale
3.9K stars64 quality82 trustExperimentalClaude Code + OpenAI Agents2mo since pushNeeds review

Scenario Coding agents

Claude Code + OpenAI Agents · 4 targets

rustclaude-code
by junhoyeoDetailsQuick view
4

All Agentic Architectures

VERIFIEDEXCELLENT · 95TRUST · 88SAFE · REVIEWEDCODING

35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider…

$ npx skills add FareedKhan-dev/all-agentic-architectures
3.7K stars64 quality88 trustReviewedOpenAI Agents + LangChain2mo since pushSafe to try

Scenario GitHub automation

OpenAI Agents + LangChain · 4 targets

jupyter-notebookai-agents
by FareedKhan-devDetailsQuick view
5

Checkup

STRONG · 72TRUST · 75SAFE · EXPERIMENTALCODING

AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title,…

$ npx skills add agentvitals/checkup
95 stars49 quality75 trustExperimentalClaude Code6d since pushNeeds review

Scenario Coding agents

Claude Code + CLI · 4 targets

shell
by agentvitalsDetailsQuick view
6

Evaluation Guidebook

VERIFIEDSTRONG · 76TRUST · 81SAFE · REVIEWEDDESIGN

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

$ npx skills add huggingface/evaluation-guidebook
2.1K stars54 quality81 trustReviewed with permission notes9mo since pushNeeds review

Scenario Design and creative

CLI + Codex · 4 targets

jupyter-notebookmachine-learning
by huggingfaceDetailsQuick view
7

AnglE

PROMISING · 69TRUST · 79SAFE · REVIEWEDRESEARCH

Train and Infer Powerful Sentence Embeddings with AnglE | 🔥 SOTA on STS and MTEB Leaderboard

$ npx skills add SeanLee97/AnglE
571 stars46 quality79 trustReviewed with permission notes5mo since pushNeeds review

Scenario RAG and knowledge

CLI + Codex · 4 targets

pythonrag
by SeanLee97DetailsQuick view
8

Awesome LLM Eval

PROMISING · 61TRUST · 76SAFE · REVIEWEDRESEARCH

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表…

$ npx skills add onejune2018/Awesome-LLM-Eval
642 stars42 quality76 trustReviewed with permission notesOpenAI Agents9mo since pushNeeds review

Scenario RAG and knowledge

OpenAI Agents + CLI · 4 targets

rag
by onejune2018DetailsQuick view
9

FLamby

NEEDS REVIEW · 44TRUST · 70SAFE · EXPERIMENTALCODING

Cross-silo Federated Learning playground in Python. Discover 7 real-world federated datasets to test your new FL strategies and try to beat the leaderboard.

$ npx skills add owkin/FLamby
239 stars35 quality70 trustExperimental2y since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

pythonhealth-data
by owkinDetailsQuick view
10

Viberank

PROMISING · 65TRUST · 76SAFE · EXPERIMENTALCODING

🏆 The AI coding usage leaderboard — Claude Code, Codex, Gemini CLI & more. Real costs and tokens from ccusage data. Submit with: npx viberank-cli

$ npx skills add sculptdotfun/viberank
102 stars45 quality76 trustExperimentalClaude Code + OpenAI Agents2mo since pushNeeds review

Scenario Coding agents

Claude Code + OpenAI Agents · 4 targets

typescriptdeveloper-tools
by sculptdotfunDetailsQuick view