AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Live registry search

Results for “webserver-benchmarking

Exact name and slug matches are checked against the live registry before ranked alternatives.

Clear search

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-16 of 26 ranked candidates matching "webserver-benchmarking"

Best blend of relevance, quality, freshness, and verified outcomes

1

Cleverhans

VERIFIEDSTRONG · 73TRUST · 82SAFE · REVIEWEDCODING

An adversarial example library for constructing attacks, building defenses, and benchmarking both

$ npx skills add cleverhans-lab/cleverhans
6.4K stars53 quality82 trustReviewed with permission notes2y since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

jupyter-notebookmachine-learning
by cleverhans-labDetailsQuick view
2

Bencher

STRONG · 70TRUST · 75SAFE · REVIEWEDCODING

🐰 Bencher - Continuous Benchmarking

$ npx skills add bencherdev/bencher
855 stars51 quality75 trustReviewed with permission notes2mo since pushNeeds review

Scenario Coding agents

CLI + Codex · 4 targets

rustci/cd
by bencherdevDetailsQuick view
3

Lidar Slam Ros2

STRONG · 74TRUST · 81SAFE · REVIEWEDCODING

ROS 2 LiDAR SLAM for pointcloud-map authoring, benchmarking, and Autoware-compatible map workflows.

$ npx skills add rsasaki0109/lidar_slam_ros2
823 stars51 quality81 trustReviewed2mo since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

htmlrobotics
by rsasaki0109DetailsQuick view
4

Bocoel

PROMISING · 66TRUST · 77SAFE · REVIEWEDDESIGN

Bayesian Optimization as a Coverage Tool for Evaluating LLMs. Accurate evaluation (benchmarking) that's 10 times faster with just a few lines of modular code.

$ npx skills add rentruewang/bocoel
289 stars44 quality77 trustReviewed with permission notes3mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonmachine-learning
by rentruewangDetailsQuick view
5

Leaderboard

NEEDS REVIEW · 43TRUST · 71SAFE · EXPERIMENTALDESIGN

SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.

$ npx skills add SpeechColab/Leaderboard
545 stars38 quality71 trustExperimental1y since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by SpeechColabDetailsQuick view
6

Powerful Benchmarker

NEEDS REVIEW · 42TRUST · 65SAFE · BLOCKEDDATA

A library for ML benchmarking. It's powerful.

$ npx skills add KevinMusgrave/powerful-benchmarker
441 stars37 quality65 trustBlocked for auto-install3y since pushRisky

Scenario Data analysis

CLI + Codex · 4 targets

jupyter-notebookcomputer-vision
by KevinMusgraveDetailsQuick view
7

Place Recognition Evaluation

NEEDS REVIEW · 36TRUST · 67SAFE · BLOCKEDAUTOMATION

Benchmarking and evaluation framework for place recognition methods, featuring SuperPoint+SuperGlue, LoGG3D-Net, Scan Context, DBoW2, MixVPR, STD

$ npx skills add 4ku/Place-recognition-evaluation
126 stars33 quality67 trustBlocked for auto-install2y since pushRisky

Scenario Workflow automation

CLI + Codex · 4 targets

c++robotics
by 4kuDetailsQuick view
8

Agentops

VERIFIEDEXCELLENT · 97TRUST · 89SAFE · REVIEWEDCODING

Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK,…

$ npx skills add AgentOps-AI/agentops
5.8K stars65 quality89 trustReviewedClaude Code + OpenAI Agents2mo since pushSafe to try

Scenario Coding agents

Claude Code + OpenAI Agents · 4 targets

pythonllm
by AgentOps-AIDetailsQuick view
9

Evalscope

VERIFIEDEXCELLENT · 94TRUST · 88SAFE · REVIEWEDCODING

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

$ npx skills add modelscope/evalscope
3.0K stars63 quality88 trustReviewed2mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

pythonrag
by modelscopeDetailsQuick view
10

Docext

VERIFIEDEXCELLENT · 88TRUST · 85SAFE · REVIEWEDRESEARCH

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

$ npx skills add NanoNets/docext
2.0K stars58 quality85 trustReviewed5mo since pushSafe to try

Scenario Document processing

CLI + Codex · 4 targets

pythonocr
by NanoNetsDetailsQuick view
11

Video ChatGPT

VERIFIEDSTRONG · 72TRUST · 81SAFE · REVIEWEDDESIGN

[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrai…

$ npx skills add mbzuai-oryx/Video-ChatGPT
1.5K stars49 quality81 trustReviewed with permission notesOpenAI Agents1y since pushNeeds review

Scenario Design and creative

OpenAI Agents + CLI · 4 targets

pythonchatbot
by mbzuai-oryxDetailsQuick view
12

golang-benchmark

STRONG · 82TRUST · 75SAFE · REVIEWEDRESEARCH

Golang benchmarking, profiling, and performance measurement. Use when writing, running, or comparing Go benchmarks, profiling hot paths with pprof, interpreting CPU/memo…

$ npx skills add samber/cc-skills-golang --skill golang-benchmark
3.0K stars48 quality75 trustReviewed with permission notesClaude Code2d since pushNeeds review

Scenario Research agents

Claude Code + CLI · 4 targets

by samberDetailsQuick view
13

WindowsAgentArena

STRONG · 71TRUST · 79SAFE · REVIEWEDRESEARCH

Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.

$ npx skills add microsoft/WindowsAgentArena
867 stars47 quality79 trustReviewed with permission notes4mo since pushNeeds review

Scenario Research agents

CLI + Codex · 4 targets

pythoncomputer-use
by microsoftDetailsQuick view
14

VectorizedMultiAgentSimulator

STRONG · 75TRUST · 79SAFE · REVIEWEDCODING

VMAS is a vectorized differentiable simulator designed for efficient Multi-Agent Reinforcement Learning benchmarking. It is comprised of a vectorized 2D physics engine w…

$ npx skills add proroklab/VectorizedMultiAgentSimulator
588 stars46 quality79 trustReviewed with permission notes3mo since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

pythonmulti-agent
by proroklabDetailsQuick view
15

BenchMARL

PROMISING · 67TRUST · 78SAFE · REVIEWEDCODING

BenchMARL is a library for benchmarking Multi-Agent Reinforcement Learning (MARL). BenchMARL allows to quickly compare different MARL algorithms, tasks, and models while…

$ npx skills add facebookresearch/BenchMARL
636 stars42 quality78 trustReviewed with permission notes7mo since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

pythonmulti-agent
by facebookresearchDetailsQuick view
16

VINE

PROMISING · 62TRUST · 73SAFE · REVIEWEDDESIGN

[ICLR 2025] "Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances" (Official Implementation)

$ npx skills add Shilin-LU/VINE
388 stars45 quality73 trustReviewed with permission notes5mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonimage-generation
by Shilin-LUDetailsQuick view

Page 1

Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.

Try the agent resolve API