AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Live registry search

Results for “benchmark

Exact name and slug matches are checked against the live registry before ranked alternatives.

Clear search

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-16 of 80 ranked candidates matching "benchmark"

Best blend of relevance, quality, freshness, and verified outcomes

1

DynamicMap Benchmark

PROMISING · 60TRUST · 75SAFE · REVIEWEDCODING

The First Dynamic Map Removal Benchmark | Included 8 SOTA methods | Continous updating

$ npx skills add KTH-RPL/DynamicMap_Benchmark
425 stars41 quality75 trustReviewed with permission notes11mo since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

jupyter-notebookrobotics
by KTH-RPLDetailsQuick view
2

Deep Visual Geo Localization Benchmark

PROMISING · 65TRUST · 77SAFE · REVIEWEDCODING

Official code for CVPR 2022 (Oral) paper "Deep Visual Geo-localization Benchmark"

$ npx skills add gmberton/deep-visual-geo-localization-benchmark
256 stars44 quality77 trustReviewed with permission notes5mo since pushNeeds review

Scenario Coding agents

CLI + Codex · 4 targets

pythoncomputer-vision
by gmbertonDetailsQuick view
3

Speech To Text Benchmark

STRONG · 70TRUST · 77SAFE · REVIEWEDDESIGN

speech to text benchmark framework

$ npx skills add Picovoice/speech-to-text-benchmark
693 stars47 quality77 trustReviewed with permission notes5mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

pythonspeech
by PicovoiceDetailsQuick view
4

Browsers Benchmark

STRONG · 70TRUST · 77SAFE · REVIEWEDCODING

Browser automation engine benchmark - Test bypass rates, performance & stealth against Cloudflare, DataDome, reCAPTCHA, Kasada, Imperva, Akamai, PerimeterX and other bot…

$ npx skills add techinz/browsers-benchmark
335 stars48 quality77 trustReviewed with permission notesBrowser agents3mo since pushNeeds review

Scenario Testing and QA

Browser agents + CLI · 4 targets

pythonplaywright
by techinzDetailsQuick view
5

Roboflow 100 Benchmark

NEEDS REVIEW · 45TRUST · 72SAFE · EXPERIMENTALDESIGN

Code for replicating Roboflow 100 benchmark results and programmatically downloading benchmark datasets

$ npx skills add roboflow/roboflow-100-benchmark
299 stars36 quality72 trustExperimental2y since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

jupyter-notebookmachine-learning
by roboflowDetailsQuick view
6

Street Tryon Benchmark

NEEDS REVIEW · 37TRUST · 69SAFE · EXPERIMENTALDESIGN

[WACV'25] StreetTryOn: A Benchmark for In-the-Wild Virtual Try-On and Cross-Domain Virtual Try-On

$ npx skills add cuiaiyu/street-tryon-benchmark
159 stars34 quality69 trustExperimental2y since pushNeeds review

Scenario Design and creative

CLI + Codex · 4 targets

jupyter-notebookimage-generation
by cuiaiyuDetailsQuick view
7

BLINK Benchmark

PROMISING · 55TRUST · 74SAFE · EXPERIMENTALCODING

This repo contains evaluation code for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive". https://arxiv.org/abs/2404.12390 [ECCV 2024]

$ npx skills add zeyofu/BLINK_Benchmark
169 stars38 quality74 trustExperimental11mo since pushNeeds review

Scenario Coding agents

CLI + Codex · 4 targets

pythoncomputer-vision
by zeyofuDetailsQuick view
8

Chinese Llm Benchmark

VERIFIEDEXCELLENT · 98TRUST · 85SAFE · REVIEWEDCODING

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…

$ npx skills add jeinlee1991/chinese-llm-benchmark
6.2K stars65 quality85 trustReviewedClaude Code + OpenAI Agents3mo since pushSafe to try

Scenario GitHub automation

Claude Code + OpenAI Agents · 4 targets

llm
by jeinlee1991DetailsQuick view
9

Deep Text Recognition Benchmark

VERIFIEDSTRONG · 70TRUST · 80SAFE · EXPERIMENTALRESEARCH

Text recognition (optical character recognition) with deep learning methods, ICCV 2019

$ npx skills add clovaai/deep-text-recognition-benchmark
3.9K stars52 quality80 trustExperimental2y since pushNeeds review

Scenario Document processing

CLI + Codex · 4 targets

jupyter-notebookocr
by clovaaiDetailsQuick view
10

Caliper Benchmarks

NEEDS REVIEW · 41TRUST · 73SAFE · EXPERIMENTALCODING

Sample benchmark files for Hyperledger Caliper https://wiki.hyperledger.org/display/caliper

$ npx skills add hyperledger-caliper/caliper-benchmarks
117 stars33 quality73 trustExperimental1y since pushNeeds review

Scenario GitHub automation

CLI + Codex · 4 targets

javascriptblockchain
by hyperledger-caliperDetailsQuick view
11

golang-benchmark

STRONG · 82TRUST · 75SAFE · REVIEWEDRESEARCH

Golang benchmarking, profiling, and performance measurement. Use when writing, running, or comparing Go benchmarks, profiling hot paths with pprof, interpreting CPU/memo…

$ npx skills add samber/cc-skills-golang --skill golang-benchmark
3.0K stars48 quality75 trustReviewed with permission notesClaude Code3d since pushNeeds review

Scenario Research agents

Claude Code + CLI · 4 targets

by samberDetailsQuick view
12

Mteb

VERIFIEDEXCELLENT · 95TRUST · 86SAFE · REVIEWEDRESEARCH

MTEB: Massive Text Embedding Benchmark

$ npx skills add embeddings-benchmark/mteb
3.3K stars63 quality86 trustReviewed2mo since pushSafe to try

Scenario RAG and knowledge

CLI + Codex · 4 targets

pythonsemantic-search
by embeddings-benchmarkDetailsQuick view
13

MMMU

PROMISING · 61TRUST · 77SAFE · REVIEWEDCODING

This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"

$ npx skills add MMMU-Benchmark/MMMU
578 stars42 quality77 trustReviewed with permission notes6mo since pushNeeds review

Scenario Coding agents

CLI + Codex · 4 targets

pythoncomputer-vision
by MMMU-BenchmarkDetailsQuick view
14

Kube Bench

VERIFIEDEXCELLENT · 99TRUST · 89SAFE · REVIEWEDCODING

Checks whether Kubernetes is deployed according to security best practices as defined in the CIS Kubernetes Benchmark

$ npx skills add aquasecurity/kube-bench
8.1K stars66 quality89 trustReviewed1mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

gokubernetes
by aquasecurityDetailsQuick view
15

Fashion Mnist

VERIFIEDSTRONG · 77TRUST · 79SAFE · EXPERIMENTALRESEARCH

A MNIST-like fashion product database. Benchmark :point_down:

$ npx skills add zalandoresearch/fashion-mnist
12.8K stars57 quality79 trustExperimental4y since pushNeeds review

Scenario Research agents

CLI + Codex · 4 targets

pythonmachine-learning
by zalandoresearchDetailsQuick view
16

VectorDBBench

VERIFIEDEXCELLENT · 90TRUST · 84SAFE · REVIEWEDRESEARCH

Benchmark for vector databases.

$ npx skills add zilliztech/VectorDBBench
1.1K stars60 quality84 trustReviewed2mo since pushSafe to try

Scenario RAG and knowledge

CLI + Codex · 4 targets

pythonvector-search
by zilliztechDetailsQuick view

Page 1

Showing the strongest 16 results to keep the registry fast for humans and agents. Refine by use case, platform, stars, or search query for a narrower shortlist.

Try the agent resolve API