AI Agent Skill Repository

AI Agent Skills Directory

Browse reusable skills for Codex, Claude Code, Cursor, finance, research, web scraping, PPT, football analytics, data, marketing, design, and more.

Live registry search

Results for “mle-bench

Exact name and slug matches are checked against the live registry before ranked alternatives.

Clear search

Decision filters

Choose by scenario, quality, and trust signals.

Showing 1-16 of 16 ranked candidates matching "mle-bench"

Best blend of relevance, quality, freshness, and verified outcomes

1

Kube Bench

VERIFIEDEXCELLENT · 99TRUST · 89SAFE · REVIEWEDCODING

Checks whether Kubernetes is deployed according to security best practices as defined in the CIS Kubernetes Benchmark

$ npx skills add aquasecurity/kube-bench
8.1K stars66 quality89 trustReviewed1mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

gokubernetes
by aquasecurityDetailsQuick view
2

MLE Agent

VERIFIEDPROMISING · 66TRUST · 81SAFE · REVIEWEDRESEARCH

🤖 MLE-Agent: Your intelligent companion for seamless AI engineering and research. 🔍 Integrate with arxiv and paper with code to provide better code/research plans 🧰 O…

$ npx skills add MLSysOps/MLE-agent
1.6K stars49 quality81 trustReviewed with permission notesClaude Code + OpenAI Agents1y since pushNeeds review

Scenario RAG and knowledge

Claude Code + OpenAI Agents · 4 targets

pythonmlops
by MLSysOpsDetailsQuick view
3

arbor

EXCELLENT · 92TRUST · 82SAFE · REVIEWEDRESEARCH

Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinemen…

$ npx skills add K-Dense-AI/scientific-agent-skills --skill arbor
34.0K stars55 quality82 trustReviewed with permission notesClaude Code3d since pushSafe to try

Scenario Research agents

Claude Code + CLI · 4 targets

by K-Dense-AIDetailsQuick view
4

Deep Research Bench

STRONG · 70TRUST · 77SAFE · REVIEWEDRESEARCH

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

$ npx skills add Ayanami0730/deep_research_bench
753 stars47 quality77 trustReviewed with permission notes3mo since pushNeeds review

Scenario Research agents

CLI + Codex · 4 targets

pythonresearch-agent
by Ayanami0730DetailsQuick view
5

EnterpriseRAG Bench

PROMISING · 67TRUST · 75SAFE · REVIEWEDRESEARCH

Dataset and benchmark for RAG on company internal documents.

$ npx skills add onyx-dot-app/EnterpriseRAG-Bench
406 stars45 quality75 trustReviewed with permission notes4mo since pushNeeds review

Scenario RAG and knowledge

CLI + Codex · 4 targets

semantic-search
by onyx-dot-appDetailsQuick view
6

Aacr Bench

STRONG · 74TRUST · 79SAFE · REVIEWEDCODING

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.

$ npx skills add alibaba/aacr-bench
209 stars50 quality79 trustReviewed with permission notes19d since pushSafe to try

Scenario Coding agents

CLI + Codex · 4 targets

pythoncode-review
by alibabaDetailsQuick view
7

Space Robotics Bench

PROMISING · 55TRUST · 73SAFE · REVIEWEDCODING

Robot Learning Beyond Earth

$ npx skills add AndrejOrsula/space_robotics_bench
157 stars38 quality73 trustReviewed with permission notes9mo since pushNeeds review

Scenario Coding agents

CLI + Codex · 4 targets

pythonrobotics
by AndrejOrsulaDetailsQuick view
8

MLE Flashcards

VERIFIEDEXCELLENT · 89TRUST · 86SAFE · REVIEWEDCODING

200+ detailed flashcards useful for reviewing topics in machine learning, computer vision, and computer science.

$ npx skills add b7leung/MLE-Flashcards
2.4K stars58 quality86 trustReviewed4mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

computer-vision
by b7leungDetailsQuick view
9

Bench

PROMISING · 68TRUST · 75SAFE · REVIEWEDDESIGN

A tool for evaluating LLMs

$ npx skills add arthur-ai/bench
428 stars45 quality75 trustReviewed with permission notes5mo since pushNeeds review

Scenario Multimodal media

CLI + Codex · 4 targets

typescriptmlops
by arthur-aiDetailsQuick view
10

VEFX Bench

PROMISING · 65TRUST · 75SAFE · REVIEWEDDESIGN

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

$ npx skills add Visko-Platform/VEFX-Bench
215 stars43 quality75 trustReviewed with permission notes3mo since pushNeeds review

Scenario Design and creative

CLI + Codex · 4 targets

pythonimage-generation
by Visko-PlatformDetailsQuick view
11

Chronos

VERIFIEDEXCELLENT · 85TRUST · 81SAFE · REVIEWEDCODING

Kodezi Chronos is a debugging-first language model that achieves state-of-the-art results on SWE-bench Lite (80.33%) and 67% real-world fix accuracy, over six times bett…

$ npx skills add Kodezi/Chronos
4.9K stars57 quality81 trustReviewed with permission notesOpenAI Agents9mo since pushNeeds review

Scenario Coding agents

OpenAI Agents + CLI · 4 targets

javamachine-learning
by KodeziDetailsQuick view
12

SE Agent

PROMISING · 64TRUST · 75SAFE · REVIEWEDCODING

SE-Agent is a self-evolution framework for LLM Code agents. It enables trajectory-level evolution to exchange information across reasoning paths via Revision, Recombinat…

$ npx skills add JARVIS-Xs/SE-Agent
280 stars40 quality75 trustReviewed with permission notesClaude Code11mo since pushNeeds review

Scenario Coding agents

Claude Code + CLI · 4 targets

pythoncoding-agent
by JARVIS-XsDetailsQuick view
13

Mini Swe Agent

VERIFIEDEXCELLENT · 97TRUST · 85SAFE · REVIEWEDCODING

The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench veri…

$ npx skills add SWE-agent/mini-swe-agent
5.3K stars65 quality85 trustReviewed with permission notes2mo since pushSafe to try

Scenario GitHub automation

CLI + Codex · 4 targets

pythonai-agents
by SWE-agentDetailsQuick view
14

Agentic Harness Engineering

STRONG · 79TRUST · 77SAFE · REVIEWEDCODING

Official AHE code — Agentic Harness Engineering: observability-driven automatic evolution of coding-agent harnesses (concurrent w/ meta-harness). NexAU-AHE reaches 84.7%…

$ npx skills add china-qijizhifeng/agentic-harness-engineering
600 stars50 quality77 trustReviewed with permission notesOpenAI Agents2mo since pushSafe to try

Scenario Coding agents

OpenAI Agents + CLI · 4 targets

pythoncoding-agent
by china-qijizhifengDetailsQuick view
15

Zhikuncode

STRONG · 70TRUST · 74SAFE · EXPERIMENTALCODING

Claude Code / Cursor 开源替代。部署在你自己的服务器上,团队用浏览器打开就能编程——包括手机。CLI & Web UI 双入口,Multi-Agent 协作,原生直连千问/DeepSeek 等国产大模型。SWE-bench Lite 56%,不是玩具。技能/插件/跨会话记忆,8 层安全沙箱,数据不离开你的机器。Doc…

$ npx skills add zhikunqingtao/zhikuncode
290 stars48 quality74 trustExperimentalClaude Code + Cursor1mo since pushNeeds review

Scenario Coding agents

Claude Code + Cursor · 4 targets

javamulti-agent
by zhikunqingtaoDetailsQuick view
16

RustyRAG

PROMISING · 65TRUST · 71SAFE · REVIEWEDRESEARCH

Production-grade RAG API built in Rust. Hybrid search with HNSW dense vectors and BM25 sparse matching, cross-encoder reranking, layout-aware document extraction via Doc…

$ npx skills add RustyRAG/RustyRAG
196 stars43 quality71 trustReviewed with permission notes3mo since pushNeeds review

Scenario RAG and knowledge

CLI + Codex · 4 targets

rustdocument-extraction
by RustyRAGDetailsQuick view