Annuaire de skills

Découvrez des skills réutilisables pour les AI agents.

Recherchez de vrais skills GitHub par tâche et vérifiez Stars, confiance, audit, catégorie et chemin d’installation avant de les utiliser.

Chaque recommandation reste clairement reliée à son dépôt, son audit et son chemin d’installation.

Résultats de recherche: benchmark

Annuaire en anglais

A Python library for anomaly detection across tabular, time series, graph, text, and image data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.

9.9K
Stars
86/100
Confiance
Catégorie: ml-automationAudit

Checks whether Kubernetes is deployed according to security best practices as defined in the CIS Kubernetes Benchmark

8.1K
Stars
86/100
Confiance
Catégorie: devopsAudit

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

6.2K
Stars
77/100
Confiance
Catégorie: agent-frameworksAudit

35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider LLM support and a 17-task benchmark leaderboard.

3.7K
Stars
85/100
Confiance
Catégorie: agent-frameworksAudit

MTEB: Massive Text Embedding Benchmark

3.3K
Stars
80/100
Confiance
Catégorie: rag-knowledgeAudit

A skill for AI agents (Claude Code, Codex, Cursor) that rewrites Traditional Chinese text to remove AI writing patterns, correct China-Taiwan localization, and fix punctuation.

691
Stars
84/100
Confiance
Catégorie: utilityAudit

AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.

1.9K
Stars
77/100
Confiance
Catégorie: coding-agentsAudit

入门资料整理:1.多因子股票量化框架开源教程 2.学界和业界的经典资料收录 3.AI + 金融的相关工作,包括LLM, Agent, benchmark(evaluation), etc.

1.5K
Stars
79/100
Confiance
Catégorie: financeAudit
Evo84

turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.

1.2K
Stars
84/100
Confiance
Catégorie: agent-skillsAudit

Benchmark for vector databases.

1.1K
Stars
76/100
Confiance
Catégorie: rag-knowledgeAudit

ClickBench: a Benchmark For Analytical Databases

1.0K
Stars
73/100
Confiance
Catégorie: data-analysisAudit

A self-learning skill layer for Claude Code that automatically distills, merges, updates, and prunes skills from real sessions.

413
Stars
75/100
Confiance
Catégorie: coding-agentsAudit