Direktori skill

Temukan skill yang dapat digunakan kembali untuk AI agents.

Cari skill GitHub nyata berdasarkan tugas lalu periksa stars, trust, audit, kategori, dan jalur pemasangan sebelum digunakan.

Setiap rekomendasi tetap terhubung dengan repositori, audit, dan jalur pemasangannya.

Hasil pencarian: benchmark

Direktori bahasa Inggris

A Python library for anomaly detection across tabular, time series, graph, text, and image data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.

9.9K
Stars
86/100
Kepercayaan
Kategori: ml-automationAudit

Checks whether Kubernetes is deployed according to security best practices as defined in the CIS Kubernetes Benchmark

8.1K
Stars
86/100
Kepercayaan
Kategori: devopsAudit

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

6.2K
Stars
77/100
Kepercayaan
Kategori: agent-frameworksAudit

35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider LLM support and a 17-task benchmark leaderboard.

3.7K
Stars
85/100
Kepercayaan
Kategori: agent-frameworksAudit

MTEB: Massive Text Embedding Benchmark

3.3K
Stars
80/100
Kepercayaan
Kategori: rag-knowledgeAudit

A skill for AI agents (Claude Code, Codex, Cursor) that rewrites Traditional Chinese text to remove AI writing patterns, correct China-Taiwan localization, and fix punctuation.

691
Stars
84/100
Kepercayaan
Kategori: utilityAudit

AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.

1.9K
Stars
77/100
Kepercayaan
Kategori: coding-agentsAudit

入门资料整理:1.多因子股票量化框架开源教程 2.学界和业界的经典资料收录 3.AI + 金融的相关工作,包括LLM, Agent, benchmark(evaluation), etc.

1.5K
Stars
79/100
Kepercayaan
Kategori: financeAudit
Evo84

turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.

1.2K
Stars
84/100
Kepercayaan
Kategori: agent-skillsAudit

Benchmark for vector databases.

1.1K
Stars
76/100
Kepercayaan
Kategori: rag-knowledgeAudit

ClickBench: a Benchmark For Analytical Databases

1.0K
Stars
73/100
Kepercayaan
Kategori: data-analysisAudit

A self-learning skill layer for Claude Code that automatically distills, merges, updates, and prunes skills from real sessions.

413
Stars
75/100
Kepercayaan
Kategori: coding-agentsAudit