Direktori skill

Temukan skill yang dapat digunakan kembali untuk AI agents.

Cari skill GitHub nyata berdasarkan tugas lalu periksa stars, trust, audit, kategori, dan jalur pemasangan sebelum digunakan.

Setiap rekomendasi tetap terhubung dengan repositori, audit, dan jalur pemasangannya.

Hasil pencarian: benchmarking

Direktori bahasa Inggris

Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI

5.8K
Stars
86/100
Kepercayaan
Kategori: agent-frameworksAudit

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

3.0K
Stars
85/100
Kepercayaan
Kategori: dataAudit

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

2.0K
Stars
80/100
Kepercayaan
Kategori: document-processingAudit

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

34K
Stars
80/100
Kepercayaan
Kategori: researchAudit
aeo77

Answer Engine Optimization (AEO) skill โ€” optimize content to be cited by AI language models (ChatGPT, Perplexity, Claude, Gemini, Mistral) as authoritative sources. Distinct from SEO โ€” AEO optimizes for citation in LLM-generated responses, not search rankings. Use when planning content for AI-first search audiences, auditing existing content for E-E-A-T signals, tracking which pages get cited by which LLMs, or building a citation-friendly content strategy. Triggers โ€” 'AEO audit', 'optimize for ChatGPT', 'get cited by Perplexity', 'LLM citation strategy', 'answer engine optimization', 'content for AI search', 'E-E-A-T audit'. Output is a markdown audit report (default) or JSON for pipeline integration. Stdlib-only Python tools.

25K
Stars
77/100
Kepercayaan
Kategori: securityAudit

An adversarial example library for constructing attacks, building defenses, and benchmarking both

6.4K
Stars
74/100
Kepercayaan
Kategori: ml-automationAudit

๐Ÿฐ Bencher - Continuous Benchmarking

855
Stars
67/100
Kepercayaan
Kategori: devopsAudit

ROS 2 LiDAR SLAM for pointcloud-map authoring, benchmarking, and Autoware-compatible map workflows.

823
Stars
73/100
Kepercayaan
Kategori: robotics-iotAudit

[ACL 2024 ๐Ÿ”ฅ] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

1.5K
Stars
73/100
Kepercayaan
Kategori: support-automationAudit

Golang benchmarking, profiling, and performance measurement. Use when writing, running, or comparing Go benchmarks, profiling hot paths with pprof, interpreting CPU/memory/trace profiles, analyzing results with benchstat, setting up CI benchmark regression detection, or investigating production performance with Prometheus runtime metrics. Also use when the developer needs deep analysis on a specific performance indicator - this skill provides the measurement methodology, while `samber/cc-skills-golang@golang-performance` provides the optimization patterns.

3.0K
Stars
67/100
Kepercayaan
Kategori: researchAudit

Windows Agent Arena (WAA) ๐ŸชŸ is a scalable OS platform for testing and benchmarking of multi-modal AI agents.

867
Stars
71/100
Kepercayaan
Kategori: automationAudit

๐Ÿฆ„ Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data for end-to-end AI benchmarking

214
Stars
70/100
Kepercayaan
Kategori: ml-automationAudit