No visual example yet
Explore the skilldesign-ai-benchmarking
Aperivue
Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm defini…
OPENAGENTSKILL / DIRECTORY
次のタスクに合うスキルを。Codex、Claude Code、Cursor などのツールを探せます。
44 Skills
検索結果: 44
No visual example yet
Explore the skillAperivue
Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm defini…
No visual example yet
Explore the skillcleverhans-lab
An adversarial example library for constructing attacks, building defenses, and benchmarking both
No visual example yet
Explore the skillbencherdev
🐰 Bencher - Continuous Benchmarking
No visual example yet
Explore the skillrsasaki0109
ROS 2 LiDAR SLAM for pointcloud-map authoring, benchmarking, and Autoware-compatible map workflows.
No visual example yet
Explore the skillsangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
No visual example yet
Explore the skillrentruewang
Bayesian Optimization as a Coverage Tool for Evaluating LLMs. Accurate evaluation (benchmarking) that's 10 times faster with just a few lines of modular code.
No visual example yet
Explore the skillSpeechColab
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
No visual example yet
Explore the skillKevinMusgrave
A library for ML benchmarking. It's powerful.
No visual example yet
Explore the skillBenchmarking and evaluation framework for place recognition methods, featuring SuperPoint+SuperGlue, LoGG3D-Net, Scan Context, DBoW2, MixVPR, STD
No visual example yet
Explore the skillAgentOps-AI
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK,…
No visual example yet
Explore the skillmodelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
No visual example yet
Explore the skillNanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
No visual example yet
Explore the skillmbzuai-oryx
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrai…
No visual example yet
Explore the skillsamber
Golang benchmarking, profiling, and performance measurement. Use when writing, running, or comparing Go benchmarks, profiling hot paths with pprof, interpreting CPU/memo…
No visual example yet
Explore the skillmicrosoft
Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
No visual example yet
Explore the skillnexscope-ai
Cross-platform ecommerce competitor analysis and strategic intelligence. Multi-channel presence evaluation, strategy assessment, positioning analysis, and competitive be…