暂未收录效果图
查看技能说明design-ai-benchmarking
Aperivue
Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm defini…
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
44 Skills
搜索结果: 44
暂未收录效果图
查看技能说明Aperivue
Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm defini…
暂未收录效果图
查看技能说明cleverhans-lab
An adversarial example library for constructing attacks, building defenses, and benchmarking both
暂未收录效果图
查看技能说明bencherdev
🐰 Bencher - Continuous Benchmarking
暂未收录效果图
查看技能说明rsasaki0109
ROS 2 LiDAR SLAM for pointcloud-map authoring, benchmarking, and Autoware-compatible map workflows.
暂未收录效果图
查看技能说明sangrokjung
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities…
暂未收录效果图
查看技能说明rentruewang
Bayesian Optimization as a Coverage Tool for Evaluating LLMs. Accurate evaluation (benchmarking) that's 10 times faster with just a few lines of modular code.
暂未收录效果图
查看技能说明SpeechColab
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
暂未收录效果图
查看技能说明KevinMusgrave
A library for ML benchmarking. It's powerful.
暂未收录效果图
查看技能说明Benchmarking and evaluation framework for place recognition methods, featuring SuperPoint+SuperGlue, LoGG3D-Net, Scan Context, DBoW2, MixVPR, STD
暂未收录效果图
查看技能说明AgentOps-AI
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK,…
暂未收录效果图
查看技能说明modelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
暂未收录效果图
查看技能说明NanoNets
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
暂未收录效果图
查看技能说明mbzuai-oryx
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrai…
暂未收录效果图
查看技能说明samber
Golang benchmarking, profiling, and performance measurement. Use when writing, running, or comparing Go benchmarks, profiling hot paths with pprof, interpreting CPU/memo…
暂未收录效果图
查看技能说明microsoft
Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
暂未收录效果图
查看技能说明nexscope-ai
Cross-platform ecommerce competitor analysis and strategic intelligence. Multi-channel presence evaluation, strategy assessment, positioning analysis, and competitive be…