Evalscope
modelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–10 / 10
Results: 10
modelscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
ablab
votchallenge
The official VOT Challenge evaluation and analysis toolkit
VoltAgent
A curated collection of AI agent research papers released in 2026, covering agent engineering, memory, evaluation, workflows, and autonomous systems.
certsocietegenerale
alibaba
A CLI evaluation framework to make your Agent Skill Up.
probabl-ai
Track your Data Science. Skore's open-source Python library accelerates ML model development with automated evaluation reports, smart methodological guidance, and compre…
ApodexAI
Evaluation harness for Apodex-1.0 on public deep-research benchmarks.
Forward-Future
Practical repeatable AI-agent workflows for engineering, evaluation, operations, content, and design.
texttron
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)