OpenAgentSkill Registry Manifest Skill: AgentBench Slug: thudm-agentbench Category: agent-frameworks Description: A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24) Agent fit: - Decision: 100/100 Production-ready - Primary fit: Coding agents - Role: Primary pick Supply profile: - Track: Coding and developer agents - Scenario: Coding agents - Applicable agents: Claude Code, OpenAI Agents, CLI, Codex, Cursor - Maintenance: 6mo since push - Risk: Safe to try Trust: - Trust score: 87/100 Production candidate - Audit: 89/100 Safe to try Attribution: - Status: Community indexed - Source: GitHub star discovery - Creator: THUDM - Claim URL: https://www.openagentskill.com/skills/thudm-agentbench#claim-this-skill Install: npx skills add THUDM/AgentBench URLs: - Web: https://www.openagentskill.com/skills/thudm-agentbench - API: https://www.openagentskill.com/api/agent/skills/thudm-agentbench - Install API: https://www.openagentskill.com/api/skills/thudm-agentbench/install - Repository: https://github.com/THUDM/AgentBench